Skip to content

Methodology

This public site is generated from completed report artifacts, not from direct raw-source reads.

Pipeline

  • Upstream source data is normalized into curated DuckDB and Parquet outputs.
  • Closed weekly and monthly report artifacts are generated from curated storage.
  • Site generation copies only public-safe CSV downloads and HTML report snapshots into the public site source.

Public boundary

  • Row-level public downloads exclude company, URL, raw descriptions, raw job IDs, and raw OpenAI payloads.
  • Aggregate report content may reference company names only in aggregated ranking sections.
  • Private drill-down remains local-only in the Streamlit app.