Methodology
This public site is generated from completed report artifacts, not from direct raw-source reads.
Pipeline
- Upstream source data is normalized into curated DuckDB and Parquet outputs.
- Closed weekly and monthly report artifacts are generated from curated storage.
- Site generation copies only public-safe CSV downloads and HTML report snapshots into the public site source.
Public boundary
- Row-level public downloads exclude company, URL, raw descriptions, raw job IDs, and raw OpenAI payloads.
- Aggregate report content may reference company names only in aggregated ranking sections.
- Private drill-down remains local-only in the Streamlit app.