Configuration
This project supports typed programmatic config and TOML-driven runtime config.
Programmatic Library Config
JobScraperConfig
JobScraperConfig captures the runtime inputs for a single scrape:
position: job title or search phraselocation: LinkedIn location textenrichment_provider:none,openai, oranthropic, defaultnoneenrichment_model: model name for the selected provider, defaults togpt-4o-minifor OpenAI andclaude-haiku-4-5-20251001for Anthropictime_posted:TimePostedenum value or matching stringremote:RemoteTypeenum value or matching stringdistance: search radiusadvanced_config: optionalJobScraperAdvancedConfigoverrides
JobScraperAdvancedConfig
Use it for optional overrides such as:
LOCATION_MAPPINGKEYWORDSSKILLS_CATEGORIES
Runtime Enums
TimePosted:ALL,MONTH,WEEK,DAYRemoteType:ALL,ON-SITE,REMOTE,HYBRID
Runtime TOML Config
The tracked template is runtime.example.toml. A real runtime.toml can define:
[logging]
level = "INFO"
file_name = "main.log"
[storage]
file_name = "linkedin_jobs.sqlite"
state_dir = "artifacts/state"
[scrape.once]
position = "Data Scientist"
location = "Monterrey"
enrichment_provider = "none"
enrichment_model = "gpt-4o-mini"
time_posted = "DAY"
remote_types = ["REMOTE", "HYBRID", "ON-SITE"]
file_name = "LinkedIn_Jobs_Data_Scientist_Monterrey.csv"
output_dir = "artifacts/jobs"
append = true
[scrape.daily]
cities = ["Monterrey", "Guadalajara", "Mexico City"]
position = "Data Scientist"
enrichment_provider = "none"
enrichment_model = "gpt-4o-mini"
time_posted = "DAY"
output_dir = "artifacts/jobs"
combined_file_name = "LinkedIn_Jobs_Data_Scientist_Mexico.csv"
[export]
run_id = ""
file_name = "linkedin_jobs_export.csv"
output_dir = "artifacts/jobs"
CLI Override Precedence
The runtime precedence is:
- CLI flags
- environment overrides
- TOML file values
- code defaults
Supported env overrides include:
LINKEDIN_WEB_SCRAPER_CONFIGLINKEDIN_WEB_SCRAPER_LOG_LEVELLINKEDIN_WEB_SCRAPER_LOG_FILELINKEDIN_WEB_SCRAPER_STORAGE_URLLINKEDIN_WEB_SCRAPER_STORAGE_FILELINKEDIN_WEB_SCRAPER_STATE_DIRLINKEDIN_WEB_SCRAPER_OUTPUT_DIRLINKEDIN_WEB_SCRAPER_ENRICHMENT_PROVIDERLINKEDIN_WEB_SCRAPER_ENRICHMENT_REQUIREDLINKEDIN_WEB_SCRAPER_ENRICHMENT_MODEL
Enrichment Runtime Behavior
Enrichment is optional and provider-selectable through enrichment_provider (none, openai, or anthropic).
- Install
LinkedInWebScraper[openai]orLinkedInWebScraper[anthropic]to match the selected provider - Set
OPENAI_API_KEYorANTHROPIC_API_KEYin the environment, matching the selected provider - Keep secrets out of TOML and out of the repo
- By default, enrichment setup or request failures fall back to the non-enriched dataset; set
enrichment_required = trueto fail the run instead - Enriched rows include audit fields such as
EnrichmentModel,EnrichmentResponseId, andEnrichmentRawPayload
Artifact And State Paths
Managed defaults resolve under artifacts/:
- bare CSV file names ->
artifacts/jobs/ - bare log file names ->
artifacts/logs/ - bare SQLite file names ->
artifacts/state/
Explicit absolute paths and explicit nested relative paths are preserved.
Storage Model
SQLite persistence is enabled by default for CLI and DailyScrapeService workflows.
scrape_runsstores run lifecycle metadatajobsstores the latest canonical job attributesjob_snapshotsstores the run-level dataframe payloadjob_enrichmentsstores structured OpenAI audit data when present- CSV exports are downstream artifacts written from persisted data
Use build_sqlite_storage_url() for a managed default URL, or inject SQLiteScrapeStorage(storage_url=...) into DailyScrapeService when you need a custom local path or DSN.
mssql / Azure SQL
SQLiteScrapeStorage runs against any SQLAlchemy-supported dialect, not only SQLite.
Point it at Azure SQL (or any mssql server) by setting LINKEDIN_WEB_SCRAPER_STORAGE_URL
to a mssql+pyodbc:// URL, or by passing that URL as runtime_config.storage.url /
storage_url=. Install the pyodbc Python package with:
pip install LinkedInWebScraper[mssql]
pyodbc also needs the system-level Microsoft ODBC Driver 18 for SQL Server
(msodbcsql18) installed and on the path, or connections fail with an
"ODBC Driver 18 ... not found" error. On Ubuntu/Debian:
curl https://packages.microsoft.com/keys/microsoft.asc | sudo tee /etc/apt/trusted.gpg.d/microsoft.asc
curl https://packages.microsoft.com/config/ubuntu/22.04/prod.list | sudo tee /etc/apt/sources.list.d/mssql-release.list
sudo apt-get update
ACCEPT_EULA=Y sudo apt-get install -y msodbcsql18 unixodbc-dev
See Microsoft's ODBC driver install docs for other platforms.
Text columns use SQLAlchemy's generic Unicode/UnicodeText types so mssql renders
them as NVARCHAR, keeping non-ASCII text (Spanish job postings, for example)
round-tripping correctly. The engine is created with pool_pre_ping=True so a stale
connection, for example an Azure SQL serverless tier resuming from auto-pause, is
retried transparently instead of surfacing as a connection error.
Base.metadata.create_all() provisions the schema on first use; there is no Alembic
migration path yet, schema changes require a fresh database or a hand-run migration.
Root Runtime Scripts
main.py and process_ds_jobs.py remain available as direct runtime entrypoints for the daily and once workflows.