Pipeline
How scraper output gets from a spider run to the Documenters platform - scheduled runs, Azure storage, and the Airtable registry.
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortxPipeline
The pipeline covers everything that happens to scraper output after you merge a spider: scheduled runs, storage, error reporting, and the Airtable registry that tracks which spiders exist.
The production flow
Scheduled run - GitHub Actions workflows in each city-scrapers repo run every spider on a daily cron trigger.
Storage - each spider's JSON output is uploaded to Azure Blob Storage containers.
Error reporting - failures are reported to Sentry.
Archival - a secondary workflow submits scraped URLs to the Internet Archive's Wayback Machine, creating a permanent public record of the source pages at scrape time.
New spiders added to city_scrapers/spiders/ are picked up by the scheduled
workflow automatically - no workflow changes required.
The Airtable registry
Alongside the run pipeline, a separate automation keeps the team's Airtable
spider registry in sync: when a PR touching spider files opens, a workflow
extracts each spider's name and agency via AST parsing and upserts the
records. See Airtable sync.
Pages
- Scheduled runs - the cron workflows, Azure storage, Sentry, and Wayback archival.
- Airtable sync - the two-workflow trust boundary that registers new spider slugs.
Last updated on