Building Scrapers
The development workflow for writing city-meeting scrapers, from source analysis to PR.
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortxBuilding scrapers
City Scrapers uses Scrapy to extract public meeting information from government websites. Each spider produces a JSON array of meeting records that flows into Documenters.org - and that you can inspect in the Meetings Viewer while you develop.
The workflow
Scope the target - decide whether the agency's site is worth scraping at all, and whether anyone has already scoped it. See Scoping a target.
Analyze the source - open the agency's meeting page in DevTools, check the Network tab for a JSON endpoint, and pick a primary URL. See Source analysis.
Pick a spider pattern - LegistarSpider for Legistar sites, a spider
factory for many similar agencies, or a singular CityScrapersSpider for
everything else. See Spider patterns.
Write the spider - implement parse and the _parse_* helpers, following
the schema field by field. See
Writing a spider.
Run and inspect - scrapy crawl <name> -O <name>.json, then load the file
in the viewer and apply the QA rubric.
Test and submit - write pytest tests against saved fixtures, run flake8/black/isort, open a PR. See Testing and Contributing.
The tight loop is steps 4 and 5: run the spider, refresh the viewer, fix, repeat. The viewer exists to make that loop fast.
Common pitfalls
- Lowercase
-o: silently appends to an existing file, duplicating records. Always use uppercase-O. - Timezone-aware datetimes:
startandendmust be naive. The spider'stimezoneclass attribute declares the local timezone; never embedtzinfoin the datetime itself. - Missing time defaults: if a source has a date but no time, default to
midnight and say so in
time_notes. - Links out of order: the priority is agenda, minutes, video, then other attachments - not agenda, video, minutes.
async def start(): consumer repos pin Scrapy 2.11.2, which does not recognize the newerstart()coroutine. Usestart_requests().
Section map
- Scoping a target - is this site worth scraping at all?
- Source analysis - read the source before writing code.
- Spider patterns - singular spiders, spider factories, mixins.
- Writing a spider - the field-by-field parse implementation.
- Bot detection - when the site fights back.
- Scraper types - the taxonomy with real examples.
- Testing - fixtures, freeze_time, two-phase tests.
- Contributing - fork, pipenv, lint, PR.
Last updated on