Source Analysis
Read the source website before writing a spider - DevTools first, JSON APIs before CSS selectors.
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortxSource analysis
Research comes before code. The right parse strategy is determined by what the source site actually does, not by the first thing that works.
DevTools first
Open the agency's meeting page and inspect the Network tab (filter XHR and Fetch) while the page loads:
- A JSON or XML endpoint feeding the page is the best outcome. APIs return typed fields; HTML parsing is brittle by comparison.
- An iCalendar feed (
.ics, RFC 5545) is nearly as good - the structure is standardized. - Rendered HTML only means CSS selectors, and a more brittle spider.
Use Postman or curl to probe the endpoint
directly - pagination params, filters, and required headers are much easier to
figure out there than in spider code.
JavaScript-heavy agency sites are usually JAM-stack frontends over a JSON API. The network tab almost always reveals it. Prefer the API over parsing the rendered DOM.
Primary and secondary URLs
Many sources split meeting data across pages:
- The primary URL is where the meeting list lives - the listing page your spider starts from.
- Secondary URLs - detail pages, document pages - supplement the primary data. Treat them as enrichment: the spider should degrade gracefully when a secondary page is missing or changes shape.
When parse yields Request objects for detail pages, pass the listing-page
data through with cb_kwargs and attach errback=self.handle_error so a dead
detail page does not kill the whole run. See
Testing for how to test this pattern.
Extraction strategy, in order of preference
Parse the endpoint response with response.json(). Fields arrive typed;
dates, titles, and document URLs need no selector guesswork. This is the
pattern for JS-rendered agency sites and the least brittle option.
def parse(self, response):
for item in response.json()["events"]:
yield self._parse_meeting(item)What to write down
Before moving on, note:
- The primary URL and any secondary URLs.
- The extraction mode (API / CSS / hybrid) and why.
- Known edge cases in the source data: variant title spellings, missing times, meetings listed in PDFs only. These belong in the scraper's spec and its tests - not discovered in review.
- Whether the site shows signs of bot detection.
Last updated on