Writing a Spider
How to write a Scrapy spider that produces valid meeting records for the Meetings Viewer.
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortxWriting a spider
Check the pinned Scrapy version in the repo's Pipfile before using any API
you saw in a recent tutorial. Consumer repos pin Scrapy 2.11.2 - the newer
async def start() coroutine is silently skipped by that version. Use
start_requests().
Spider structure
Every spider extends CityScrapersSpider (or LegistarSpider for Legistar
sites) and implements a parse method that yields meeting records. See
Spider patterns for the architecture choice.
from city_scrapers_core.spiders import CityScrapersSpider
from city_scrapers_core.items import Meeting
class SandieCityCouncilSpider(CityScrapersSpider):
name = "sandie_city_council"
agency = "San Diego City Council"
timezone = "America/Los_Angeles"
start_urls = ["https://example.gov/meetings"]
def parse(self, response):
for item in response.css(".meeting-item"):
meeting = Meeting(
title=self._parse_title(item),
description=self._parse_description(item),
classification=self._parse_classification(item),
start=self._parse_start(item),
end=self._parse_end(item),
all_day=self._parse_all_day(item),
time_notes=self._parse_time_notes(item),
location=self._parse_location(item),
links=self._parse_links(item),
source=response.url,
id=self._parse_id(item),
status=self._parse_status(item),
)
yield meetingField-by-field guide
The authoritative field list and expected behavior live in the schema reference. Notes below cover the fields that cause the most review comments.
title
The official meeting name including the body - "City Council Regular Meeting", not "Regular Meeting". Strip embedded dates and times: a title like "Board Meeting - January 15, 2026 at 6:00 PM" needs normalizing down to the body name.
description
Usually an empty string. Populate only when the source provides one - do not synthesize a description from the title.
start and end
Timezone-naive datetime objects. Never embed tzinfo - the spider's
timezone class attribute declares the local timezone for the framework.
startis required. If the source has a date but no time, default to midnight and explain intime_notes("see agenda for time").endisNonein most cases. The framework defaults it tostart + 2 hours; set it explicitly only when the source provides a confirmed end time.- For date comparisons against meeting times, use
datetime.now(tz=ZoneInfo(self.timezone))- notdatetime.utcnow()(deprecated) and notdatetime.now()without a timezone, which diverges between your machine and CI runners. - Never hardcode year ranges in
range()calls - they break silently at rollover.
status
One of passed, tentative, confirmed, cancelled. The usual pattern is
self._get_status(meeting), which derives passed/tentative from start
relative to now, plus optional cancellation text the method scans for.
cancelledrequires explicit evidence on the source - a missing meeting is not sufficient evidence.- Account for both spellings on source sites: "cancelled" and "canceled".
location
An object with name (room/floor/building) and address (full street
address). Empty strings are acceptable when unknown - say so in time_notes
and point readers at the agenda attachment.
links
A list of { href, title } dicts in priority order:
- Agenda PDF
- Minutes PDF
- Video / recording
- Other attachments
Always prefer the PDF format; include HTML or other formats when no PDF
exists. The title describes the document type ("Agenda", "Minutes",
"Video").
id
A unique identifier within the spider's output. The base class's _get_id
builds one from the meeting fields - use it unless the source has a stable
native ID.
source
The URL of the page the meeting came from. Prefer the meeting's detail page; the listing page URL is acceptable when no detail page exists.
Error handling
- Log, don't swallow: every
exceptblock should emit at least alogger.warning(...)with enough context to diagnose the failure later. A barepassproduces a spider that exits cleanly and yields nothing. - Attach
errback=self.handle_errortoscrapy.Requestcalls, especially detail pages. - Compile regex patterns at class level, not inside per-item methods - the patterns get recompiled on every call otherwise.
- Use explicit keys for deduplication, not array position.
Running your spider
scrapy crawl {spider} -O data/scrapers/{spider}.jsonSuppress Scrapy's INFO logs to see only your own output:
scrapy crawl {spider} -O test.json -s LOG_LEVEL=WARNINGThen validate the output against the schema before writing tests:
scrapy validate {spider}Load the result in the Meetings Viewer and check it against the QA rubric - or paste the JSON straight into the rubric report on that page.
Last updated on