Building Scrapers

Writing a Spider

How to write a Scrapy spider that produces valid meeting records for the Meetings Viewer.

Repo
city-scrapers (core + consumer repos)city-scrapers-core and the per-city repos like city-scrapers-fortx

Writing a spider

Check the pinned Scrapy version in the repo's Pipfile before using any API you saw in a recent tutorial. Consumer repos pin Scrapy 2.11.2 - the newer async def start() coroutine is silently skipped by that version. Use start_requests().

Spider structure

Every spider extends CityScrapersSpider (or LegistarSpider for Legistar sites) and implements a parse method that yields meeting records. See Spider patterns for the architecture choice.

from city_scrapers_core.spiders import CityScrapersSpider
from city_scrapers_core.items import Meeting


class SandieCityCouncilSpider(CityScrapersSpider):
    name = "sandie_city_council"
    agency = "San Diego City Council"
    timezone = "America/Los_Angeles"
    start_urls = ["https://example.gov/meetings"]

    def parse(self, response):
        for item in response.css(".meeting-item"):
            meeting = Meeting(
                title=self._parse_title(item),
                description=self._parse_description(item),
                classification=self._parse_classification(item),
                start=self._parse_start(item),
                end=self._parse_end(item),
                all_day=self._parse_all_day(item),
                time_notes=self._parse_time_notes(item),
                location=self._parse_location(item),
                links=self._parse_links(item),
                source=response.url,
                id=self._parse_id(item),
                status=self._parse_status(item),
            )
            yield meeting

Field-by-field guide

The authoritative field list and expected behavior live in the schema reference. Notes below cover the fields that cause the most review comments.

title

The official meeting name including the body - "City Council Regular Meeting", not "Regular Meeting". Strip embedded dates and times: a title like "Board Meeting - January 15, 2026 at 6:00 PM" needs normalizing down to the body name.

description

Usually an empty string. Populate only when the source provides one - do not synthesize a description from the title.

start and end

Timezone-naive datetime objects. Never embed tzinfo - the spider's timezone class attribute declares the local timezone for the framework.

  • start is required. If the source has a date but no time, default to midnight and explain in time_notes ("see agenda for time").
  • end is None in most cases. The framework defaults it to start + 2 hours; set it explicitly only when the source provides a confirmed end time.
  • For date comparisons against meeting times, use datetime.now(tz=ZoneInfo(self.timezone)) - not datetime.utcnow() (deprecated) and not datetime.now() without a timezone, which diverges between your machine and CI runners.
  • Never hardcode year ranges in range() calls - they break silently at rollover.

status

One of passed, tentative, confirmed, cancelled. The usual pattern is self._get_status(meeting), which derives passed/tentative from start relative to now, plus optional cancellation text the method scans for.

  • cancelled requires explicit evidence on the source - a missing meeting is not sufficient evidence.
  • Account for both spellings on source sites: "cancelled" and "canceled".

location

An object with name (room/floor/building) and address (full street address). Empty strings are acceptable when unknown - say so in time_notes and point readers at the agenda attachment.

A list of { href, title } dicts in priority order:

  1. Agenda PDF
  2. Minutes PDF
  3. Video / recording
  4. Other attachments

Always prefer the PDF format; include HTML or other formats when no PDF exists. The title describes the document type ("Agenda", "Minutes", "Video").

id

A unique identifier within the spider's output. The base class's _get_id builds one from the meeting fields - use it unless the source has a stable native ID.

source

The URL of the page the meeting came from. Prefer the meeting's detail page; the listing page URL is acceptable when no detail page exists.

Error handling

  • Log, don't swallow: every except block should emit at least a logger.warning(...) with enough context to diagnose the failure later. A bare pass produces a spider that exits cleanly and yields nothing.
  • Attach errback=self.handle_error to scrapy.Request calls, especially detail pages.
  • Compile regex patterns at class level, not inside per-item methods - the patterns get recompiled on every call otherwise.
  • Use explicit keys for deduplication, not array position.

Running your spider

scrapy crawl {spider} -O data/scrapers/{spider}.json

Suppress Scrapy's INFO logs to see only your own output:

scrapy crawl {spider} -O test.json -s LOG_LEVEL=WARNING

Then validate the output against the schema before writing tests:

scrapy validate {spider}

Load the result in the Meetings Viewer and check it against the QA rubric - or paste the JSON straight into the rubric report on that page.

Last updated on