Conversation
…ata scale
Three related Firestore document/field size-limit bugs surfaced when the
weekly incremental scraper was run for the first time against the full
production-scale dataset (300K+ filings):
1. The live weekly cursor (scrapers/lobbying) stored the entire processed-URL
history and summary cache as two fields on one document. That document
exceeded Firestore's 1MB limit partway through a run, silently failing
(and thus skipping) every registrant processed afterward. Moved to
subcollections — one small doc per URL — mirroring the pattern the
backfill cursor already used, with point lookups instead of an in-memory
set/dict.
2. compute_stats() streamed the full lobbyingFilings/lobbyingRegistrants
collections (300K+ docs) in one unbounded query, which timed out
server-side; the client library's automatic stream-retry then crashed on
an internal AttributeError instead of recovering. Replaced with
cursor-paginated batches (50K docs/request) and a manual retry that
re-issues a fresh query rather than resuming a broken stream.
3. Once (2) was fixed, compute_stats() reached a third limit: the
billSummaries_{court} JSON blob itself exceeded Firestore's 1MB
field-size limit for the current session (1,057KB for court 194's ~5,600
bills), with courts 192/193 close behind. Restructured to one small doc
per bill in a bills subcollection instead of one JSON blob per court —
same fix pattern as (1), applied to writer.py, seedLobbyingStats.ts, and
the frontend fetcher in components/db/lobbying.ts.
All three fixes validated end-to-end against dev Firestore at current scale
(373K filings, 25.6K registrants, 11 courts including the previously-failing
194th).
Also includes scripts/firebase-admin/checkLobbyingFreshness.ts, a read-only
diagnostic for checking scraper cursor state and data recency.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
run_backfill() marked a year "complete" after one pass and skipped it forever on every future run. That's wrong for the current (still-accruing) year: a run partway through the year would mark it complete based on whatever existed at that moment, silently missing every disclosure filed afterward — no future backfill run would ever see it again. This is exactly what happened to 2026 in production: marked complete in July with 0 disclosures captured. run_backfill already has a fully correct, granular completeness check — _is_backfill_processed, a per-URL subcollection lookup. The year-level flag only ever bought a coarse fast-path (skip re-listing a year's registrants entirely) and it's what caused the bug. Removed it: every run now always re-lists every requested year (one cheap HTTP request per year) and relies solely on the per-URL cursor for correctness, so no year can ever be skipped wholesale again. Added tests/test_scrape.py with a small in-memory Firestore fake (real enough to simulate write-then-read-back across calls, unlike a plain mock) covering both cursor systems: - Regression test reproducing the exact bug scenario (empty pass, then real data appears for the same year) — fails against the old code with 4/4 backfill tests red, passes with the fix, confirmed by checking out the pre-fix scrape.py and rerunning the suite against it. - Backfill always re-lists every year, per-URL dedup still works, dry-run never touches Firestore, completedYears is never written anywhere. - Weekly-mode cursor sanity checks (prior-year caching, current-year always live, parent doc stays small — all state in subcollections). Also validated live against dev: --mode backfill --year 2005 --limit 2 run twice confirms the year is re-listed both times while already-processed disclosures are correctly not reprocessed (0 new on both runs, as expected since 2005 was already backfilled in July). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ction
The billSummaries_{court} restructure (previous commit) moved per-bill
counts from a single document field to a `bills` subcollection, but the
existing `match /lobbyingMeta/{id}` rule only covers documents directly in
that collection — Firestore rules aren't recursive, so it never covered the
new subcollection. The bills index page was failing with "Missing or
insufficient permissions" as a result; caught on a preview deployment,
since all earlier validation ran through the Admin SDK, which bypasses
security rules entirely.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…rminology, fix incomplete client/firm lists Four related frontend fixes to the lobbying explorer: 1. Embedded lobbying card (shown on regular MAPLE bill pages) previously capped at 5 filings with a "view all" link out to a separate page. Now shows the full paginated list (10/page) inline, matching the pagination pattern already used elsewhere in the explorer. 2. Client and lobbyist names in the shared LobbyingFilingsTable were plain text everywhere it's used (the embed, and the firm/client detail pages). Now clickable, linking to their respective detail pages — same href pattern already used on the lobbying explorer's own bill detail page. 3. Renamed "Lobbying Firm(s)" to "Lobbyist(s)" throughout the visible UI (stat cards, column headers, page titles, explainer text) — matches the MA Secretary of State's own "Lobbyist Public Search" terminology, which this data is sourced from. Left the /lobbying/firms URL and internal file/variable names unchanged (route stability). 4. Audited data loading across the lobbying feature. Found a real correctness bug, not just a performance one: the clients and firms index pages, and the client detail page's "linked firms" list, all read via useLobbyingAllRegistrants(), which caps at 2,000 of 25,000+ registrant docs (Firestore query limit) — silently showing an incomplete list. Confirmed at the actual data: only 4,815 distinct firms surface via the old registrant scan vs. 7,360 that exist dataset-wide. Fixed by precomputing per-client and per-firm summary docs server-side (writer.py's compute_stats(), mirrored in seedLobbyingStats.ts) over the full, paginated registrants collection — same one-small-doc-per-item subcollection pattern already used for billSummaries, needed here too since ~5,300 clients and ~7,360 firms are already close to Firestore's 1MB single-document limit as flat blobs. Client summary docs also carry a per-firm compensation breakdown, so the client detail page can do a single-document lookup instead of scanning all registrants. Verified live: dev Firestore reseeded, all four changes checked in an actual browser (Playwright against the running dev server) — pagination controls, clickable links, "Lobbyists" terminology, and the corrected 5,135 clients / 7,360 firms counts, zero console errors. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
nesanders
marked this pull request as ready for review
September 16, 2026 00:11
nesanders
requested review from
Mephistic,
alexjball,
kiminkim724,
mertbagt,
mvictor55,
sashamaryl and
timblais
as code owners
September 16, 2026 00:11
Collaborator
Author
|
Additional feedback captured during hack night in discussion with Grace:
|
…agination Follow-up visual/UX fixes to the embedded lobbying card (LobbyingBillCard) after the earlier pagination/link changes: - Table background was inheriting Bootstrap's --bs-table-bg from --bs-body-bg (MAPLE's page background, #eae7e7), reading as a heavy grey bleed-through inside the white card. Overridden to var(--bs-gray-100), the correct way to theme a single Bootstrap 5 table instance via its own CSS custom properties rather than fighting cell-level specificity. - Row hover was similarly inheriting a Bootstrap default; set explicitly to a 6% opacity dark tint (rgba(15, 23, 42, 0.06)) instead of a flat color, consistent with the subtle shadow tone already used elsewhere in this component. - Added an optional `bordered` prop to LobbyingFilingsTable (opt-in, default off, so firm/client detail pages are unaffected) and enabled it on the embedded card with var(--bs-border-color). - The "page X / Y" pagination indicator looked like a disabled button but wasn't interactive. Replaced with an actual editable text input — typing a number and pressing Enter/blurring jumps to that page, clamped to [1, totalPages]. This is the shared LobbyingPaginationBar component, so the improvement applies everywhere it's used, not just the embedded card. - Added an optional `itemLabel` prop to LobbyingPaginationBar so a page can append a label after the "X–Y of Z" count (e.g. "Lobbying filings") without hardcoding page-specific text into the shared component. - Card header: replaced the top-right filing count with the "Open in Lobbying Explorer" link (now visible without scrolling to the bottom of a paginated table); the count moved down to the pagination bar's "of Z Lobbying filings" text instead. Title changed from the shared "Lobbying Explorer" string (used by the overview page's <h1>) to a new, card-specific "Lobbying Activity" label. Verified live: dev server + headless browser — confirmed computed background (#f8f9fa), border color (#dee2e6), hover tint, typing a page number to jump (including out-of-range clamping to totalPages), and zero console errors. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Collaborator
Author
|
Most recent commit implements:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Four fixes to the lobbying explorer:
/lobbying/firmsURL and internal names unchanged.useLobbyingAllRegistrants(), which caps at 2,000 of 25,000+ registrant docs (Firestore query limit) — silently showing an incomplete list. Confirmed against real data: only 4,815 distinct firms surfaced via the old registrant scan vs. 7,360 that actually exist. Fixed by precomputing per-client and per-firm summary docs server-side instead.Checklist
firestore.indexes.json. — N/A, the new queries are unfiltered full-subcollection reads (no composite index needed)Screenshots
Updated lobbying card on bill page
Now features clickable entity links, pagination, and renames entity type string to "lobbyist". Updates background color, makes pagination page editable, retitles card, moves lobbying explorer link.
Lobbyist naming on lobbying explorer
Entity renamed from Lobbying Firms to Lobbyist
Also note: shows the correct 7360 number as described in statistics fixes below.
Statistics fixes
/lobbying/clients/1199SEIU): a "LOBBYISTS" sidebar section listing all firms that represented this client with compensation, all clickable; bills table shows clickable client/lobbyist links./bills/194/H4000): "1–10 of 582" with working Prev/Next controls, clickable client and lobbyist links, "Open in Lobbying Explorer →" link at the bottom.Known issues
compute_stats()/seedLobbyingStatsrun (same as the existingbillSummariesdocs) — fine at current scale (~5,300 clients / ~7,400 firms), but worth watching as the dataset grows.Steps to test/reproduce
yarn firebase-admin run-script seedLobbyingStats --env dev(or wait for the next scrape run) to populate the newlobbyingMeta/clientSummariesandlobbyingMeta/firmSummariessubcollections.firestore.rulesto dev./lobbying/clients— table should show real names, "Lobbyists" (not "Firms") column header, correct total count in the "N clients" line above the table./lobbying/firms— same check; page title/subnav should read "Lobbyists"./bills/194/H4000) and scroll to the lobbying card — should show all filings paginated 10/page (not capped at 5), with clickable client and lobbyist links, and an "Open in Lobbying Explorer →" link at the bottom.