Skip to content

feat(pesacheck_meedan_bridge): Make the source CMS configurable (Ghost + Superdesk) - #1211

Open
koechkevin wants to merge 4 commits into
mainfrom
feat/pesacheck-provider-switch
Open

koechkevin wants to merge 4 commits into
mainfrom
feat/pesacheck-provider-switch

Conversation

@koechkevin

@koechkevin koechkevin commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Why

PesaCheck is moving off Ghost: the new site is built on Superdesk + Publisher, read over a Hasura GraphQL API (pesacheck-ui). The bridge was Ghost-shaped throughout — fetching, parsing, summary extraction and language detection were all inlined in main.py.

After this, the CMS is a config choice, so the cutover is an env change plus a restart, and rollback is flipping it back.

What changed

A provider interface

main.py no longer knows about any CMS. Each provider exposes name, fetch(since, limit, max_articles) returning raw records, and parse(record) returning a normalized Article (guid, title, url, published_at, plain-text summary, categories, language, author, thumbnail).

File Role
provider_base.py Article + the shared User-Agent + html_to_text
provider_ghost.py today's Content API behaviour
provider_superdesk.py Publisher's GraphQL API
providers.py registry + get_provider()

Each provider validates its own settings, so running Superdesk no longer requires a Ghost API key (settings.py used to demand it at import).

Providers are flat modules, not a providers/ package: py/BUILD globs *.py in that directory only, and a subpackage would need its own BUILD file plus correct pex dependency inference — unverifiable without Pants locally, and a mistake would surface as an ImportError in the deployed cron rather than at build time.

Superdesk mapping

One query against swp_article, filtered by tenant and to articles carrying a Debunk verdict — that clause is what separates fact-checks from homepage blocks, team profiles and Media Centre entries on the same routes. metadata arrives as a JSON-encoded string and is parsed.

Bridge field Source
guid metadata.guid (survives a re-import), falling back to id
url PESACHECK_ARTICLE_URL_TEMPLATE over {site}, {desk}, {slug}
summary lead, HTML-stripped
language metadata.language — no more guessing from tag names
categories subject names for Debunklang, countries, content_type, Harm_type

Cutover safety (from review)

Two problems with switching provider, both fixed:

  1. The first Superdesk run would have re-posted fact-checks Check already has. With no checkpoint it fetched the newest page and posted it. Migrated articles carry new guids, so feed_exists doesn't match, and Check does not reject them: its signature is MD5([set_fact_check.to_json, team_id]) (project_media_creators.rb) and set_fact_check carries the URL, which differs between the two sites. That would have published duplicate reports. Anything published beyond one page since the last Ghost run would also have been skipped.

    A provider with no articles of its own now starts from the newest article stored under any source — where the previous provider stopped. This relies on migrated articles keeping their original published_at, which they do: staging holds fact-checks back to 2018-11-15.

  2. That fallback can reach far back, so catch-up is now bounded. Articles are fetched oldest first when catching up (newest page when there's no checkpoint at all), so a partial run still advances the checkpoint, and a run stops at PESACHECK_MAX_ARTICLES (default 100) and continues on the next one. Without this, a stale checkpoint walked years of archive in a single run: a local test with only 2024-era rows paginated ~2 years of Ghost posts and would have posted all of them to Check.

The new tests also caught a latent crash: get_checkpoint() used UTC, which main.py stopped importing during the refactor, so any legacy Medium row (naive timestamps) raised NameError.

Article URL

Publisher has no canonical-URL field — swp_redirect_route exists but is empty — so the bridge builds the URL, and the shape matters beyond the link: Check's duplicate signature covers it, so a new shape makes already-imported articles look new.

PESACHECK_ARTICLE_URL_TEMPLATE defaults to {site}/{slug}/, the Ghost-era shape every fact-check already in Check links to, and switches to {site}/fact-checks/{desk}/{slug} with one setting once the new site serves that canonically. Publisher slugs match the Ghost slugs today (two live URLs spot-checked, both 200).

⚠️ Related, outside this PR: old Ghost URLs currently 404 on the new site, and every existing Check report and tipline response links to one. Keeping the default on the Ghost shape keeps new reports consistent with the old ones, but the 404s need redirects on pesacheck-ui.

Database

guid NOT LIKE 'http%' meant "Ghost", which breaks once Superdesk guids (UUIDs) arrive. Added source and language columns via an idempotent migration that backfills medium/ghost for existing rows. The fetch checkpoint is scoped to the active provider, falling back to all sources as described above. Legacy Medium summary handling keys off source instead of sniffing the guid.

Settings

Variable Default Notes
PESACHECK_PROVIDER ghost ghost | superdesk
PESACHECK_POSTS_LIMIT PESACHECK_GHOST_POSTS_LIMIT, else 15 page size; deployments keep their value
PESACHECK_MAX_ARTICLES 100 ceiling on one catch-up run
PESACHECK_SITE_URL https://pesacheck.org
PESACHECK_ARTICLE_URL_TEMPLATE {site}/{slug}/
PESACHECK_SUPERDESK_GRAPHQL_URL — required for superdesk
PESACHECK_SUPERDESK_TENANT_CODE — required for superdesk
PESACHECK_SUPERDESK_PRESHARED_AUTH unset sent as x-preshared-auth

Testing

  • 70 tests (pants test pesacheck_meedan_bridge/py::, green in CI): provider registry and per-provider settings validation; Ghost behaviour; Superdesk mapping, tenant/Debunk filter, paging, GraphQL errors, preshared-auth header, unparseable metadata, guid fallback, URL template incl. an article with no route; the Ghost→Superdesk cutover; the catch-up cap and the oldest-first ordering contract; legacy naive timestamps in the global checkpoint; the schema migration; plus everything inherited from feat(pesacheck_meedan_bridge): Fetch articles from Ghost Content API #1208/fix(pesacheck_meedan_bridge): Stop retrying fact-checks Check already has #1210.
  • Live against graphql-staging.pesacheck.org (tenant 123abc, 14,773 fact-checks): mapping verified on real rows; generated URLs return 200.
  • Migration on a copy of the production database: 25 legacy rows correctly become source=medium.
  • End-to-end into the Check sandbox: a first Superdesk run created articles with the right language and tags (media 3900334–3900336, 3900350–3900351); re-runs posted 0; switching back to ghost resumed from the Ghost checkpoint.
  • Live cutover test: seeded with a Ghost checkpoint of 2026-09-20, the first Superdesk run fetched only the 5 articles published after it, oldest first, rather than the archive.
  • flake8, ruff format, isort, bandit pass.

Deploying the switch

  1. Set PESACHECK_SUPERDESK_GRAPHQL_URL and PESACHECK_SUPERDESK_TENANT_CODE (production, not staging).
  2. Set PESACHECK_SITE_URL, and PESACHECK_ARTICLE_URL_TEMPLATE if the new site's shape should be canonical.
  3. Set PESACHECK_PROVIDER=superdesk.

Production already has Ghost articles stored, so the first Superdesk run resumes from the last Ghost one rather than re-importing. Keep an eye on the first run's Sentry summary.

Not included

  • Thumbnails for Superdesk are left empty: Check never receives one, and resolving renditions needs the media base URL.
  • No dual-source running — one provider per run, by design.
  • PESACHECK_SUPERDESK_PRESHARED_AUTH is the sturdier answer to the Cloudflare problem we hit in production, where the WAF rule is pinned to the Dokku host's non-elastic IP. Worth adopting for Ghost too, separately.

🤖 Generated with Claude Code

Base automatically changed from fix/pesacheck-duplicate-factchecks to main September 30, 2026 07:01
@koechkevin
koechkevin requested a review from a team September 30, 2026 16:06
@kilemensi

This comment was marked as resolved.

@chatgpt-codex-connector

This comment was marked as resolved.

@chatgpt-codex-connector

This comment was marked as resolved.

@kilemensi

This comment was marked as outdated.

PesaCheck is moving from Ghost to Superdesk, so the bridge can no longer
be Ghost-shaped throughout. Fetching, parsing, summaries and language
detection move behind a provider interface, and PESACHECK_PROVIDER picks
which CMS a run reads.

- provider_base.py: the normalized Article the rest of the bridge sees
- provider_ghost.py: today's Content API behaviour, unchanged
- provider_superdesk.py: Publisher's GraphQL API, filtered to the tenant
  and to articles carrying a Debunk verdict, paginated by limit/offset
- providers.py: the registry; each provider validates its own settings,
  so Superdesk no longer needs a Ghost key
- database: source and language columns, added by an idempotent
  migration that backfills medium/ghost for existing rows; the fetch
  checkpoint is now per provider, so switching back and forth neither
  re-posts nor loses place
- PESACHECK_POSTS_LIMIT replaces PESACHECK_GHOST_POSTS_LIMIT, falling
  back to it so deployments keep their value

Providers are flat modules rather than a package: py/BUILD globs *.py in
that directory only, and a subpackage needs its own BUILD plus correct
pex dependency inference, which can't be verified without Pants locally.
@koechkevin
koechkevin force-pushed the feat/pesacheck-provider-switch branch from bfcb45f to 7a83eac Compare October 1, 2026 07:11

@kilemensi kilemensi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice refactor — the provider split is clean and the per-provider settings validation is a real improvement.

One problem with the cutover itself: the first Superdesk run will re-post fact-checks Check already has, and can skip articles. Details inline on main.py. The other two comments are a question about URLs and a docs fix.

Separately from this PR, and worth raising on pesacheck-ui: old Ghost URLs (pesacheck.org/<slug>/) currently 404 on the new site. They fall through to the catch-all app/[...slug] route, which only renders Publisher pages, and nothing uses swp_redirect_route yet. Every existing Check report and tipline response links to those URLs.

Comment thread pesacheck_meedan_bridge/py/main.py Outdated
Comment thread pesacheck_meedan_bridge/py/provider_superdesk.py Outdated
Comment thread pesacheck_meedan_bridge/README.md Outdated
@koechkevin

Copy link
Copy Markdown
Contributor Author

Nice refactor — the provider split is clean and the per-provider settings validation is a real improvement.

One problem with the cutover itself: the first Superdesk run will re-post fact-checks Check already has, and can skip articles. Details inline on main.py. The other two comments are a question about URLs and a docs fix.

Separately from this PR, and worth raising on pesacheck-ui: old Ghost URLs (pesacheck.org/<slug>/) currently 404 on the new site. They fall through to the catch-all app/[...slug] route, which only renders Publisher pages, and nothing uses swp_redirect_route yet. Every existing Check report and tipline response links to those URLs.

People need to sleep😂

…tover

Review feedback on #1211.

A provider's first run had no checkpoint, so it fetched the newest page
and posted it. At the Ghost -> Superdesk cutover those same fact-checks
are already in Check under Ghost's guids and URLs: the new guids don't
match feed_exists, and Check doesn't reject them either, because its
signature is an MD5 over set_fact_check (which carries the URL) plus the
team. We'd have published duplicate reports. Anything published beyond
one page since the last Ghost run would also have been skipped.

A provider with no articles of its own now starts from the newest
article stored under any source, i.e. where the previous provider
stopped. Confirmed migrated articles keep their original published_at:
staging holds fact-checks back to 2018-11-15.

That fallback can reach far back, so catch-up is now bounded:
- fetch oldest first when catching up (newest page when there is no
  checkpoint at all), so a partial run still advances the checkpoint
- stop at PESACHECK_MAX_ARTICLES (default 100) per run and continue on
  the next one

Without the bound, a stale checkpoint walked years of archive in one
run: a local test with only 2024-era rows paginated ~2 years of Ghost
posts and would have posted them all to Check.

Also fixes a crash the new tests caught: get_checkpoint() used UTC,
which main.py no longer imported after the provider refactor, so any
legacy Medium row (naive timestamps) would raise NameError.
Review feedback on #1211. Publisher has no canonical-URL field
(swp_redirect_route is empty), so the bridge builds the URL. The shape
decides more than the link: Check's duplicate signature covers the
fact-check URL, so a new shape makes already-imported articles look new,
and existing Check reports all link to the Ghost-era form.

PESACHECK_ARTICLE_URL_TEMPLATE defaults to {site}/{slug}/, which matches
those existing reports, and can be switched to
{site}/fact-checks/{desk}/{slug} once the new site serves that
canonically. Publisher slugs match the Ghost slugs today (spot-checked
two live URLs, both 200).
@koechkevin
koechkevin requested a review from kilemensi October 5, 2026 06:59

@kilemensi kilemensi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One new bug — both providers

provider_ghost.py:70 and provider_superdesk.py:137: the max_articles early return bypasses the oldest-first reversal
at the end of the function.

if max_articles is not None and len(posts) >= max_articles:
return posts[:max_articles] # no reversal
...
return posts if since else list(reversed(posts))

On the no-checkpoint path the fetch is order=desc, so if that early return fires it hands back newest-first,
contradicting the function's own docstring ("Returns oldest first either way"). main.py then processes newest-first,
the checkpoint jumps straight to the newest article, and a mid-batch store failure (break) permanently strands every
older article in the batch — the exact invariant the comment at main.py:157 asserts.

It needs PESACHECK_POSTS_LIMIT >= PESACHECK_MAX_ARTICLES to trigger, so the shipped defaults (15 vs 100) are safe and
it's latent today. But it's config-reachable by anyone raising the page size to cut round-trips, it's silent, and the
failure mode is permanent article loss. Fix is to break with the slice instead of returning, letting the single exit
handle ordering.

… order

Review feedback on #1211. The max_articles early return skipped the
reversal at the single exit, so on the no-checkpoint path (which fetches
newest first) a capped run handed back newest-first. main() would then
advance the checkpoint to the newest article, and a mid-batch store
failure would strand every older article in the batch for good.

Latent at the shipped defaults (page size 15, cap 100) since the cap
can't fire on the first page, but reachable by raising the page size to
the cap or beyond. Break with the slice instead, so ordering stays with
the single exit.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants