Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 22 additions & 1 deletion pesacheck_meedan_bridge/.env.example
Original file line number Diff line number Diff line change
@@ -1,6 +1,27 @@
# Which CMS to read fact-checks from: ghost or superdesk.
PESACHECK_PROVIDER=ghost
# Page size, and how many articles a run with no stored article fetches.
PESACHECK_POSTS_LIMIT=15
# Ceiling on one catch-up run, so a checkpoint far in the past (a long outage,
# or the first run after switching provider) can't import years at once.
PESACHECK_MAX_ARTICLES=100
# Public site, used to build the article URL posted to Check.
PESACHECK_SITE_URL=https://pesacheck.org
# How a Superdesk article's URL is built. The default is the Ghost-era shape,
# which every fact-check already in Check links to. Switch to
# "{site}/fact-checks/{desk}/{slug}" once the new site serves that canonically.
PESACHECK_ARTICLE_URL_TEMPLATE={site}/{slug}/

# Ghost (provider: ghost)
PESACHECK_URL=https://pesacheck.org
PESACHECK_GHOST_CONTENT_API_KEY=
PESACHECK_GHOST_POSTS_LIMIT=15

# Superdesk/Publisher (provider: superdesk)
PESACHECK_SUPERDESK_GRAPHQL_URL=https://graphql-staging.pesacheck.org/v1/graphql
PESACHECK_SUPERDESK_TENANT_CODE=
# Optional: shared secret a Cloudflare WAF rule matches to skip bot protection.
PESACHECK_SUPERDESK_PRESHARED_AUTH=

PESACHECK_CHECK_URL=https://check-api.checkmedia.org/api/graphql
PESACHECK_CHECK_TOKEN=
PESACHECK_CHECK_WORKSPACE_SLUG=pesacheck-tipline-sandbox
Expand Down
48 changes: 46 additions & 2 deletions pesacheck_meedan_bridge/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,10 +41,48 @@ To run `pex` binary, execute:
docker compose exec api-pesacheck_meedan_bridge ./pex
```

## Where articles come from

`PESACHECK_PROVIDER` selects the CMS the bridge reads:

| Provider | Reads | Needs |
|---|---|---|
| `ghost` | Ghost Content API at `{PESACHECK_URL}/ghost/api/content/posts/` | `PESACHECK_GHOST_CONTENT_API_KEY` |
| `superdesk` | Publisher's GraphQL API (`swp_article`), filtered to the tenant and to articles carrying a `Debunk` verdict | `PESACHECK_SUPERDESK_GRAPHQL_URL`, `PESACHECK_SUPERDESK_TENANT_CODE` |

Only the active provider's settings are required, so switching is a config
change plus a restart. Each article records the provider it came from in the
`source` column, and the fetch position is tracked per provider.

A provider that has no articles of its own yet — the first run after a switch —
starts from the newest article stored under *any* source, i.e. wherever the
previous provider stopped. Both CMSes hold the same fact-checks under different
ids and URLs, so starting from scratch would re-post them: Check would not
reject them as duplicates, because its signature covers the fact-check URL,
which differs between the two sites. It would also skip anything published
beyond one page since the last run.

Superdesk articles are posted to Check with a URL built from
`PESACHECK_SITE_URL` and `PESACHECK_ARTICLE_URL_TEMPLATE` (`{site}`, `{desk}`,
`{slug}`); point the site at a preview deployment to test against one.
Publisher has no canonical-URL field — `swp_redirect_route` is empty — so the
URL is assembled here, and the default keeps the Ghost-era `{site}/{slug}/`
shape that every fact-check already in Check links to. The shape matters
beyond the link: Check's duplicate signature covers the fact-check URL, so
changing it makes already-imported articles look new.

Their language comes from Superdesk directly, and the Check tags are the
language, country, content type and harm type.

## How articles are picked up

Each run fetches everything published since the newest PesaCheck (Ghost) article
already stored in SQLite, paginating in pages of `PESACHECK_GHOST_POSTS_LIMIT`.
Each run fetches everything published since the newest article already stored
for the active provider, oldest first, in pages of `PESACHECK_POSTS_LIMIT` and
at most `PESACHECK_MAX_ARTICLES` per run. The cap bounds a checkpoint far in
the past — a long outage, or the first run after a provider switch — and
costs nothing: the checkpoint advances as articles are stored, so the next run
picks up where this one stopped.

When nothing is stored yet, only the newest page is fetched, so the first run
against a fresh database does not backfill PesaCheck's whole archive.

Expand All @@ -57,6 +95,12 @@ Articles are tracked by `status`:
| `Completed` | Accepted by Check, with the Check ids stored. |
| `Duplicate` | Check already has this fact-check (it rejects a repeat of the same content). Terminal: never retried. The Check ids are not recorded, so find the article in Check by title if you need them. |

### Adding another provider

Add a module exposing `name`, `fetch(since, limit)` returning raw records, and
`parse(record)` returning a `provider_base.Article`; register it in
`providers.py`. `main.py` knows nothing about any particular CMS.

### Known limitation: backdated articles

The cursor is the newest `published_at` we have stored, so an article published
Expand Down
2 changes: 1 addition & 1 deletion pesacheck_meedan_bridge/py/VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
0.1.22
0.1.26
73 changes: 63 additions & 10 deletions pesacheck_meedan_bridge/py/database.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,12 +19,22 @@ class PesacheckFeed:
check_project_media_id: str = ""
check_full_url: str = ""
claim_description_id: str = ""
# Which CMS the article came from: "medium", "ghost" or "superdesk".
source: str = ""
language: str = ""


class PesacheckDatabase:
# Appended to pesacheck_feeds after the original columns, in this order.
ADDED_COLUMNS = (
("source", "TEXT NOT NULL DEFAULT ''"),
("language", "TEXT NOT NULL DEFAULT ''"),
)

def __init__(self):
self.db_file = settings.PESACHECK_DATABASE_NAME
self.create_table()
self.migrate()

def create_connection(self):
return sqlite3.connect(self.db_file)
Expand All @@ -46,18 +56,49 @@ def create_table(self):
categories TEXT DEFAULT '[]',
check_project_media_id TEXT,
check_full_url TEXT,
claim_description_id TEXT)"""
claim_description_id TEXT,
source TEXT NOT NULL DEFAULT '',
language TEXT NOT NULL DEFAULT '')"""
)
conn.commit()
finally:
conn.close()

def migrate(self):
"""Add columns a database created by an older version is missing."""
conn = self.create_connection()
try:
cur = conn.cursor()
existing = {
row[1] for row in cur.execute("PRAGMA table_info(pesacheck_feeds)")
}
added = []
for column, definition in self.ADDED_COLUMNS:
if column not in existing:
cur.execute(
f"ALTER TABLE pesacheck_feeds ADD COLUMN {column} {definition}"
)
added.append(column)
if "source" in added:
# Rows predating the column: Medium stored the post URL as the
# guid, Ghost stored the post id.
cur.execute(
"""UPDATE pesacheck_feeds
SET source = CASE WHEN guid LIKE 'http%' THEN 'medium'
ELSE 'ghost' END
WHERE source = ''"""
)
conn.commit()
finally:
conn.close()

def insert_pesacheck_feed(self, feed):
conn = self.create_connection()
sql = """INSERT INTO pesacheck_feeds (title, pubDate, author,
guid, link, thumbnail, description, status, categories,
check_project_media_id, check_full_url, claim_description_id)
VALUES(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)"""
check_project_media_id, check_full_url, claim_description_id,
source, language)
VALUES(?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)"""
try:
cur = conn.cursor()
cur.execute(
Expand All @@ -75,6 +116,8 @@ def insert_pesacheck_feed(self, feed):
feed.check_project_media_id,
feed.check_full_url,
feed.claim_description_id,
feed.source,
feed.language,
),
)
conn.commit()
Expand All @@ -87,7 +130,8 @@ def update_pesacheck_feed(self, guid, new_feed, expected_status=None):
SET title = ?, pubDate = ?, author = ?, link = ?, thumbnail = ?,
description = ?, status = ?, categories = ?,
check_project_media_id = ?, check_full_url = ?,
claim_description_id = ? WHERE guid = ?"""
claim_description_id = ?, source = ?, language = ?
WHERE guid = ?"""
params_tail = []
if expected_status is not None:
sql += " AND status = ?"
Expand All @@ -108,6 +152,8 @@ def update_pesacheck_feed(self, guid, new_feed, expected_status=None):
new_feed.check_project_media_id,
new_feed.check_full_url,
new_feed.claim_description_id,
new_feed.source,
new_feed.language,
guid,
*params_tail,
),
Expand Down Expand Up @@ -174,15 +220,22 @@ def feed_exists(self, guid):
finally:
conn.close()

def get_ghost_pub_dates(self):
# Legacy Medium rows use the post URL as guid; Ghost rows use the post id.
def get_pub_dates(self, source=None):
"""Publication dates already stored, for one provider or for all.

All sources is what the first run of a newly configured provider needs:
the articles exist under the old provider's source, so its position is
the only thing that says where the new one should start.
"""
conn = self.create_connection()
sql = "SELECT pubDate FROM pesacheck_feeds WHERE pubDate != ''"
params = ()
if source is not None:
sql += " AND source = ?"
params = (source,)
try:
cur = conn.cursor()
cur.execute(
"SELECT pubDate FROM pesacheck_feeds "
"WHERE guid NOT LIKE 'http%' AND pubDate != ''"
)
cur.execute(sql, params)
return [row[0] for row in cur.fetchall()]
finally:
conn.close()
Expand Down
Loading
Loading