Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Frede — Fake (Machine-Generated) Review Detector

Detects whether a restaurant or local-business review was written by a person or generated by a language model, and explains each decision.

Each review is scored on two signal families computed over raw text: GPT-2 surprisal statistics (including Fast-DetectGPT's conditional probability curvature) and stylometric surface features. Explanations are exact rather than sampled — token highlighting is each token's probability under the reference model, which is the same quantity the detector reads, and feature attributions are SHAP values from the classifier's own trees. A request lands in 60–300 ms.

The web UI also carries a Human-in-the-Loop study measuring whether those explanations help people tell machine-written reviews from human ones.

Where this came from. The project began as a port of ref/guided_study.ipynb, which trained a sentiment classifier on a star-rating proxy and called it fake-review detection. That pipeline is archived — see Archived. The replacement detects authorship against ground truth rather than sentiment against a label of convenience, and the reasoning is in PLAN.md and the stage sections below.

Stack

  • Backend: FastAPI (single worker), serving a JSON API + the built frontend SPA.
  • Frontend: React + Ant Design, built with Vite (frontend/ → frontend/dist, served by FastAPI). Node.js + npm required to build.
  • ML: PyTorch + HuggingFace transformers (GPT-2 surprisal), scikit-learn (gradient-boosted classifier), shap for live feature attribution.

Quick start

# 1. Install Python dependencies (into the existing .venv)
uv pip install --python .venv/bin/python -e ".[all]"
cd frontend && npm install && cd ..

# 2. Build the generation packet (downloads the Yelp corpus on first run)
make gen-packet

# 3. Generate the machine-written half — see prompts/GENERATION_BRIEF.md.
#    This step needs a model and is not automated here. Hand the brief and the
#    packet to a generator, then fold the results back in:
make check-pairs RESULTS=data/gen_results/<tag> GENERATOR=<tag>
make ingest-gen  RESULTS=data/gen_results/<tag> GENERATOR=<tag>

# 4. Train the detector, then build the study artifacts
make stage2
make stage4

# 5. Build the frontend, then run the app
make run            # -> http://127.0.0.1:8000

Step 3 is the one that cannot be shortcut, and the one worth reading the brief for: three separate generation runs have been produced that passed every per-record check and were still unusable.

What you get

Page URL What it does
Analyzer /#/ Paste a review → human/machine verdict at a calibrated threshold, token-predictability highlighting, reason codes with evidence, and a per-feature table (value, position in the human distribution, SHAP contribution).
HITL /#/hitl Two-pass study: judge the same reviews first raw, then with explanations. Ground truth is actual authorship. Results show human accuracy/precision/recall/F1 and Cohen's κ against the model, per pass.
Examples /#/examples One worked case per confusion-matrix quadrant, from the held-out split.

Frontend dev mode (hot reload against the same API): cd frontend && npm run dev, then open http://localhost:5173. The dev server proxies /api only if you configure server.proxy in frontend/vite.config.js.

API

  • GET /api/health — detector readiness, feature count, training generators.
  • POST /api/analyze {text} → prediction, machine_prob, flagged, threshold, features (value / human percentile / SHAP), tokens/logprobs, highlighted_html, reason_codes, summary, scope_note, elapsed_ms.
  • GET /api/hitl/session / POST /api/hitl/judgment / GET /api/hitl/results / POST /api/hitl/reset — the two-pass HITL flow. judgment is Human or Machine.
  • GET /api/examples — precomputed worked examples.

Reading the verdict

prediction is taken at the threshold calibrated to hold false positives on real reviews at 1%, not at 0.5 — a 0.5 boundary carries no false-positive guarantee and would imply roughly an order of magnitude more false accusations. A review can therefore sit above 0.5 and still be reported as human; the summary says so explicitly when it happens ("leans slightly machine but not enough to flag") rather than silently rounding the verdict.

Archived: the original sentiment-proxy study

Kept for reference. make train still builds it, but nothing serves it and the app no longer loads it at startup.

The notebook labelled Yelp 1–2★ reviews as fake and 4–5★ as real, so its model detects harsh sentiment, not fabricated content — a praise-heavy promotional review is confidently predicted REAL. make train fine-tunes distilbert on that label and writes models/fake_review_distilbert/; the old LIME / Integrated Gradients explanation stack that went with it has been removed.

Smoke-test the archived pipeline on a tiny subset: make smoke.

How the port maps to the notebook

  • Phase 1 (data/preprocessing): training/data.py + app/preprocess.py. The canonical preprocess() = clean_text() → remove_stopwords() is applied identically at train and serve time (the notebook does the same in cells 10+12→16 and cell 36).
  • Phase 2 (models): training/train.py — plain-PyTorch fine-tune loop with AdamW + linear warmup, best-val-F1 checkpoint (the notebook used RoBERTa; we use distilbert-base for CPU-friendly training).
  • Phase 3 (XAI): the notebook's LIME + IG stack, which lived in app/services/xai_service.py and app/reason_codes.py. Both files have since been rewritten for the current detector — they now hold the SHAP and token-predictability explanation layer, not the notebook's. The old code is gone; this entry records only what the port originally mapped to.
  • Phase 4 (dashboard + HITL): app/routers/ + frontend/. A bug in the notebook (cell 33 never appends its generated explanations, leaving its CSV empty) is not reproduced.

What "fake" means here

Faithful to the notebook, the training label is a star-rating proxy: Yelp 1–2★ reviews are labelled fake (1) and 4–5★ real (0), 3★ dropped. So the model is really detecting low-star / harsh reviews, and a praise-heavy promotional review ("best ever! five stars!") is usually predicted REAL — even though the rule-based Reason Codes (e.g. RC-01 promotional language) will still flag its promotional wording. That tension is inherent to the notebook's methodology and is reproduced as-is.

Dataset trap: Yelp/yelp_review_full stores \n literally

The corpus encodes paragraph breaks as the two literal characters \ n, not as newlines: 52% of reviews contain them and 0% contain a real newline.

Left alone this silently destroys machine-text detection. Freshly generated machine text never carries the artefact, so "contains \n" separates the two classes at ~0.96 AUC while measuring nothing about authorship — the detector would learn a dataset quirk and fail on any real input. It also leaked into the generation prompts, so rewrites were being asked to paraphrase escape litter.

app.preprocess.normalize_escapes() converts them to real characters and is applied on both sides: in detect.stage0.load_yelp_raw (so packets are built from clean text) and in scripts.ingest_gen_results.py (so already-generated text is normalised identically). Anything new that reads Yelp text directly must apply it too.

A second-order effect remains and is worth watching: after normalisation, 52% of human reviews carry real paragraph breaks versus 7% of machine output. Part of that is genuine — LLM reviews really are usually one block — but part was induced by feeding the models escape litter. Packets regenerated after the fix no longer induce it.

Machine-generated detection (Stage 0 + Stage 1)

The star-proxy label above detects sentiment, not authorship. detect/ and scripts/ build the replacement: real detection of machine-written reviews, using no external labelled dataset.

Stage 0 — zero-shot baseline (make stage0, → data/stage0_report.json):

  • detect/features.py computes two signal families on raw text — reference-LM surprisal statistics including Fast-DetectGPT's conditional probability curvature, plus stylometric/burstiness features.
  • The runner measures per-feature AUC against local generators from two families (matched = optimistic bound, cross-family = pessimistic bound), so the signal is bracketed rather than assumed.
  • A preprocessing ablation re-runs the same comparison after app.preprocess.preprocess. It exists to quantify what the current pipeline destroys: lowercasing, punctuation stripping and stopword removal delete exactly the features that separate human from machine writing.
  • Finally it correlates the shipped model's P(fake) against star rating and against curvature on human reviews only.

Stage 1 — paired generation packet (see prompts/generate_machine_reviews.md):

make gen-packet                                   # -> data/gen_packet/
make check-pairs RESULTS=data/gen_results/model-a GENERATOR=model-a
make ingest-gen  RESULTS=data/gen_results/model-a GENERATOR=model-a

The generator's operator brief is prompts/GENERATION_BRIEF.md — hand that to the generating agent; it is self-contained. make check-pairs is the gate that decides whether a run is usable: it catches template collapse, rewrites that ignored their source, lost opinions, and truncation — failure modes that all score 100% coverage and are invisible from a completion count.

Writes tasks.jsonl (all records), shards/ (the work split into 50-record files, so an agent run is resumable per shard), and a held-out human_eval.jsonl for false-positive measurement. Add --vendor-files for pre-wrapped requests.anthropic.jsonl / requests.openai.jsonl batch files.

Three conditions, all star-balanced across 1–5★:

condition source text role
machine_rewrite given core paired training signal (content-matched)
machine_generate withheld generalisation beyond rewriting
human_edit given negative control — separates "machine-written" from "has no typos"

Run the packet through at least two different generators: Stage 2 trains on one and evaluates on the other, which is the only honest measure of whether the detector learned machine authorship or one model's quirks.

Stage 2 — detector (make stage2, → data/stage2_report.json):

Trains on the detect.features zero-shot signals over raw text and reports per-generator. Three rules make the eventual second generator a new row rather than a rewrite of the analysis:

  • detect/dataset.py draws the human class from the packet's own source_text, never from a generator's output, so adding a generator never perturbs the negative class.
  • Splits are keyed to custom_id, which the frozen packet already assigns — so two generators running the same packet inherit identical train/dev membership and the cross-generator number is just a difference of rows.
  • human_edit is a probe, not training data. It is human text a model has copy-edited; keeping it out of training preserves a clean measurement of how often the detector flags a real reviewer just because the text is now clean.

Headline metric is TPR at a fixed FPR, not F1: a false positive means accusing a real customer, so the operating point is set by the false-positive rate. The frozen dev split holds only ~150 machine records, hence the cross-validated error bars.

make stage2                                        # as shipped — INVALID, see below
make stage2 GENERATORS="<run-a> <run-b>"           # once real generator runs exist

Provenance: one generation run was synthetic, and how it was caught

The first "MiMo V2.5" run (data/gen_results/mimo-v2.5-full/) was not model output. It came from scripts/gen_reviews.py, a 971-line rule-based generator that reproduces all three conditions byte-for-byte. A detector trained on it separates template-generated text from human text — AUROC 0.997 is what a deterministic paraphrase rule against real prose yields — and every downstream number from it is void. The file is kept as the provenance record for that dataset.

What eventually caught it was a cross-record check, because every per-record check passed: 33.9% of its 5-grams repeated ten or more times, against 0.11% for genuine MiMo output and 0.01% for DeepSeek. Two earlier signals were visible and not pursued — an AUROC of 0.997 on an open problem is implausible (published cross-domain machine-text detection sits around 0.85–0.95), and a 90% positional verbatim rate is a substitution rule, not a rewrite.

Two gates now block this class of failure:

  • check_pairs reports per-condition health and a verdict.
  • ingest_gen_results refuses to write a run that fails, and scopes the refusal: a misaligned packet blocks everything, while a single bad condition costs exactly that condition rather than the good data beside it.

Template detection gates on concentration (share of 5-grams repeated 10+ times), not on the plain reuse rate. Reuse conflates text assembled from a phrase pool with text that is individually written but formulaic; only concentration separates them. Reuse is still reported, as a style statistic relevant to single-generator risk, but it does not gate.

Results on the genuine data

make stage2 GENERATORS=mimo-v2.5-rerun — MiMo V2.5 output, verified model-written:

CV AUROC (out-of-fold, n=2706) 0.9595 ± 0.0084
Held-out test AUROC (n=439) 0.9572
TPR @ 1% FPR 0.655
Length-matched CV (n=2047) 0.9619
FPR, copy-edited human reviews 0.49%
FPR, held-out real reviews 0.75%

Both conditions are now measurable separately, and the split is the expected one — a model writing from scratch is far easier to catch than a model rewriting a real review while preserving it:

condition AUROC TPR@1%FPR n (test)
machine_generate 0.9936 0.841 88
machine_rewrite 0.8978 0.352 54

Against a strong modern generator, surface style beats perplexity. The feature families invert relative to Stage 0's zero-shot result: stylometric features alone reach CV 0.9430, surprisal features alone only 0.8189, and the top contributors are style_avg_word_len (dAUC +0.077), style_has_em_dash, style_excl_per_sent and style_contraction_rate. Stage 0 found GPT-2 surprisal dominant — but it was scored against GPT-2 and Pythia, not against a model of MiMo's class. Perplexity separates weak generators well and strong ones poorly.

This makes the single-generator caveat concrete rather than theoretical: if avg_word_len and em-dash usage are MiMo's habits rather than properties of machine authorship, a large part of the 0.96 is a house-style detector. The cross-generator row is the only thing that distinguishes the two.

Scope of the shipped results — read before quoting any number

Three limits apply to data/stage2_report.json as produced by make stage2:

  1. One generator (MiMo V2.5), with a measured house style. Its 5-gram reuse rate is 17.5% against DeepSeek's 6.7% — genuine output, but formulaic. Given that the detector's top features are stylometric, some of the 0.96 may be detecting that formula rather than authorship. Only a second generator can separate them.

  2. machine_rewrite is a biased sample. Only 371 of 999 rewrites passed the substitution check; the rest were MiMo swapping synonyms rather than rewriting. The survivors are the better end of a poor distribution, so the 0.8978 for that condition is optimistic, and the pooled 0.9572 is weighted toward the easier machine_generate.

  3. A length confound exists but does not drive the result. The generate prompt pins output to 60–160 words while real Yelp reviews run 32–371, so word count alone separates the classes at AUROC ~0.60. Restricting both classes to the overlap window (72–110 words) leaves performance unchanged (length_matched in the report: CV 0.9967 at n=1135) — the shortcut is available but unused. Any future packet should draw each record's target length from the human distribution instead of a fixed band.

  4. The detector is biased against formal human writing, and the measured FPR does not cover it. Bucketing the held-out human reviews by the detector's own score gives a monotone gradient in formality — mean word length rises from 4.13 in the lowest decile to 4.52 in the highest, and the top 5% of real human reviews reach P(machine) = 0.70–0.96. The top human review sits at 0.965 against a 0.9715 threshold.

    None of the 297 held-out human reviews crosses the line, so the 0.49% probe FPR holds on this sample. But the margin is one bad sample wide, and the probe itself was measured on copy-edited Yelp reviews, which skew casual — it does not represent a carefully written formal review. A hand-written formal review tested during Stage 4 scored P = 0.996 and was flagged. Treat the false-positive rate as a lower bound that applies to casual prose.

    The cause is structural: the human class is raw Yelp (typos, contractions, digits, uneven sentences) and the machine class is model output (clean, formal, even). Careful human writing lands nearer the machine side of that gap.

Stage 3 — serving the new detector

The Analyzer page now runs the Stage 2 detector, not the sentiment proxy: app/services/detector_service.py loads the GPT-2 surprisal extractor, the gradient-boosted classifier and the human/machine reference profile; app/services/xai_service.py explains each decision.

Explanations are exact, not sampled. Token highlighting is each token's probability under the reference LM — the same quantity the model reads, so a highlighted token is part of the evidence rather than an approximation of it. Feature attributions are SHAP values computed from the classifier's own trees (~30 ms, exact for tree ensembles). The old LIME-150-samples + Integrated Gradients path cost 3–8 s per request and explained a surrogate; this one lands in 60–300 ms.

Reason codes were rewritten for authorship (app/reason_codes.py). The old RC-01…RC-05 described promotional language and urgency phrasing — artifacts of reading 1-2★ reviews as "fake" — and are preserved in reason_codes_legacy.py for the previous pipeline. Each new code is anchored to a feature whose separation was measured in Stage 2 and fires only when the value falls outside the range real reviews occupy, expressed as a percentile of the human training distribution. A code therefore asserts something checkable — "mean surprisal sits at the ≥99th percentile of real reviews" — rather than a heuristic.

Two conventions worth knowing when reading the UI:

  • Red means predictable, the opposite of the legacy sentiment view. Machine authorship shows up as text the reference model found unsurprising.
  • A reason code can contradict the verdict. It reports a fact about the text, which is independent of which way the balance tipped; those are greyed and labelled "present but outweighed" rather than hidden.

HITL and Examples are still the old task. Their precomputed data (data/hitl_pool.json, data/examples.json) was built for the sentiment proxy and still loads against the legacy model. Regenerating them for machine-authorship is not done yet — the Analyzer is the page that reflects the current system.

Stage 4 — HITL and Examples on the new task

make stage4          # -> data/hitl_pool.json, data/examples.json

Both pages previously answered "fake or real?" with the old sentiment model while the Analyzer answered "human or machine?". They now match.

The HITL question changed to human-written vs machine-written, and ground truth is actual authorship — real Yelp reviews on one side, verified model output on the other. The previous study's "fake" meant 1–2 stars, so its accuracy measured agreement with a sentiment proxy; this one measures what it says. The two-pass design (raw, then with XAI) is unchanged, so the research question — do explanations help people — carries over intact.

Both builders call the same run_analysis the Analyzer calls, so what a participant sees is the live tool rather than a lookalike. The old versions reimplemented the explanation stack offline with LIME and Integrated Gradients, which existed only because the old model needed sampling-based attribution; this detector's explanations are exact and cheap, so there is nothing to reimplement.

Examples picks by quadrant deliberately: correct calls show the most confident instance, errors show the one nearest the decision boundary. An earlier version took the boundary case for all four, which made every example look alike and taught nothing.

A missing quadrant is shown, not hidden. At the calibrated threshold the detector produced no false positives across the held-out split, so the page says so rather than quietly displaying three cases.

Notes / tuning

  • Latency is 60–300 ms per request, dominated by the GPT-2 forward pass. SURPRISAL_MAX_LENGTH in app/config.py caps the tokens scored and is the lever that matters; the SHAP step is ~30 ms and not worth tuning.
  • SURPRISAL_LM is fixed at training time. Every surprisal feature shifts if you change it, so pointing the app at a different reference model invalidates the saved detector. Retrain if you change it.
  • The operating threshold ships inside the model bundle (threshold_at_fpr_1pct), not in config — it is calibrated on training folds at fit time and would be meaningless hard-coded.
  • shap is a hard dependency of the serving path, not optional: the detector service builds a TreeExplainer at startup. The optional extra in pyproject.toml predates that and make install should be used.
  • Single worker is intentional — the detector is a process-wide singleton with a threading.Lock around loading.
  • lime and captum are no longer used and remain in pyproject.toml only because the archived pipeline listed them. They can be dropped.

About

Fake (Machine-Generated) Review Detector

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages