Site-wide linked project repositories rechecked: 47/47 — 27 September 2026.
This pass rebuilt the repository list directly from the portfolio pages, covering the 39 AI repositories plus eight Research / Learning & Development empirical repositories. It rechecked documented formulas, machine-readable result values, headline arithmetic, and portfolio text against the current committed evidence.
- mini_transformers_sequences — corrected stale documentation from an earlier Transformer run. The current 6-month three-seed Transformer mean MAE is 19.691348204741605 (displayed as 19.691) versus HGB 20.537293415179263, giving a mean delta of -0.8459452104376588 (displayed as -0.846). The calculation guide, paper, review figure, and portfolio page were synchronized.
- knowledge_tracing_benchmark — synchronized the exact seed-42 GRU Brier value in the calculation guide and paper to the current machine-readable result: 0.1799516879180387. Rounded public text remains 0.1800.
- The eight Research / L&D empirical repositories were independently spot-recomputed from their released headline quantities (frontier fractions, Cramér's V, network gaps, triage enrichment, Jaccard overlap, forecast-error summaries, treatment-effect arithmetic, and training-investment percentage changes). No additional headline arithmetic error was found.
This site-wide pass verifies internal numerical consistency between committed result files, formulas, calculation guides, papers, and portfolio text. It does not claim that every one of the 47 projects was retrained or rebuilt from raw external data during this pass. Existing project-level reproducibility workflows and tests remain the stronger check for full source-to-result reproduction.
Repository completion: 39/39.
LAST REPOSITORY COMPLETED: Feedback Quality Evaluator — 39/39.
The last repository commit for the main pass is 21ba2cfefa7309c1919df484ac20fcc8e6fe3aff. Subsequent follow-up commits integrate reproduced LSTM and TalkMoves results. Review date: 26 September 2026.
- Both portfolio pages revised: 9 AI Engineering entries and 30 AI in Education entries.
- Eight featured empirical entries use two paragraphs; the other 31 use one paragraph.
- Every repository has two new scientific SVG figures, a calculation guide, source-linked evidence and reproducible figure generation.
- All 39 README front sections revised; 13 existing working papers expanded or corrected. Existing references and detailed documentation retained where applicable.
- Scientific figures distinguish external-data results, synthetic demonstrations and illustrative arithmetic. They do not fabricate measurements.
| Repository | Correction and evidence |
|---|---|
| Multimodal Self-Regulation Lab | Replaced in-sample classifier AUC with a shared stratified 70/30 holdout; imputation/scaling fit only on training data. The rank-fusion comparator is explicitly transductive. |
| LSTM Time Series | Replaced future-borrowing interpolation with past-only forward fill; full 20-epoch run and CI succeeded. Updated results, figures and working paper. Run. |
| Feedback Quality Evaluator | Constant identical ratings now produce undefined kappa rather than an unjustified value of one; exact agreement remains one. Regression tests pass. |
| Classroom Discourse Intelligence | Completed the empirical study, detected duplicate archive copies, added content deduplication/grouping, and reran. Final analysis: 565 groups and 175,129 teacher utterances. Corrected run. |
- 27 bundled demonstrations executed successfully.
- 441 checks passed in 17 existing unittest suites during the initial review. Additional regression checks cover the three calculation fixes; the feedback suite subsequently passed 43 tests.
- 28 existing pytest-style test functions were executed directly, with supported temporary-directory arguments; no failures. This is not reported as a full pytest-suite run.
- 121 independent stored-result arithmetic checks passed within declared floating-point tolerances. These cover selected summary statistics, confusion-derived metrics and paired differences, not every possible mathematical claim.
- All 78 SVGs parse and regenerate from their stated source files. Text-bound checks found no overflow in the main pass; representative figures were rendered and inspected.
- Page structure checks confirm paragraph and figure counts. Local browser screenshots could not be produced because the browser executable was unavailable; no local desktop/mobile rendering pass is claimed.
The review does not claim that all 39 projects were independently reproduced from raw data or scientifically validated. Synthetic policy rules remain prototypes. Most pre-existing empirical results were inspected against their committed evidence; the corrected LSTM and TalkMoves experiments were fully rerun during this review.