Lightweight proof-of-concept for oversight-centered metrology in coding agents: workflow-aware evaluation, interrupt channels, and claim-margin reporting beyond raw success scores.
-
Updated
Mar 12, 2026 - Python
Lightweight proof-of-concept for oversight-centered metrology in coding agents: workflow-aware evaluation, interrupt channels, and claim-margin reporting beyond raw success scores.
Reproducible benchmark of security-oracle construct validity on 140 real CVE fixes.
Two benchmark-validity studies of transcriptome-based metabolic reaction activity scoring on Recon3D (joint patient-and-reaction hold-out with input substitution; language-model interventions and external cohorts), the metabench checks, and the MetaGNN scorer they audit
MQ-EVA — Independent AI Evaluation Assurance: a framework for determining whether AI performance claims can be trusted
To associate your repository with the benchmark-validity topic, visit your repo's landing page and select "manage topics."