Provide a native Python API for Quickdash's configurable eval interpretation and scoring, with a calculated aggregation tree, visible structured warnings, and shared behavioral tests against the browser engine.
Scope: extract the browser processing API, implement native Python CSV/YAML processing and all current scoring/coverage rules, add a tree/JSON CLI, and document the contracts. Configuration sizes, contents, and ordering must not be assumed. Sampo's summarize-evals is a consumer/reference for presentation, not the scoring specification.
Status: ready for review in draft PR #4 (9d6ff92). Implementation role: Codex. Next action: review the API and before/after dashboard behavior together. Builds on weighted components in #3. Public CI passed: https://github.com/OpenEuroLLM/quickdash/actions/runs/36887115024
Acceptance: both engines satisfy the same independently specified cases and agree on trees, weights, coverage, and diagnostic conditions; Python runs without Node; dashboard behavior is compared with the pre-change version, with intentional bug fixes identified; CLI warnings are visible on stderr and returned separately from score trees. No private result data is published.
Provide a native Python API for Quickdash's configurable eval interpretation and scoring, with a calculated aggregation tree, visible structured warnings, and shared behavioral tests against the browser engine.
Scope: extract the browser processing API, implement native Python CSV/YAML processing and all current scoring/coverage rules, add a tree/JSON CLI, and document the contracts. Configuration sizes, contents, and ordering must not be assumed. Sampo's summarize-evals is a consumer/reference for presentation, not the scoring specification.
Status: ready for review in draft PR #4 (9d6ff92). Implementation role: Codex. Next action: review the API and before/after dashboard behavior together. Builds on weighted components in #3. Public CI passed: https://github.com/OpenEuroLLM/quickdash/actions/runs/36887115024
Acceptance: both engines satisfy the same independently specified cases and agree on trees, weights, coverage, and diagnostic conditions; Python runs without Node; dashboard behavior is compared with the pre-change version, with intentional bug fixes identified; CLI warnings are visible on stderr and returned separately from score trees. No private result data is published.