Add native Python analysis, browser parity tests, and weighted eval components - #4
Merged
Merged
Conversation
7 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Quickdash can now calculate and explain evaluation scores from Python using the same CSVs and YAML configurations as the browser. The API returns calculated aggregation trees, effective weights, contributions, coverage and structured diagnostics; a renderer only formats the result. It also includes weighted eval components, PolyMath aggregation, and sample data so the published dashboard can be explored immediately.
Closes #3. Closes #5.
What to review
load_config,read_results,analyzefor independent model summaries, andcomparefor A/B scores on shared valid coverage. See the API guide.QuickdashWarningby default. Applications can explicitly collect diagnostics; strict mode raisesDiagnosticError. Invalid input/configuration remains an error. Diagnostic conditions, affected identities and calculation effects are tested; message wording is not a cross-language contract.quickdash/python -m quickdashprints trees or JSON. Warnings go to stderr;--strictrejects warning-bearing results. Native Python requires PyYAML, with no Node/pandas runtime. The standalone HTML still operates offline.app/analysis.js; the UI consumes that engine and its diagnostics. The Python builder uses the native library. The Node subprocess bridge is removed.Pages startup data
When
results/contains no CSVs, the main Pages dashboard embedsexamples/sample-evals.csv: the development export's 2,124 metric rows, with the model labelled SAMPLE and original filesystem paths, timestamps, and file-count metadata removed. Scores and scoring settings are unchanged. Shared CSVs automatically replace this fallback. Invalid shared CSVs still fail the build.The browser supplies three clearly labelled synthetic choices: deterministic perturbations with a 2-point standard deviation, plus higher/lower options shifted by ±3 raw percentage points. These are exploration aids, not measured model runs. Sample/fallback handling, synthetic scale preservation, and actual Pages startup are tested.
AGENTS.mdrequires every scoring, interpretation, exclusion, validation, or diagnostic-condition change to add shared regression coverage in both engines, including independent expected values or invariants. The full public sample now runs in CI across all shipped profile/set/mode combinations with independent analysis, matched comparisons, and missing-data comparisons (36 combinations today). Tests discover configuration files instead of assuming their counts.Weighted components
The YAML catalogue defines positive
relative_weightvalues. PolyMath uses 1, 2, 4 and 8, normalized by their sum of 15, within each language/protocol. Complete groups are averaged within the eval before category/English weighting. Expanded views retain individual raw scores and show component contributions.Incompatible component matches, language mappings or named-set selections are errors. Missing results in otherwise valid configurations produce diagnostics and exclude the entire affected group from both comparison scores. Component counts and sums come from configuration.
Behavior comparison and intentional fixes
The pre-library baseline is commit
05acf0bon this branch, which already includes the component work. Relative to that baseline, the full local export agrees across 48 combinations of weighting profiles, eval sets, aggregation modes and missing coverage: included rows, scores, weights, contributions, and descriptive category/language trees. Both baseline and candidate pass the public browser suite. The candidate also passes the full-export browser suite, including sorting, scroll preservation, imports, warnings, missing-field highlights and mobile layout.Relative to
main, PolyMath's configured component weighting is an intentional scoring change. The library refactor preserves the existing scoring policy. Additional fixes found during review:constructorortoString.Warning presentation can differ; shared tests assert codes, contexts and effects. Existing missing-metric highlighting and manually adjusted category-weight warnings are exercised in browser tests.
Validation
The library revision passed CI; the latest sample and parity additions are checked by the required build on this PR.
tests/compare_baseline.cjs.git diff --checkpassed.CI runs the public Python/JavaScript cases, an installed-wheel trial without Node on PATH, browser tests, and standalone builds. Only the sanitized, explicitly requested sample export is added; original private source data and generated review artifacts remain untracked.
Documentation impact
Updated README setup/common commands, configuration and development guides, results/sample instructions, and added the Python API/CLI guide and contributor parity requirements. Building from source now requires installing the Python package; using a generated HTML file still needs only a browser.