Add configurable component weights so PolyMath contributes its difficulty-weighted score while preserving individual results for inspection.
Status: implemented and awaiting review in draft PR #4. Owner: implementation agent; maintainer reviews the PR. Next action: review the component policy and UI, then merge when accepted. Independent of #2 (overlapping original/multilingual evals).
PolyMath's official per-language score is (low + 2*medium + 4*high + 8*top)/15. The oellm-eval task template emits unweighted per-split mean accuracy; the weighting must happen when combining difficulties. Store relative_weight values (1, 2, 4, 8); divide by their sum to obtain the effective shares (1/15, 2/15, 4/15, 8/15).
Scope and acceptance:
Sources:
Validation at commit 82828aa: 36 Python tests, 99 Node tests, public-fixture and full-export browser suites, and local documentation links passed. PR CI passed, including public tests, browser checks, and shared/demo builds without private exports. No production deployment has been performed; the issue stays open until the PR is accepted and merged.
The relative_weight rename at 4886d35 passed 17 public Python tests, 97 public Node tests, the public browser suite, and CI. Scoring is unchanged; component weight is rejected by schema validation.
Compatibility validation at 05acf0b passed 36 Python tests, 102 Node tests, and both browser suites. Named sets cannot select partial component groups or incompatible shot settings; valid selections with missing CSV results retain the warning/exclusion policy.
Compatibility-check CI passed; deployment was skipped because PR #4 remains draft and unmerged.
Add configurable component weights so PolyMath contributes its difficulty-weighted score while preserving individual results for inspection.
Status: implemented and awaiting review in draft PR #4. Owner: implementation agent; maintainer reviews the PR. Next action: review the component policy and UI, then merge when accepted. Independent of #2 (overlapping original/multilingual evals).
PolyMath's official per-language score is
(low + 2*medium + 4*high + 8*top)/15. The oellm-eval task template emits unweighted per-split mean accuracy; the weighting must happen when combining difficulties. Storerelative_weightvalues (1, 2, 4, 8); divide by their sum to obtain the effective shares (1/15, 2/15, 4/15, 8/15).Scope and acceptance:
Sources:
Validation at commit
82828aa: 36 Python tests, 99 Node tests, public-fixture and full-export browser suites, and local documentation links passed. PR CI passed, including public tests, browser checks, and shared/demo builds without private exports. No production deployment has been performed; the issue stays open until the PR is accepted and merged.The
relative_weightrename at4886d35passed 17 public Python tests, 97 public Node tests, the public browser suite, and CI. Scoring is unchanged; componentweightis rejected by schema validation.Compatibility validation at
05acf0bpassed 36 Python tests, 102 Node tests, and both browser suites. Named sets cannot select partial component groups or incompatible shot settings; valid selections with missing CSV results retain the warning/exclusion policy.Compatibility-check CI passed; deployment was skipped because PR #4 remains draft and unmerged.