Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 17 additions & 3 deletions .github/workflows/pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,24 @@ jobs:
- uses: actions/setup-node@v7
with:
node-version: '22'
- name: Install the native Python library
run: python3 -m pip install -e .
- name: Check scoring, imports, and configuration
run: |
python3 -m unittest tests.test_data
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs
python3 -m unittest tests.test_data tests.test_engines
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs tests/test_components.cjs
- name: Check the installed package without Node on PATH
run: |
python3 -m pip wheel . --wheel-dir "$RUNNER_TEMP/quickdash-wheels"
python3 -m venv "$RUNNER_TEMP/quickdash-package"
"$RUNNER_TEMP/quickdash-package/bin/python" -m pip install --no-index \
--find-links "$RUNNER_TEMP/quickdash-wheels" oellm-quickdash
cd "$RUNNER_TEMP"
PATH="$RUNNER_TEMP/quickdash-package/bin" \
"$RUNNER_TEMP/quickdash-package/bin/quickdash" "$GITHUB_WORKSPACE/examples/scores.csv" \
--catalogue "$GITHUB_WORKSPACE/configs/examples/catalogue.yaml" \
--weights "$GITHUB_WORKSPACE/configs/examples/weights.yaml" \
--compare 'Example A' 'Example B' --format json > quickdash-installed.json
- name: Check the browser with public fixtures
run: |
google-chrome --headless --no-sandbox --disable-gpu \
Expand All @@ -40,7 +54,7 @@ jobs:
node tests/test_public_browser.mjs || { cat "$RUNNER_TEMP/quickdash-chrome.log"; exit 1; }
- name: Build shared results and the fictional demo
run: |
python3 -m app.build --results-dir results \
python3 -m app.build --results-dir results --sample-csv examples/sample-evals.csv \
--output output/shared > "$RUNNER_TEMP/quickdash-shared.json"
python3 -m app.build examples/scores.csv --catalogue configs/examples/catalogue.yaml \
--weights configs/examples/weights.yaml --eval-set configs/sets/any-available.yaml \
Expand Down
36 changes: 36 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# Working on Quickdash

Quickdash has two implementations of the scoring contract: the native Python
library in `quickdash/` and the browser engine in `app/analysis.js` with its
configuration parsers. Neither implementation is the reference for the other.

## Keep Python and JavaScript in sync

- Every change to scoring, configuration interpretation, input validation,
coverage/exclusion policy, or diagnostic conditions must include a shared
regression case in `tests/test_engines.py` that runs through both engines.
A bug fix should fail before the fix and pass in both implementations afterward.
- Compare scores, trees, effective weights, contributions, included/excluded
measurement identities, and diagnostic codes and context. Human-facing warning
wording does not need to match. Invalid inputs must be rejected by both engines.
- Include independently calculated expectations or invariants: agreement alone
does not prove correctness if both implementations share the same mistake.
- Do not assume fixed eval, category, language, component, profile, or set counts.
Exercise changing contents and sizes. The full sample test discovers shipped
profiles and sets automatically and covers every supported aggregation mode.
- Keep these checks in CI. Do not skip parity tests when changing only one engine,
and do not weaken comparisons merely to accommodate a disagreement.
- For dashboard behavior changes, also extend the public browser tests. Node
exercises the actual browser calculation module; browser tests check that the
UI supplies and renders those calculations correctly.

Run the shared contract and scoring checks before publishing:

```sh
python -m unittest tests.test_data tests.test_engines
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs tests/test_components.cjs
```

See [development and publishing](docs/development.md) for browser checks and
standalone builds. Keep README setup and common commands accurate, and update the
relevant guide when public interfaces or scoring policies change.
37 changes: 28 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,11 @@ A standalone, offline dashboard for comparing model evaluation scores. Explore c

**[Open the dashboard](https://openeurollm.github.io/quickdash/)** or **[try the fictional example](https://openeurollm.github.io/quickdash/demo.html)**. No installation is needed to use either page.

While no shared CSVs have been added to `results/`, the main page opens with our [sample eval export](examples/README.md) and synthetic comparison choices. Shared CSVs replace that fallback automatically on the next successful deployment.

## Compare models

1. Select shared models as A and B, or use **Add model CSV** to open your exports. With one real model, a labelled synthetic comparison is supplied for exploring the interface.
1. Select shared models as A and B, or use **Add model CSV** to open your exports. Three labelled synthetic comparisons (perturbed, higher, and lower scores) are supplied for exploring the interface.
2. Choose a **Weighting profile**. Leave **Eval set** on **Any available** to compare the measurements both models have, or select **flagship-1** to check an expected set. Weights and eval sets are independent. The supplied sets exclude prompted Global PIQA pending scoring validation.
3. Review **Warnings**, then explore the scores and breakdowns. The global catalogue determines how to interpret each eval: category, scoring field, normalization, and language assignments.

Expand All @@ -26,12 +28,15 @@ The [Pages workflow](.github/workflows/pages.yml) publishes only the generated d

## Build a standalone file

Building requires Python 3.8+ and Node.js 18+. The YAML parser is [bundled with its license](app/vendor/README.md); no package installation is needed.
Building requires Python 3.8+ and the Python package below. Node.js is needed only for development tests; the generated dashboard remains standalone and offline.

```sh
git clone https://github.com/OpenEuroLLM/quickdash.git
cd quickdash
python3 -m app.build examples/scores.csv --catalogue configs/examples/catalogue.yaml \
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
python -m app.build examples/scores.csv --catalogue configs/examples/catalogue.yaml \
--weights configs/examples/weights.yaml --eval-set configs/sets/any-available.yaml --output output/example
open output/example/index.html # macOS; elsewhere, open it in your browser
```
Expand All @@ -41,10 +46,10 @@ The example compares two fictional models with a multilingual reasoning eval and
To build the shared dashboard, including all shared results and config choices:

```sh
python3 -m app.build --results-dir results --output output/shared
python3 -m app.build --results-dir results --sample-csv examples/sample-evals.csv --output output/shared
```

An empty `results/` directory produces a page ready for local CSV imports. To start without embedded models regardless of the directory’s contents, omit `--results-dir`.
The sample is used only when `results/` contains no CSVs. Omit `--sample-csv` to leave an empty results directory upload-ready; omit both options to start without embedded models regardless of shared data.

For private exports, put your CSV in the ignored `data/` directory:

Expand All @@ -67,15 +72,29 @@ The builder replaces the files it generates in the chosen output directory, incl

Filters affect inspection views, while composite scores and contribution weights use shared coverage within the selected eval set. A named set with missing requirements is labelled incomplete; it uses the shared subset with redistributed weights. Unmatched measurements are excluded from both compared scores with warnings. Malformed input and invalid selected scores are rejected; failed imports preserve the active dashboard. See the [data-handling policy](docs/configuration.md#data-validation-and-failure-behavior).

Language/category breakdowns show descriptive raw averages. Weighted scores use configured normalization, whose baselines and limitations are visible per eval. Shared numerical scales do not establish comparable difficulty across benchmarks. Unknown/mixed-language scores use the documented English fallback for balancing; this does not change their language labels.
Language/category breakdowns show descriptive raw averages for ordinary evals. Evals with configured components, including PolyMath, show calculated scores when collapsed; expand them to see individual raw scores, relative component weights, and contributions. PolyMath stores `relative_weight` values of 1, 2, 4, and 8, divided by their total of 15 when scoring; incomplete language/protocol groups are excluded with warnings. Incompatible component configurations are rejected before taking effect. See [component aggregation](docs/configuration.md#weighted-components-within-an-eval). Weighted scores use configured normalization, whose baselines and limitations are visible per eval. Shared numerical scales do not establish comparable difficulty across benchmarks. Unknown/mixed-language scores use the documented English fallback for balancing; this does not change their language labels.

## Analyze from Python or the command line

The native Python library uses the same CSVs, YAML rules, weighting profiles, and optional eval sets as the dashboard. It returns calculated trees, effective weights, contributions, coverage, and structured diagnostics. Python emits warnings by default; applications can explicitly collect them or reject results with warnings. See the [Python API and CLI guide](docs/python-api.md).

After installing the package as above:

```sh
quickdash examples/scores.csv --catalogue configs/examples/catalogue.yaml \
--weights configs/examples/weights.yaml --eval-set configs/sets/any-available.yaml \
--compare 'Example A' 'Example B'
```

Add `--format json` for a complete report. Diagnostics go to stderr. Python and browser engines run against the same behavioral test cases.

## Development

Application code lives in `app/`, tests in `tests/`, and contributor documentation in `docs/`. Common checks:
Application code lives in `app/` and `quickdash/`, tests in `tests/`, and contributor documentation in `docs/`. Common checks:

```sh
python3 -m unittest tests.test_data
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs
python -m unittest tests.test_data tests.test_engines
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs tests/test_components.cjs
```

See [development and publishing](docs/development.md) for the source layout, browser tests, and GitHub Pages workflow.
Expand Down
Loading