Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions .github/workflows/pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ jobs:
- name: Check scoring, imports, and configuration
run: |
python3 -m unittest tests.test_data
node --test tests/test_data.cjs tests/test_yaml.cjs
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs
- name: Check the browser with public fixtures
run: |
google-chrome --headless --no-sandbox --disable-gpu \
Expand All @@ -40,9 +40,10 @@ jobs:
node tests/test_public_browser.mjs || { cat "$RUNNER_TEMP/quickdash-chrome.log"; exit 1; }
- name: Build shared results and the fictional demo
run: |
python3 -m app.build --results-dir results --configs-dir configs \
python3 -m app.build --results-dir results \
--output output/shared > "$RUNNER_TEMP/quickdash-shared.json"
python3 -m app.build examples/scores.csv --config configs/example.yaml \
python3 -m app.build examples/scores.csv --catalogue configs/examples/catalogue.yaml \
--weights configs/examples/weights.yaml --eval-set configs/sets/any-available.yaml \
--output output/demo > "$RUNNER_TEMP/quickdash-demo.json"
mkdir -p output/site
cp output/shared/index.html output/site/index.html
Expand Down
22 changes: 13 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,16 +6,19 @@ A standalone, offline dashboard for comparing model evaluation scores. Explore c

## Compare models

1. Choose an **Eval configuration** at the top. The supplied OELLM config defines evals, selected metrics, normalization, category weights, and explicit language assignments.
2. Select shared models as A and B, or use **Add model CSV** to open your own exports. With one real model, a clearly labelled synthetic comparison is supplied for exploring the interface.
3. Review **Warnings**, then explore the score and breakdown tabs. To use a temporary YAML config, open **Eval configuration → Load config**.
1. Select shared models as A and B, or use **Add model CSV** to open your exports. With one real model, a labelled synthetic comparison is supplied for exploring the interface.
2. Choose a **Weighting profile**. Leave **Eval set** on **Any available** to compare the measurements both models have, or select **flagship-1** to check an expected set. Weights and eval sets are independent. The supplied sets exclude prompted Global PIQA pending scoring validation.
3. Review **Warnings**, then explore the scores and breakdowns. The global catalogue determines how to interpret each eval: category, scoring field, normalization, and language assignments.

Files opened in the dashboard stay in your browser; they are not uploaded. Changes last until reload, which restores the published models and settings. **Clear models** removes the loaded models from your session while keeping config choices. **Export config YAML** saves edited scoring settings, but does not save model data or modify the repository.
**Original** is the startup weighting profile. **Code & math emphasis** gives Code and Math 20% each, with the other category weights adjusted as shown in the [configuration reference](docs/configuration.md#choose-weights-and-expected-coverage).

Files opened here stay in your browser; they are not uploaded. Changes last until reload. **Clear models** removes the loaded models while keeping settings. Under **Eval configuration**, load or export the catalogue, weights, and eval set as separate YAML files. Weight exports include your edits and active score calculation.

## Share results and scoring configs

- Add public CSV exports to [results/](results/README.md) to offer their models in the shared dashboard.
- Add YAML scoring configs to [configs/](configs/README.md) to offer them in the configuration selector. Each config needs a unique `name`. [configs/default.txt](configs/default.txt) names the one selected on startup, currently `oellm.yaml`.
- Add interpretation rules to [configs/catalogue.yaml](configs/catalogue.yaml).
- Add weighting profiles to [configs/weights/](configs/weights/), or optional named eval sets to [configs/sets/](configs/sets/). Each directory has a `default.txt` choosing its startup selection. See [contributing configs](configs/README.md).

Use a pull request or GitHub’s **Add file → Upload files**. Changes on `main` trigger tests and a GitHub Pages rebuild; pull requests are checked without publishing. Invalid inputs stop the update and leave the last successful site online. The repository and dashboard are public, so use browser imports for private comparisons.

Expand All @@ -28,7 +31,8 @@ Building requires Python 3.8+ and Node.js 18+. The YAML parser is [bundled with
```sh
git clone https://github.com/OpenEuroLLM/quickdash.git
cd quickdash
python3 -m app.build examples/scores.csv --config configs/example.yaml --output output/example
python3 -m app.build examples/scores.csv --catalogue configs/examples/catalogue.yaml \
--weights configs/examples/weights.yaml --eval-set configs/sets/any-available.yaml --output output/example
open output/example/index.html # macOS; elsewhere, open it in your browser
```

Expand All @@ -47,7 +51,7 @@ For private exports, put your CSV in the ignored `data/` directory:
```sh
mkdir -p data
# Copy your export to data/evals.csv, then:
python3 -m app.build data/evals.csv --config configs/oellm.yaml --output output/private
python3 -m app.build data/evals.csv --output output/private
```

The builder replaces the files it generates in the chosen output directory, including audits and score summaries. Use separate output directories to keep builds. `data/` and `output/` are ignored by Git. CSV columns and configuration rules are described in the [configuration reference](docs/configuration.md).
Expand All @@ -61,7 +65,7 @@ The builder replaces the files it generates in the chosen output directory, incl
- **Eval configuration:** inspect every task's category, languages, selected field, raw alternate fields, normalization, and source notes. Load or export YAML here.
- **Warnings:** review missing coverage, scoring inconsistencies, language fallbacks, sample-count issues, and caveats stored in the config.

Filters affect inspection views, while composite scores and contribution weights use the models' full shared coverage. Unmatched measurements are excluded from both compared scores with warnings. Malformed input and invalid selected scores are rejected; failed imports preserve the active dashboard. See the [data-handling policy](docs/configuration.md#data-validation-and-failure-behavior).
Filters affect inspection views, while composite scores and contribution weights use shared coverage within the selected eval set. A named set with missing requirements is labelled incomplete; it uses the shared subset with redistributed weights. Unmatched measurements are excluded from both compared scores with warnings. Malformed input and invalid selected scores are rejected; failed imports preserve the active dashboard. See the [data-handling policy](docs/configuration.md#data-validation-and-failure-behavior).

Language/category breakdowns show descriptive raw averages. Weighted scores use configured normalization, whose baselines and limitations are visible per eval. Shared numerical scales do not establish comparable difficulty across benchmarks. Unknown/mixed-language scores use the documented English fallback for balancing; this does not change their language labels.

Expand All @@ -71,7 +75,7 @@ Application code lives in `app/`, tests in `tests/`, and contributor documentati

```sh
python3 -m unittest tests.test_data
node --test tests/test_data.cjs tests/test_yaml.cjs
node --test tests/test_data.cjs tests/test_yaml.cjs tests/test_suites.cjs tests/test_warning_policy.cjs
```

See [development and publishing](docs/development.md) for the source layout, browser tests, and GitHub Pages workflow.
Expand Down
Loading