Repository navigation
feat(eval): add reproducible discovery evaluation reports - #424
Conversation
|
Completed the remaining #409 evaluation work in af1aff3. The two generated reports below are a synthetic harness control, not a product improvement. Both ran on the clean committed tree: the before command exits 0; the after command intentionally exits 1 when only the normalized REWE branch is changed. Data/provider projections, presentation judgment and navigation replay results stay unchanged. The real read-only pilot is separate: adapted API captures and inspected browser labels, uncontrolled API cache, unknown deployment/geocoder revisions, no raw-upstream or installed-device claim. Its misses demonstrate useful evaluation outcomes rather than needing search/product fixes in this PR.
All four viewports: 430×932, DPR 1, English, dark, zoom 15, zero pitch/bearing; zero observed POI/park label overlaps. Rendered feature counts were not treated as readable-label counts. Operator capture files/screenshots stay outside Git. Public identity judgments and source-specific rights/redaction limits are documented in the developer guide. Before: generated reportMarkdown SHA-256: Discovery evaluation reportProtocol: 2; catalog: cdeba7569bffeb983bf5a1580d80d456d24e67069ef853134bc7119a0330ae18 Automated cases are offline recorded/synthetic/contract evidence. Optional reviewed observations retain their own live/recorded/synthetic provenance; absent layers and installed-device cases are unavailable. Assertion counts can overlap between cases; they are not independent observations or production performance metrics.
Independently reviewed observationsDefinitions and frozen pilot budgets: ee56f34187b97bd5f7f34dafcb8693ae1d4a6bfa754a027400dcdfa0a51f8a40; 41 case/layer observations unavailable. {
"region": "Germany: Berlin pilot",
"extractDate": {
"value": null,
"reason": "Synthetic records, no extract used"
},
"style": {
"value": null,
"reason": "Synthetic display judgment, no rendered style used"
},
"deployment": {
"value": null,
"reason": "No deployment used in this synthetic example"
},
"sources": {
"osm": {
"value": null,
"reason": "Coordinates from reviewed case, no imported OSM generation"
},
"overture": {
"value": null,
"reason": "Not used"
}
},
"provider": {
"id": "synthetic-control",
"capabilities": [
"projected search rows"
]
},
"configuration": {
"language": "en",
"theme": "dark",
"viewport": [
430,
932
],
"dpr": 1
},
"queryOrder": [
"business/rewe-invalidenstrasse"
],
"cacheIsolation": "isolated"
}
Absent layers (listed individually in report.json):
After: generated comparison report (intentional failure)Markdown SHA-256: Discovery evaluation reportProtocol: 2; catalog: cdeba7569bffeb983bf5a1580d80d456d24e67069ef853134bc7119a0330ae18 Automated cases are offline recorded/synthetic/contract evidence. Optional reviewed observations retain their own live/recorded/synthetic provenance; absent layers and installed-device cases are unavailable. Assertion counts can overlap between cases; they are not independent observations or production performance metrics. Comparison{
"reviewed": {
"contextChanged": false,
"captureConditionsChanged": false,
"changedProviderInputs": [],
"regressions": [
"business/rewe-invalidenstrasse/normalization"
],
"changes": [
{
"caseId": "business/rewe-invalidenstrasse",
"layer": "normalization",
"metric": "recall",
"before": 1,
"after": 0
},
{
"caseId": "business/rewe-invalidenstrasse",
"layer": "normalization",
"metric": "firstRank",
"before": 1,
"after": null
},
{
"caseId": "business/rewe-invalidenstrasse",
"layer": "normalization",
"metric": "wrongBranch",
"before": 0,
"after": 1
}
]
},
"regressions": [],
"improvements": [],
"changedInputs": [],
"appOnlyComparison": false,
"runRegression": true
}Changed inputs and dirty trees prevent an application-only comparison; comparisons do not establish causality.
Independently reviewed observationsDefinitions and frozen pilot budgets: ee56f34187b97bd5f7f34dafcb8693ae1d4a6bfa754a027400dcdfa0a51f8a40; 41 case/layer observations unavailable. {
"region": "Germany: Berlin pilot",
"extractDate": {
"value": null,
"reason": "Synthetic records, no extract used"
},
"style": {
"value": null,
"reason": "Synthetic display judgment, no rendered style used"
},
"deployment": {
"value": null,
"reason": "No deployment used in this synthetic example"
},
"sources": {
"osm": {
"value": null,
"reason": "Coordinates from reviewed case, no imported OSM generation"
},
"overture": {
"value": null,
"reason": "Not used"
}
},
"provider": {
"id": "synthetic-control",
"capabilities": [
"projected search rows"
]
},
"configuration": {
"language": "en",
"theme": "dark",
"viewport": [
430,
932
],
"dpr": 1
},
"queryOrder": [
"business/rewe-invalidenstrasse"
],
"cacheIsolation": "isolated"
}
Absent layers (listed individually in report.json):
|
Summary
pnpm discovery-eval, with versioned JSON/Markdown reports and application/input/style fingerprints.Related issues
Fixes #409
Supports #397 and later discovery/navigation work. Station repair remains #389 / #392. Installed navigation/offline evidence remains explicitly unavailable until #398 / #403 / #404 provide those capabilities.
How was this tested?
pnpm lint, fullpnpm check-types,pnpm check:policy, fullpnpm test(1,614 files / 17,720 tests passed; 138 opt-in tests skipped), andpnpm -C docs buildverified locally. Final independent review found no remaining Critical/Important issues after three integrity findings were reproduced and fixed.Legacy search fixtures retain unknown provenance. This is a bounded evaluation pilot, not population-wide coverage or native/offline readiness. No visual product behavior changed; no screenshots or capture archives added to Git. Sample full generated comparison is provided in a PR comment and repeatable commands/schema fixtures are in the developer guide.
Checklist
pnpm lint && pnpm check-types && pnpm testpass locally.envfiles committed