Skip to content

Extend the baseline into a versioned discovery and navigation evaluation corpus #409

Description

@Medformatik

Before you start

  • Searched current issues and discussions; no equivalent task was found. Related work is linked below.

Objective

Make the later discovery, enrichment and navigation work measurable with reviewed expectations and reproducible source/configuration metadata.

Scope

Extend the completed small visual baseline, existing search ranking fixtures, Overture quality gates and navigation replays into a documented shared evaluation protocol. Begin with existing online/browser capabilities; add native/offline cases as those capabilities become available. This is not another baseline screenshot-only task or the station ranking fix itself.

Verified current state

Rechecked against main at a8e03bc3efc49207d91f7c1855402c6cf149b8fd on October 7, 2026.

  • Baseline guide contains 20 maps, four search views and two sheets with explicit deployment/data limits; it is intentionally small.
  • Search recorder records API outputs for deterministic ranking evaluation; these outputs are not necessarily raw upstream payloads.
  • Overture evaluation already provides conflation/search quality tooling and reviewed regional release gates.
  • Mobile navigation replay tests already exercise deterministic navigation fixtures. The task coordinates and expands these tools rather than inventing evaluation from scratch.

Completion criteria

  • Define versioned cases for urban/rural browsing, stations, aliases/multilingual names, chain branches, malls/tenants, closed/missing businesses and place-sheet partial states.
  • Record expected entities/coordinates and assessor judgments independently of provider ordering; keep raw upstream responses, adapted API results and final UI ranking separate where diagnosing provider behavior.
  • Pin app/style/source revisions, region/extract dates, provider capabilities and non-secret configuration; explicitly record unavailable deployment/data revisions.
  • Define cold/warm cache isolation and synonym/query-order cases, including the regression tracked in Station searches prefer unrelated POIs and depend on synonym cache order #389 without duplicating its repair scope.
  • Reuse GPS/navigation replays for route alternatives, mode transitions, recovery/GPS loss and off-route behavior; add network-denial offline search/routing cases only when supported and report unavailable cases explicitly.
  • Agree metrics and budgets before comparing changes: useful visible labels/crowding, relevance/duplicates, branch correctness, requests/latency, partial loading and navigation correctness.
  • Provide repeatable commands, fixtures and a before/after report format; meaningful regressions run in appropriate CI without live paid-provider dependency.
  • Document evidence licensing/redaction and operator-controlled live capture; visual snapshots complement behavior assertions rather than serving as ground truth.

Verification approach

Run existing suites and new independently judged fixtures with pinned inputs. Demonstrate a deliberately changed result/order or navigation event is detected; verify replay determinism and redaction. Require a sample comparison report that separates data, provider, normalization, presentation and runtime outcomes.

Dependencies

#393 is the completed starting baseline. #389 remains the dedicated station-search repair. #297 is a semantic-recall experiment, not the shared evaluation protocol. The corpus supports #396/#397/#399 now and #398/#403/#404 as capabilities arrive; it need not wait for all those features.

This is backlog planning, not authorization to expand the implementation beyond this scope. Resolve named decisions during triage before scheduling work.

Activity

  1. added
    documentationImprovements or additions to documentation
    needs-triageNeeds initial review and categorization
    on Oct 6, 2026
  2. added theissue type on Oct 6, 2026
  3. Medformatik commented on Oct 7, 2026

    @Medformatik
    CollaboratorAuthor

    Implemented in #424, branch feat/discovery-evaluation-protocol, head 0725f76cc018764f339a27622372dffa26ee18b9 (capture provenance 318fe89c0, catalog/report command 0ebfcc185, evidence completeness fixes/docs 0725f76cc).

    The protocol reuses current search, cache, conflation, enrichment, Overture-gate and navigation fixtures. Reports pin application/style/input identity and retain raw/adapted/UI distinctions; manual/live/device/offline evidence stays explicitly unavailable. The guide defines independent judgments, predeclared budgets, source/configuration manifests, cache-order isolation and external evidence storage.

    Validation: 26 new RED/GREEN regressions; repeated corpus 99 passing / seven unavailable cases; deliberate missing-station control detected as a case regression and changed input, with exact bytes restored. Full final push checks: 1,612 files / 17,693 tests passed, 138 opt-in tests skipped; lint/types/policy/docs build passed. Fresh review findings were reproduced and fixed. PR CI is running.

    This establishes the shared evaluation protocol; it does not claim current regional recall, live-provider performance or installed/offline readiness.

  4. Medformatik commented on Oct 7, 2026

    @Medformatik
    CollaboratorAuthor

    Completed by #424, squash-merged as 080766a after final independent review and green CI. Completion criteria are checked for the documented pilot scope; installed-device and offline cases remain explicitly unavailable pending #398/#403/#404. Fresh verification passed 53 evaluator tests and the evaluator type check; the documented synthetic comparison detected the intended wrong-branch regression.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationneeds-triageNeeds initial review and categorization

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions