Skip to content

feat(dashboards): analysis panels, drill-down and deploy context (milestone 2) - #286

Open
vishr wants to merge 26 commits into
mainfrom
feat/agent-dashboards-m2
Open

vishr wants to merge 26 commits into
mainfrom
feat/agent-dashboards-m2

Conversation

@vishr

@vishr vishr commented Oct 7, 2026

Copy link
Copy Markdown
Member

What changed

Milestone 2 of agent-built dashboards. Milestone 1 (#285) gave dashboards a typed spec, six panel types and an agent that writes them. This milestone adds the analysis and investigation layer, so one sentence produces a dashboard you can investigate from, not just look at.

Nine new panel types:

  • heatmap and histogram (log2 duration buckets; metric histograms by per-series increase);
  • scatter, with separate x and y units;
  • state timeline, driven by thresholds;
  • logs, log patterns and traces;
  • service map and service health.

All fifteen types are drawn with ECharts on canvas, in light and dark themes.

Investigation:

  • Click a value to filter the dashboard.
  • A linked crosshair across time charts.
  • Brush to zoom, with one history entry per brush, so Back restores the range.
  • Per-panel time range or shift.
  • A drill drawer addressed by URL. It opens exemplar traces or logs, the shared trace waterfall and correlated logs, and survives reload.

Context on the charts:

  • Deploy markers, derived from service.version first-seen through a new version rollup.
  • Detector anomalies drawn as shaded areas, scoped to the panel's service filters.
  • A before-and-since-deploy split for bar panels.
  • One annotations request per refresh, regardless of panel count.

Tables: column formats (unit, bar, status, sparkline, trace link, service link, log template), with one trend query per panel.

Agent: the create_dashboard guide covers the new types. On replayed demo data with claude-sonnet-5-5, the share of panels that return rows (or explain why they're empty) rose from 90–95% to 97%. Held-out prompts the agent was never tuned on reached 11 of 11 dashboards, up from 10 of 11, with full intent accuracy.

Security: every log read in dashboards, including SQL panels (through AST substitution), now goes through the redacted log source. Before this, dashboards on main could read raw log bodies.

Verification

  • just check passes.

  • just test-race passes.

  • Real-browser verification on about 17M replayed demo records, driven by a scripted Playwright collector in both themes:

    • all 15 types render with accessible names and Inspect;
    • every refresh of a 23-panel dashboard makes exactly one panels request and one annotations request;
    • all 17 interaction checks pass.

    The results are in docs/benchmarks/2026-10-agent-dashboards-m2.md. The check found and fixed defects the unit tests missed:

    • the version rollup skipped compacted batches, which lost deploy history;
    • panels scrolled into view mid-refresh stayed loading forever;
    • per-panel clocks drifted;
    • several presentation issues.
  • Final review:

    • A broad review, plus a focused security and concurrency review. The focused review ran about 100 SQL-panel shapes against the pinned engine and found no path to unredacted bodies and no injection or owner-scope bypass. A concurrent ingest, compaction and rollup probe converged.
    • The findings were fixed in 2c5a686 and re-reviewed.

Checklist:

  • just check
  • just test-race when auth, API, ingest, query, MCP, or agent paths changed
  • User-facing behavior and configuration docs are current (generated docs regenerated and checked)
  • No credentials, private telemetry, host details, or enterprise-only source are included
  • API, migration, ingest, MCP/AG-UI, or release-contract changes are called out below

Contract changes

New routes:

  • POST /api/annotations: deploys and anomalies for a window and service set.
  • POST /api/panels/exemplars: drill exemplars.

Changed routes and MCP:

  • GET /api/observability/trace accepts from and to to bound the trace lookup.
  • The panel query response carries annotation scope and table trends.
  • MCP create_dashboard carries the spec guide; preview_panels points to it. No tools were added or renamed.

Panel spec v1 additions:

  • the nine viz types;
  • histogram, x_unit, drill, click, time;
  • options.split, options.columns, options.highlight, annotations.

Validation now rejects grouped timeseries, and two-dimension bars, with more than one measure; these previously dropped the extra measures silently.

Request decoding: panel request bodies reject trailing data.

Storage:

  • No control-database migrations.
  • New disposable DuckDB read-cache tables: version_rollup, version_rollup_batches and anomaly_log. They rebuild on schema mismatch.

Known limits

  • An explicitly empty multi-value selection does not survive a shared URL. It is rare, and it shows no data either way.
  • Schema-qualified log column references in SQL panels fail with an error rather than binding. They fail closed.
  • Chart clicks and brush are mouse-only. Keyboard access is planned for a later milestone.

vishr added 26 commits October 5, 2026 17:09
Fifteen tasks for drill-down, annotations, interactions and the nine
remaining visualizations, reviewed twice against the merged code and the
pinned engine.
- version_rollup records each service version's first and last sighting
  from new batches in its own analytical table and watermark. It drains
  continuously under a time budget, each pass with its own deadline,
  skips oversized batches instead of reading them unbounded, and carries
  compaction outputs whose inputs are all processed.
- The detector publishes its snapshot first and persists findings in
  the background; anomaly_log coalesces each open episode per service and
  kind, keeps at most 10,000 rows and 30 days, and writes set-based.
- POST /api/annotations returns deploys and anomalies for a window from
  those tables only, with its own error mapping (400, 503, 504).
…d redacted log bodies

- POST /api/panels/exemplars returns up to 20 traces behind one bucket,
  bar or row of a checked panel. The selection is bound, clipped to the
  panel window the user saw, and narrowed to the slowest 1,000 span
  candidates; deadlines map to 504 and cancellations end quietly.
- Every dashboard read of logs now sees redacted bodies, as log search
  already did: structured panels, variable options, empty-panel
  diagnosis, exemplars, and SQL panels, which read a redacted logs
  relation substituted inside the read-only boundary. Before this,
  dashboards could show raw log bodies to any reader.
- Request bodies with trailing data are rejected.
- histogram and heatmap panels bucket span duration_ms into log2
  buckets (negative durations in the first bucket, infinite ones in the
  overflow bucket, NaN excluded) or read metric histograms.
- Cumulative metric histograms count each series' increase over the
  window, clamped at zero on a counter reset; delta series sum; unknown
  temporality is treated as cumulative, the OpenTelemetry default.
- Malformed histogram rows are skipped instead of failing the panel.
- Heatmaps honour the cell budget by dropping their oldest buckets.
- scatter plots two measures per item, coloured by one dimension, with
  log or linear axes and independent x and y units (x_unit), as when
  plotting rows returned against query duration.
- state_timeline grades each item per bucket by thresholds, which it
  requires; share() divides by every item's traffic before the top N are
  chosen, so shown shares are not inflated.
- The agent's guide explains scales, units, ranking and what these panel
  types do not draw.
- logs streams redacted log rows, newest first, with an optional
  highlight matched against redacted text; bodies are capped at 2,000
  characters after redaction.
- log_patterns ranks body_template groups by count (up to 50) with a
  bounded trend per pattern and reports the interval it really used.
- traces lists the slowest or the slowest erroring traces from at most
  1,001 candidates chosen by the sort, so cost stays bounded.
- Validation names the exact field with a hint, rejects options that do
  not apply, and empty panels say whether a highlight or the erroring
  condition emptied them.
- service_map draws services and their calls from the shared topology
  read, with p95 and span counts on nodes and calls, latency and error
  rate on edges; failing edges are thicker and darker (#232 item 22).
- health ports the old overview widget onto a panel frame: health and
  error-rate tiles, the error trend and the health distribution with
  Fanout's health shapes.
- Both accept namespace and service equality filters, reuse the
  observability services without new SQL, and cap at 400 services.
…d health panels

Six chart panels render the server's distribution, item and rollup
frames on canvas ECharts: numeric bucket order, per-axis units on
scatter, threshold-graded timelines with health shapes (unknown cells
stay unknown), service maps whose failing edges (5% errors or more) are
thicker and darker, and a health panel ported from the old overview.
Each has an Inspect table, an accessible summary and memoized options.
The row panels render through the shared table: a log stream with
severity badges and a highlight that matches like the server (escaped,
case-insensitive, accent-sensitive), log patterns with counts and a trend
labelled at the interval the server used, and trace lists. Trace IDs are
accessible text until the drill drawer links them.
…nel time

- Clicking a bar, series or row of a panel with click.set_variable sets
  that variable, shown as a removable chip; the grouped Other series is
  never a filter value.
- Time panels share a crosshair and can be brushed to zoom: one history
  entry per selection, the selection cleared afterwards, the range kept
  as absolute URL times.
- A panel's own time range or shift drives its read window and its
  comparison, and the panel shows a badge; results report their window.
- Only actionable panels look clickable.
…terfall and its logs

- Panels with drill open a drawer of the traces (or logs) behind the
  clicked bucket, bar or row, then the trace's waterfall and logs. The
  target lives in the URL (drill=) and survives the router, so a copied
  link reopens it, using the window captured at click time.
- Row-panel trace IDs open the same drawer.
- GET /api/observability/trace accepts an absolute from/to window,
  bounded by retention, so older traces and stale links still resolve.
- The waterfall is an adapted copy of the chat view's, kept inside the
  host bundle.
- A bar with split: deploy shows each category before and after the
  latest deploy of the filtered service inside the panel window, each
  side normalised over its own sub-window. With service set to All, or
  no deploy in the window, it falls back to the plain bar with a note.
- Each time panel reports which services its deploy and anomaly markers
  cover, computed once per checked spec from the version, anomaly and
  service rollups, so services that never report a version still get
  their anomaly markers.
- Each refresh makes one annotations request covering the visible time
  panels' windows (none when no time panel is shown); deploys draw as
  dashed lines and anomalies as shaded bands, each with a tooltip, on
  panels whose annotation scope matches, or on all time panels when the
  panel has no service scope. Annotation failures leave the chart intact
  with a quiet note.
- The drill drawer announces loading, keeps waterfall rows inert, and
  offers Back to traces.
- Dashboard tests wait for requests to finish instead of a fixed delay,
  which made them fail under load.
- Table columns can render as a bar, a threshold status (shape and
  text, never colour alone), a per-row sparkline, a trace link that opens
  the drill drawer, a service link that sets the service variable, or a
  log template with its placeholders marked.
- The executor attaches every row's trend in one query per panel,
  within the response budget; trimmed trends and failures are noted
  without hiding other notes.
- Stat panels draw a sparkline.
The spec guide on create_dashboard now covers only what the schema does
not say about the nine new panel types, drill, annotations, the deploy
split, table formats, click and panel time, with two short reference
examples; preview_panels points to it instead of repeating it. Together
the descriptions are smaller than before this milestone. The agent prompt
adds when to drill, annotate or split, to use only the types a question
needs, and never to name schema fields to the user.

On replayed demo data with claude-sonnet-5-5, panels that return rows or
explain their emptiness rose from 90-95% to 97%, and held-out requests
the agent was never tuned on reached 11 of 11 dashboards and full intent
accuracy, up from 10 of 11.
Driving the milestone 2 dashboards in a real browser on replayed demo data
found defects the unit tests missed:

- The version rollup skipped any batch over 64,000 rows, which includes
  every compaction output. Versions in that data were lost and every
  dashboard showed "Annotation history is limited" forever. An oversized
  batch now gets a pass of its own, oldest first; a batch that keeps
  failing backs off without blocking newer ones.
- A panel scrolled into view while a batch was in flight never loaded.
  A panel the server leaves out now shows an error instead of spinning,
  and lazy batches keep the dashboard's markers.
- Panels in one batch read the clock separately, so a shifted panel was
  off by milliseconds. One instant now serves the whole batch.
- An empty highlighted logs panel blamed the highlight even when a filter
  caused it.
- Polish: status cells use the measure's unit; the traces panel truncates
  trace IDs and shows Error/OK; sparkline cells show their value;
  heatmap and histogram buckets carry units; the heatmap colour scale is
  capped so the distribution shows; single-series charts drop the legend;
  day boundaries read "Oct 6"; deploy labels stay clear of the zoom tools.

Adds the evidence gate that the verification report is generated from.
Long bucket labels narrow the heatmap's plot, and its two-hourly time
labels ran together. Overlapping labels are now hidden.
All fifteen panel types render with accessible names and Inspect in both
themes; every refresh of a 23-panel dashboard makes exactly one panels
request and one annotations request; all seventeen interaction checks
pass on replayed demo data. Lists the defects the check found and fixed.
- Grouped timeseries and two-dimension bars accept one measure; extra
  measures were silently dropped, so validation now says so.
- Empty variable options serialize as [] and the dashboard tolerates
  null, so a saved selection no longer crashes an empty window.
- Bar drills carry every grouping value, and deploy-split bars drill
  into the selected period rather than the whole window.
- Share sparklines in limited tables divide by the whole scope, not by
  the rows shown.
- A version rollup pass that fails retries its batches singly, so a bad
  batch no longer holds back healthy ones; waiting for the write gate
  no longer counts against the pass budget; idle passes no longer
  write; rebuilding either version table rebuilds both.
- Annotations read deploys, anomalies and the history flag in one
  transaction.
- SQL panel CTEs may not shadow telemetry relations.
- A refresh with no visible panels requests nothing.
…view

Re-run on the build with the final-review fixes: all fifteen panel types
in both themes, one panels and one annotations request per refresh, and
all seventeen interaction checks. Summarises the final review and what
was deferred.
Using the agent in a real browser showed problems around it:

- A dashboard the agent created or edited had no link in the chat and
  did not appear in the sidebar until a reload. The chat now shows a
  compact "Created"/"Updated" card that opens it, and the list
  refreshes at once.
- Opening a drill, changing a variable or brushing scrolled the
  dashboard to the top. In-page changes now keep the scroll position.
- Chart panels were a few pixels taller than their box, which painted a
  scrollbar in every chart on systems that always show scrollbars and
  clipped the deploy-split axis. Charts now size to their body.
- The deploy-split note showed a raw nanosecond timestamp; it now reads
  in the viewer's locale.
- The grid could first lay out at a guessed 1280px width. It now waits
  for its container's measured width and follows it.
Series colours came from six shades of two hue families, assigned by a
hash of the series name, so services in one chart often shared a colour
and failed colour-vision checks outright (lavender against blue measured
ΔE 0.6 for protanopia). Charts now use a validated six-slot palette -
blue, orange, aqua, violet, yellow, magenta - that passes the lightness,
chroma, adjacent-pair colour-vision, normal-vision and contrast checks
on both themes, assigned in each chart's series order.

A chart shows at most six series: `top` defaults to six and may not
exceed it. A server-computed Other is drawn muted; further series are
left out with a note rather than summed, since p95s and rates do not
add up.
A manual refresh pressed while a partial lazy batch was in flight was
dropped, so the panels outside that batch were never refreshed. It now
runs once the lazy batch settles; during a full refresh it is still
covered by the batch already running. A page test that clicked a chart
before the measured-width grid had rendered now waits for it.
The dashboards were used by hand in Chrome with both providers. Records
what that found and fixed beyond the scripted checks.
The fixture built its old episodes with TIMESTAMP_NS minus an INTERVAL,
which DuckDB evaluates at microsecond precision. On Linux, whose clock
has nanoseconds, the merged start no longer matched the bound value and
the test failed in CI while passing on macOS. Bind the start directly,
and pin a nanosecond component so a microsecond clock exercises it too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant