Repository navigation
Conversation
Fifteen tasks for drill-down, annotations, interactions and the nine remaining visualizations, reviewed twice against the merged code and the pinned engine.
- version_rollup records each service version's first and last sighting from new batches in its own analytical table and watermark. It drains continuously under a time budget, each pass with its own deadline, skips oversized batches instead of reading them unbounded, and carries compaction outputs whose inputs are all processed. - The detector publishes its snapshot first and persists findings in the background; anomaly_log coalesces each open episode per service and kind, keeps at most 10,000 rows and 30 days, and writes set-based. - POST /api/annotations returns deploys and anomalies for a window from those tables only, with its own error mapping (400, 503, 504).
…d redacted log bodies - POST /api/panels/exemplars returns up to 20 traces behind one bucket, bar or row of a checked panel. The selection is bound, clipped to the panel window the user saw, and narrowed to the slowest 1,000 span candidates; deadlines map to 504 and cancellations end quietly. - Every dashboard read of logs now sees redacted bodies, as log search already did: structured panels, variable options, empty-panel diagnosis, exemplars, and SQL panels, which read a redacted logs relation substituted inside the read-only boundary. Before this, dashboards could show raw log bodies to any reader. - Request bodies with trailing data are rejected.
- histogram and heatmap panels bucket span duration_ms into log2 buckets (negative durations in the first bucket, infinite ones in the overflow bucket, NaN excluded) or read metric histograms. - Cumulative metric histograms count each series' increase over the window, clamped at zero on a counter reset; delta series sum; unknown temporality is treated as cumulative, the OpenTelemetry default. - Malformed histogram rows are skipped instead of failing the panel. - Heatmaps honour the cell budget by dropping their oldest buckets.
- scatter plots two measures per item, coloured by one dimension, with log or linear axes and independent x and y units (x_unit), as when plotting rows returned against query duration. - state_timeline grades each item per bucket by thresholds, which it requires; share() divides by every item's traffic before the top N are chosen, so shown shares are not inflated. - The agent's guide explains scales, units, ranking and what these panel types do not draw.
- logs streams redacted log rows, newest first, with an optional highlight matched against redacted text; bodies are capped at 2,000 characters after redaction. - log_patterns ranks body_template groups by count (up to 50) with a bounded trend per pattern and reports the interval it really used. - traces lists the slowest or the slowest erroring traces from at most 1,001 candidates chosen by the sort, so cost stays bounded. - Validation names the exact field with a hint, rejects options that do not apply, and empty panels say whether a highlight or the erroring condition emptied them.
- service_map draws services and their calls from the shared topology read, with p95 and span counts on nodes and calls, latency and error rate on edges; failing edges are thicker and darker (#232 item 22). - health ports the old overview widget onto a panel frame: health and error-rate tiles, the error trend and the health distribution with Fanout's health shapes. - Both accept namespace and service equality filters, reuse the observability services without new SQL, and cap at 400 services.
…d health panels Six chart panels render the server's distribution, item and rollup frames on canvas ECharts: numeric bucket order, per-axis units on scatter, threshold-graded timelines with health shapes (unknown cells stay unknown), service maps whose failing edges (5% errors or more) are thicker and darker, and a health panel ported from the old overview. Each has an Inspect table, an accessible summary and memoized options.
The row panels render through the shared table: a log stream with severity badges and a highlight that matches like the server (escaped, case-insensitive, accent-sensitive), log patterns with counts and a trend labelled at the interval the server used, and trace lists. Trace IDs are accessible text until the drill drawer links them.
…nel time - Clicking a bar, series or row of a panel with click.set_variable sets that variable, shown as a removable chip; the grouped Other series is never a filter value. - Time panels share a crosshair and can be brushed to zoom: one history entry per selection, the selection cleared afterwards, the range kept as absolute URL times. - A panel's own time range or shift drives its read window and its comparison, and the panel shows a badge; results report their window. - Only actionable panels look clickable.
…terfall and its logs - Panels with drill open a drawer of the traces (or logs) behind the clicked bucket, bar or row, then the trace's waterfall and logs. The target lives in the URL (drill=) and survives the router, so a copied link reopens it, using the window captured at click time. - Row-panel trace IDs open the same drawer. - GET /api/observability/trace accepts an absolute from/to window, bounded by retention, so older traces and stale links still resolve. - The waterfall is an adapted copy of the chat view's, kept inside the host bundle.
- A bar with split: deploy shows each category before and after the latest deploy of the filtered service inside the panel window, each side normalised over its own sub-window. With service set to All, or no deploy in the window, it falls back to the plain bar with a note. - Each time panel reports which services its deploy and anomaly markers cover, computed once per checked spec from the version, anomaly and service rollups, so services that never report a version still get their anomaly markers.
- Each refresh makes one annotations request covering the visible time panels' windows (none when no time panel is shown); deploys draw as dashed lines and anomalies as shaded bands, each with a tooltip, on panels whose annotation scope matches, or on all time panels when the panel has no service scope. Annotation failures leave the chart intact with a quiet note. - The drill drawer announces loading, keeps waterfall rows inert, and offers Back to traces. - Dashboard tests wait for requests to finish instead of a fixed delay, which made them fail under load.
- Table columns can render as a bar, a threshold status (shape and text, never colour alone), a per-row sparkline, a trace link that opens the drill drawer, a service link that sets the service variable, or a log template with its placeholders marked. - The executor attaches every row's trend in one query per panel, within the response budget; trimmed trends and failures are noted without hiding other notes. - Stat panels draw a sparkline.
The spec guide on create_dashboard now covers only what the schema does not say about the nine new panel types, drill, annotations, the deploy split, table formats, click and panel time, with two short reference examples; preview_panels points to it instead of repeating it. Together the descriptions are smaller than before this milestone. The agent prompt adds when to drill, annotate or split, to use only the types a question needs, and never to name schema fields to the user. On replayed demo data with claude-sonnet-5-5, panels that return rows or explain their emptiness rose from 90-95% to 97%, and held-out requests the agent was never tuned on reached 11 of 11 dashboards and full intent accuracy, up from 10 of 11.
Driving the milestone 2 dashboards in a real browser on replayed demo data found defects the unit tests missed: - The version rollup skipped any batch over 64,000 rows, which includes every compaction output. Versions in that data were lost and every dashboard showed "Annotation history is limited" forever. An oversized batch now gets a pass of its own, oldest first; a batch that keeps failing backs off without blocking newer ones. - A panel scrolled into view while a batch was in flight never loaded. A panel the server leaves out now shows an error instead of spinning, and lazy batches keep the dashboard's markers. - Panels in one batch read the clock separately, so a shifted panel was off by milliseconds. One instant now serves the whole batch. - An empty highlighted logs panel blamed the highlight even when a filter caused it. - Polish: status cells use the measure's unit; the traces panel truncates trace IDs and shows Error/OK; sparkline cells show their value; heatmap and histogram buckets carry units; the heatmap colour scale is capped so the distribution shows; single-series charts drop the legend; day boundaries read "Oct 6"; deploy labels stay clear of the zoom tools. Adds the evidence gate that the verification report is generated from.
Long bucket labels narrow the heatmap's plot, and its two-hourly time labels ran together. Overlapping labels are now hidden.
All fifteen panel types render with accessible names and Inspect in both themes; every refresh of a 23-panel dashboard makes exactly one panels request and one annotations request; all seventeen interaction checks pass on replayed demo data. Lists the defects the check found and fixed.
- Grouped timeseries and two-dimension bars accept one measure; extra measures were silently dropped, so validation now says so. - Empty variable options serialize as [] and the dashboard tolerates null, so a saved selection no longer crashes an empty window. - Bar drills carry every grouping value, and deploy-split bars drill into the selected period rather than the whole window. - Share sparklines in limited tables divide by the whole scope, not by the rows shown. - A version rollup pass that fails retries its batches singly, so a bad batch no longer holds back healthy ones; waiting for the write gate no longer counts against the pass budget; idle passes no longer write; rebuilding either version table rebuilds both. - Annotations read deploys, anomalies and the history flag in one transaction. - SQL panel CTEs may not shadow telemetry relations. - A refresh with no visible panels requests nothing.
…view Re-run on the build with the final-review fixes: all fifteen panel types in both themes, one panels and one annotations request per refresh, and all seventeen interaction checks. Summarises the final review and what was deferred.
Using the agent in a real browser showed problems around it: - A dashboard the agent created or edited had no link in the chat and did not appear in the sidebar until a reload. The chat now shows a compact "Created"/"Updated" card that opens it, and the list refreshes at once. - Opening a drill, changing a variable or brushing scrolled the dashboard to the top. In-page changes now keep the scroll position. - Chart panels were a few pixels taller than their box, which painted a scrollbar in every chart on systems that always show scrollbars and clipped the deploy-split axis. Charts now size to their body. - The deploy-split note showed a raw nanosecond timestamp; it now reads in the viewer's locale. - The grid could first lay out at a guessed 1280px width. It now waits for its container's measured width and follows it.
Series colours came from six shades of two hue families, assigned by a hash of the series name, so services in one chart often shared a colour and failed colour-vision checks outright (lavender against blue measured ΔE 0.6 for protanopia). Charts now use a validated six-slot palette - blue, orange, aqua, violet, yellow, magenta - that passes the lightness, chroma, adjacent-pair colour-vision, normal-vision and contrast checks on both themes, assigned in each chart's series order. A chart shows at most six series: `top` defaults to six and may not exceed it. A server-computed Other is drawn muted; further series are left out with a note rather than summed, since p95s and rates do not add up.
A manual refresh pressed while a partial lazy batch was in flight was dropped, so the panels outside that batch were never refreshed. It now runs once the lazy batch settles; during a full refresh it is still covered by the batch already running. A page test that clicked a chart before the measured-width grid had rendered now waits for it.
The dashboards were used by hand in Chrome with both providers. Records what that found and fixed beyond the scripted checks.
The fixture built its old episodes with TIMESTAMP_NS minus an INTERVAL, which DuckDB evaluates at microsecond precision. On Linux, whose clock has nanoseconds, the merged start no longer matched the bound value and the test failed in CI while passing on macOS. Bind the start directly, and pin a nanosecond component so a microsecond clock exercises it too.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Milestone 2 of agent-built dashboards. Milestone 1 (#285) gave dashboards a typed spec, six panel types and an agent that writes them. This milestone adds the analysis and investigation layer, so one sentence produces a dashboard you can investigate from, not just look at.
Nine new panel types:
All fifteen types are drawn with ECharts on canvas, in light and dark themes.
Investigation:
Context on the charts:
service.versionfirst-seen through a new version rollup.Tables: column formats (unit, bar, status, sparkline, trace link, service link, log template), with one trend query per panel.
Agent: the
create_dashboardguide covers the new types. On replayed demo data withclaude-sonnet-5-5, the share of panels that return rows (or explain why they're empty) rose from 90–95% to 97%. Held-out prompts the agent was never tuned on reached 11 of 11 dashboards, up from 10 of 11, with full intent accuracy.Security: every log read in dashboards, including SQL panels (through AST substitution), now goes through the redacted log source. Before this, dashboards on main could read raw log bodies.
Verification
just checkpasses.just test-racepasses.Real-browser verification on about 17M replayed demo records, driven by a scripted Playwright collector in both themes:
The results are in
docs/benchmarks/2026-10-agent-dashboards-m2.md. The check found and fixed defects the unit tests missed:Final review:
Checklist:
just checkjust test-racewhen auth, API, ingest, query, MCP, or agent paths changedContract changes
New routes:
POST /api/annotations: deploys and anomalies for a window and service set.POST /api/panels/exemplars: drill exemplars.Changed routes and MCP:
GET /api/observability/traceacceptsfromandtoto bound the trace lookup.create_dashboardcarries the spec guide;preview_panelspoints to it. No tools were added or renamed.Panel spec v1 additions:
histogram,x_unit,drill,click,time;options.split,options.columns,options.highlight,annotations.Validation now rejects grouped timeseries, and two-dimension bars, with more than one measure; these previously dropped the extra measures silently.
Request decoding: panel request bodies reject trailing data.
Storage:
version_rollup,version_rollup_batchesandanomaly_log. They rebuild on schema mismatch.Known limits