Skip to content

feat(agent-sessions): decide every aggregate/filter fact at ingest, the index projects the stamps - #1160

Merged
JeremyFunk merged 17 commits into
feat/ingest-usage-bucketsfrom
feat/ingest-ai-span-stamps
Sep 29, 2026
Merged

JeremyFunk merged 17 commits into
feat/ingest-usage-bucketsfrom
feat/ingest-ai-span-stamps

Conversation

@JeremyFunk

@JeremyFunk JeremyFunk commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1143

The base is feat/ingest-usage-buckets (#1143), so the diff shows this change only. CI for a stacked PR runs once #1143 merges and this PR's base is retargeted to main.

This is the pattern PR: integration rules live in the ingest gateway, the warehouse projects the gateway's maple_ai.* stamps, and the read path keeps only display-level decoding. If it is approved, #1140, #1141, #1142 and #1149 get reworked onto it (see the last section).

What

Ingest (apps/ingest/src/ai_session/facts.rs, new). For every span the gateway already stamps, it decides each fact the list aggregates or filters on and writes it next to #1143's maple_ai.llm_call and maple_ai.usage.*:

stamp value
maple_ai.tool_call 1 on a tool call: op execute_tool, or an op outside the convention with a tool name or a span name saying "tool"
maple_ai.error 1 when the span failed: status Error, error.type, or gen_ai.response.status failed/error
maple_ai.tool.paused 1 on a tool call's copy that recorded no outcome (no result, or, from Google ADK, its confirmation request)
maple_ai.model response model, request model, ai.response.model, ai.model.id, llm.model_name
maple_ai.agent.name gen_ai.agent.name, ai.telemetry.functionId; OpenAI Agents via OpenInference: graph.node.id on an AGENT span when it equals the span name
maple_ai.tool.name gen_ai.tool.name, ai.toolCall.name, tool.name
maple_ai.tool.call_id gen_ai.tool.call.id, ai.toolCall.id
maple_ai.response.id gen_ai.response.id, ai.response.id, on model calls
maple_ai.tool.description gen_ai.tool.description, tool.description, cut to 2000 chars, on tool calls
maple_ai.tool.error_result gen_ai.tool.call.result, ai.toolCall.result, cut to 1000 chars, on failed tool calls
  • Flags are written only when they hold and values only when present; maple_ai.llm_call (on every stamped span) is the marker that the rest were decided.
  • One pass over the span's attributes fills every fact, including feat(ingest): stamp model-call usage buckets and the llm-call marker #1143's usage, which used to scan the attributes once per key. Keys are looked up in a table bucketed by key length. No regex, no JSON parsing, no allocation except the stamps themselves.
  • Claude Code's folded tool failure (fold_tool_failures) also sets maple_ai.error and maple_ai.tool.error_result and drops maple_ai.tool.paused.
  • Gateway-owned: stripped from customer input with the rest of maple_ai.*, re-derived on re-ingest.

Warehouse (migration 0035, local schema v26). ai_trace_index_mv projects the stamps and nothing else, plus generic columns: the environment, error.type, the status message (cut to 400) and the failure fingerprint's redaction chain, which is the one error_events uses and holds no vendor knowledge. The DDL goes from ~25k to ~2.9k characters. No column changes, no backfill, requiredForIngest: false. gen-ai-columns.ts loses every key list, the op/name rules and the per-provider usage multiIf (-500 lines).

Read path

  • Detail page and MCP get_agent_session (@maple/agent-sessions): on a stamped span, classifyAiSpan, isLlmCall and spanFailed take the gateway's verdicts, spanTokenBuckets and the new spanCost its buckets and cost. These are the same facts the index sums, so list and detail cannot disagree for new rows. New catalog fields: mapleLlmCall, mapleToolCall, mapleError, maple*Tokens, mapleCost.
  • /summary SQL (oversized sessions, MCP): reads the verdicts, model, agent and buckets on stamped spans.
  • The span mapper prefers maple_ai.agent.name, so every detail-page reader of the agent name shows the agent the list shows (OpenAI Agents' graph node included); the prompt-cache check counts the gateway's cache buckets.
  • Spans ingested before the stamps keep today's rules until the 30-day TTL.

Inventory: integration rules on the DB layer before this PR

"Moved" means decided at ingest and stamped; the SQL now projects the stamp.

# Rule (where) Decides Vendors / dialects Also in TS Kind Now
1 MV WHERE maple_ai.vendor.id != '' which spans are agent spans all (gateway vendor detection) mapAiSpan isAiSpan filter already a stamp; unchanged
2 GENAI_MODEL_KEYS coalesce (MV) Model column, model facet/filter semconv, Vercel ai.*, OpenInference llm.* spanModel, aiFieldSourceKeys, /summary filter/facet moved: maple_ai.model
3 GENAI_AGENT_NAME_KEYS (MV) AgentName, agent facet, row heading semconv, Vercel functionId vendor agentName sources, /summary filter/facet moved: maple_ai.agent.name, plus the OpenAI Agents graph.node.id rule
4 GENAI_TOOL_NAME_KEYS (MV) ToolName, tool facet, tools pages semconv, Vercel, OpenInference vendor toolName sources filter/facet moved: maple_ai.tool.name
5 GENAI_RESPONSE_ID_KEYS (MV) ResponseId, dedupe of mirrored calls semconv, Vercel responseId sources aggregate moved: maple_ai.response.id
6 genAiOperationExpr (MV) op, OpenInference span kind to op semconv, OpenInference OPENINFERENCE_SPAN_KIND_OPERATIONS refine input to 7/8 moved (ingest Facts::operation)
7 genAiIsLlmCallCond: INFERENCE_OPS, KNOWN_OPS, nameLooks needles tool/agent/workflow/chat/completion, model presence (MV) IsLlmCall, call counts all; name fallback for unknown ops classifyAiSpan/isLlmCall; a third variant in /summary aggregate moved: #1143's maple_ai.llm_call (with its LiteLLM, Semantic Kernel, legacy Vercel, ADK call_llm and server-span rules)
8 genAiIsToolCallCond (MV) IsToolCall, tool counts, tools pages all classifyAiSpan tool branch; variant in /summary aggregate/filter moved: maple_ai.tool_call
9 genAiIsErrorCond (MV) IsError, error badges, failing filter semconv error.type, gen_ai.response.status failed/error dialect spanFailed (case-insensitive), /summary aggregate/filter moved: maple_ai.error
10 GENAI_USAGE_KEYS + byConvention multiIf by vendor (vercel_ai_sdk, maple) and provider (anthropic, openai, gcp.gemini/gemini, gcp.vertex_ai/vertex_ai, openrouter), greatest() guards (MV) Tokens and the five buckets every dialect; conventions keyed on gen_ai.provider.name spanTokenBuckets + genAiUsageConvention; raw sums in /summary aggregate moved: #1143's maple_ai.usage.*, projected
11 GENAI_COST_KEYS (MV) Cost OpenLLMetry, OpenRouter total_cost, OpenInference usageCost sources aggregate moved: maple_ai.usage.cost
12 GENAI_TOOL_DESCRIPTION_KEYS, cut 2000 (MV) ToolDescription (tool page header) semconv, OpenInference toolDescription sources detail carried by the index moved: maple_ai.tool.description
13 genAiFailedToolCallResultExpr, cut 1000 (MV) FailedToolCallResult, fingerprint input semconv, Vercel toolCallResult sources grouping moved: maple_ai.tool.error_result
14 error.type (MV ErrorType) failure type grouping plain OTel semconv errorType aggregate stays: generic OTel field
15 status message cut to 400, fingerprint = cityHash64 of the MSG_TEXT_REDACTIONS chain (MV) StatusMessage, ErrorFingerprint none (generic text) genAiErrorFingerprintText mirror grouping stays in SQL, now reads maple_ai.tool.error_result and maple_ai.error
16 DEPLOYMENT_ENV_SQL (MV) environment facet generic OTel resource attr, both semconv spellings none filter stays: generic
17 /summary summaryMeasures_ (ai-sessions.ts): aiFieldSourceKeys coalesce, op-only llm/tool rules, failure rule, raw usage sums, models, agents oversized-session and MCP totals every dialect classifyAiSpan, spanTokenBuckets aggregate stamped spans read the stamps; old path kept for old rows
18 list usage netting (ai-span-columns.ts: reporters, links, netted claims, response-id maxMap) wrapper roll-ups, provider attempts, mirrored calls structural, no vendor keys countableUsageSpans, countedLlmCalls aggregate stays: wrappers no longer carry usage, so the roll-up netting is a no-op for new rows; attempt and mirror handling still apply
19 /summary turn key: maple_ai.turn.id, conversation id sources, eve.turn.id turn grouping of the summary maple, eve, semconv buildSessionTurns detail structure stays on the read path (candidate for a later maple_ai.turn.id stamp)
20 tool error payloads (aiToolErrorPayloadsQuery): args/result via aiFieldSourceKeys modal payloads every dialect span detail detail stays on the read path
21 span read allowlist (aiSpanAttributeKeys) which keys the detail page receives every dialect mapAiSpan detail stays; now includes the new stamp keys

In-flight SQL rules this PR makes unnecessary: #1140's tool needle change to genAiIsToolCallCond, #1141's genAiToolCallIdExpr and genAiIsPausedToolCallCond, #1142's memory ops in KNOWN_OPS (the ingest KNOWN_OPS already has them, and the tool rule uses it), #1149's MV read of the usage buckets.

Stays on the read path by design: message and transcript decoding (ai-messages.ts, OpenInference input.value/output.value and flattened llm.*_messages.N), vendor refine hooks, turn segmentation, TTFT (gen_ai.response.time_to_first_chunk and the Vercel v7 key), provider name and its legacy spellings (display only now), tool arguments and results, span attribute labels.

Performance

In short: +21% CPU per AI span in the stamping pass (286 → 346 ns), zero for non-AI spans. GenAI spans are ~0.01% of production rows, so the average per ingested span rises by well under 1 ns.

Criterion, apps/ingest/benches/ai_session_bench.rs, M-series laptop, before = #1143 HEAD, after = this branch:

bench before after per span
ai_stamp_ai_batch/ai_1200_spans_10_vendors (new: 1200 agent spans, 10 vendors, wrapper/call/tool shapes) 343 µs 415 µs 286 → 346 ns (+21%)
ai_stamp_trace/trace_20_spans_2_ai 742 ns 1125 ns +190 ns per AI span (20-attribute Vercel wrapper)
ai_stamp/mixed_100_spans_95pct_non_ai 2.36 µs 3.20 µs 23.6 → 32.0 ns (5% AI spans)
ai_stamp_trace/trace_20_spans_non_ai 155 ns 140 ns non-AI path unchanged (noise)
  • Non-AI spans are untouched: they never reach the new code.
  • An AI span pays for its new stamps. The single facts pass is cheaper than feat(ingest): stamp model-call usage buckets and the llm-call marker #1143's per-key scans (with the stamps disabled the AI batch measured 344 µs, the same as before). Most of the +60 ns/span is allocating the stamps (two Strings each); the review fixes (text facts of any OTLP type, a $0 cost stamp) added about 5%.
  • The 5%-AI mix is pessimistic: GenAI spans are about 0.01% of production rows (ai_trace_index datasource doc), so the average cost per ingested span rises by well under 1 ns.
  • cargo bench --bench ai_session_bench -- ai_stamp --save-baseline before on feat(ingest): stamp model-call usage buckets and the llm-call marker #1143, then --baseline before here.

Deploy

  1. feat(ingest): stamp model-call usage buckets and the llm-call marker #1143 merges, then this PR (retarget to main first).
  2. Ingest first. Check fresh prod spans for maple_ai.model, maple_ai.tool_call, maple_ai.error (inspect_span). If the view goes first, spans ingested in between materialize as neither a call nor a tool, with no usage.
  3. Then the view: ClickHouse migration 0035 applies through the normal path (BYO-ClickHouse orgs too, requiredForIngest: false). Tinybird is not deployed by CI: tinybird:deploy per workspace (staging, prod EU, maple_us).
  4. The API/web read path can ship with either step: it falls back to the old rules where the marker is absent.

Cleanup after the 30-day TTL (30 days after step 3): delete reportedTokenBuckets, genAiUsageConvention and the convention tables, the unstamped branches of classifyAiSpan/isLlmCall/spanFailed/spanCost and of /summary, and the list's wrapper roll-up netting.

Migration numbering

This PR takes 0035 and local v26: migration versions must be contiguous, and a BYO-ClickHouse instance past a higher number would skip a lower one landing later. #1140, #1141, #1142 and #1149 rebase onto this PR and drop their own ai_trace_index_mv migrations and local schema bumps; a PR that still needs a schema change (#1141's two columns) takes the next number after this one and re-renders the view from the latest snapshot.

What the other PRs change to follow the pattern

How verified

  • cargo test --lib ai_session (apps/ingest): 71 pass, 11 new (incl. non-string values): facts per vendor (Vercel v7 and legacy, OpenAI Agents graph node, agno hash, ADK confirmation, failures by status/error.type/response status, unknown-op tools, memory ops, cuts by character), the Claude Code folded failure. feat(ingest): stamp model-call usage buckets and the llm-call marker #1143's 26 usage tests pass unchanged on the single pass. cargo clippy --all-targets: no warnings in the touched files.
  • Vitest, one file at a time: packages/domain migrations (40), gen-ai-columns, projection order, manifest, datasource contract, ingest schema compatibility; packages/agent-sessions session-turns (50), session-summary (77); packages/query-engine-integrations ai-sessions (100), ai-span-columns, ai-integrations, ai-vendors, ai-tools, SQL catalog baseline (updated); packages/query-engine catalog baseline.
  • apps/cli: local-store-migrations, store-version, inserts, store-marker-schema. clickhouse:schema:check and tinybird:manifest:check up to date.
  • Not run (no Docker on this machine): the ClickHouse e2e. ai-trace-index-materialization seeds now carry the stamps the gateway would write; its assertions are unchanged except that the turn span's roll-up now reads no usage. WarehouseQueryService e2e (migration replay) and the SQL catalog sweep run in CI.

… the span

The gateway now decides, per stamped span, whether it is a tool call,
whether it failed, whether it is a tool call's paused copy, and its model,
agent, tool, tool call id, response id, tool description and a failed tool
call's result, and writes each as a maple_ai.* stamp beside the llm-call
marker and usage buckets. One pass over the span's attributes feeds every
fact, usage included, in place of a scan per key.

OpenAI Agents' OpenInference instrumentor names an agent only in
graph.node.id on its AGENT span; that id is the agent name where it equals
the span name.
…ai.* stamps

Migration 0039 (local schema v26; both placeholders, renumbered at merge)
recreates ai_trace_index_mv so Model, AgentName, ToolName, ResponseId,
IsLlmCall, IsToolCall, IsError, the five token buckets, Tokens, Cost,
ToolDescription and FailedToolCallResult each read the fact the ingest
gateway stamped on the span. The operation lists, span-name needles,
dialect key lists and per-provider usage conventions leave the view; what
stays is generic: the environment, error.type, the status message and the
failure fingerprint's redaction chain.
… stamps

On a span the ingest gateway stamped (maple_ai.llm_call present),
classifyAiSpan, isLlmCall and spanFailed take its verdicts, spanTokenBuckets
and the new spanCost its buckets and cost, and the /summary SQL its
verdicts, names and buckets: the facts ai_trace_index sums, so the list and
the page cannot disagree. Spans ingested before keep the op/name rules and
usage conventions until the 30-day TTL. The materialization e2e seeds carry
the stamps the gateway would have written.
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: c2c2efd3-1bb1-42cf-af07-7744a928a1f0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@maple-review-bot

maple-review-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Note

A newer push replaced 7a763ea before its review finished. The latest commit is reviewed in a new comment.

Migration versions must be contiguous, and a BYO-ClickHouse instance at 0039
would skip a lower number landing later. The in-flight view changes rebase
onto this one and drop their own migrations of the view.
The span mapper prefers maple_ai.agent.name over the decoded dialect keys,
so every reader of genAi.agentName (header, turns, waterfall, filters)
shows the agent the list and its facets show, OpenAI Agents' graph node
included.
@maple-review-bot

maple-review-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Maple review

Confidence 3/5 · needs attention
quality 98/100 · 1 note · tests partial · risk medium

Warning

This review ended early; what follows is what it established.

Moves every AI-session fact the list and detail derive — call/tool/error kind, model, agent, usage buckets and cost — into ingest-gateway maple_ai.* stamps that the recreated ai_trace_index_mv projects. Internally consistent and tested; the one gap is that nothing pins the Rust stamp keys to the catalog. The generated local-schema-v26.sql was not read in full (it is the generated snapshot of the DDL I did verify byte-for-byte against 0035_ai_trace_index_gateway_stamps.ts and generated/clickhouse-schema.ts).

  • Ingest gateway decides the maple_ai.* facts in one attribute pass (facts.rs, usage.rs)
  • ai_trace_index_mv projects those stamps; migration 0035 drops and recreates it with no backfill
  • Read paths prefer the stamps: classifyAiSpan, isLlmCall, spanFailed, spanTokenBuckets, spanCost
  • Local store schema bumps to v26 with a new snapshot and history entry

Findings

Note · F1 · Nothing pins the Rust stamp keys to MAPLE_AI_STAMP_ATTRS

tests · packages/domain/src/gen-ai.ts:146-170

MAPLE_AI_STAMP_ATTRS and its hand-copied twins in apps/ingest/src/ai_session/facts.rs / usage.rs must hold the same literal strings (the comment on facts.rs:24 says so), but no test compares the two. This PR deleted ai-span-columns.test.ts, the only test that pinned the SQL key lists to the catalog, and the e2e seed builds its stamps from MAPLE_AI_STAMP_ATTRS too, so a rename on one side leaves every test green while IsLlmCall, Model and the token buckets go empty for every span ingested afterwards. Add a guard that reads the Rust sources and asserts each stamp key appears, or assert the literals against a shared checked-in list.

What was checked
  • Rust usage key tables (usage.rs:54-109) are supersets of every alias the integrations decode (ai-vendors.ts:51-115, ai-integrations.ts:168)
  • Migration 0035's DDL is byte-identical to the generated snapshot, and index.test.ts:934 pins that
  • /summary stamped-vs-fallback predicates match the MV columns (integrations.sql:820-833); claude_code phase spans stay unstamped and take the fallback
Files not reviewed (3)

The review ended before it read these diffs, so nothing above vouches for them.

  • apps/cli/src/server/schema/local-schema-v26.sql
  • apps/cli/test/local-store-migrations.test.ts
  • apps/cli/test/native-local-store-migration.sh
Copy all findings (1)
Findings from an automated review of commit 0410c4dac62600dc10a58e732e3b56e212603c18. Verify each one against the current code before changing anything, fix only those that still apply, and keep each fix to the lines it names.

---

F1 · Note · tests · packages/domain/src/gen-ai.ts:146-170
Nothing pins the Rust stamp keys to `MAPLE_AI_STAMP_ATTRS`
`MAPLE_AI_STAMP_ATTRS` and its hand-copied twins in `apps/ingest/src/ai_session/facts.rs` / `usage.rs` must hold the same literal strings (the comment on `facts.rs:24` says so), but no test compares the two. This PR deleted `ai-span-columns.test.ts`, the only test that pinned the SQL key lists to the catalog, and the e2e seed builds its stamps from `MAPLE_AI_STAMP_ATTRS` too, so a rename on one side leaves every test green while `IsLlmCall`, `Model` and the token buckets go empty for every span ingested afterwards. Add a guard that reads the Rust sources and asserts each stamp key appears, or assert the literals against a shared checked-in list.

0410c4d · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple to ask about one.

@maple-review-bot

maple-review-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Maple review

Confidence 4/5 · likely safe to merge
The head only refactors the e2e seed helper; the remaining risk is the unstamped-span fallback in ai-sessions.ts, which no seed exercises.
quality 98/100 · 1 note · tests covered · risk medium · 1/1 new units observable

Since the last review the pull request only rewrites the e2e test's gateway seed helper to build the stamp map from one entries list, which yields the same attributes. No defect in that change; the open finding about the Rust stamp keys being unpinned still stands.

  • gateway() builds its stamp map from one entries list instead of mutating it

Still open from earlier reviews

What was checked
  • gateway() at 280f997 emits the same keys and values as at 0410c4d (git diff)
  • Every stamped seed carries gateway(...); AGENT_CHILD_SPAN and PLAIN_SPAN stay unstamped
  • indexRow expectations match the 0035 DDL projection column for column
Observability coverage: 1 of 1 changes observable
Change Kind Observable Evidence
ingest span stamping (apps/ingest/src/ai_session/facts.rs) in-process attribute computation on the existing ingest path yes no new inbound entrypoint, outbound call or background work; the pass runs inside the ingest request

280f997 · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple to ask about one.

…holds it

An integer tool call id or response id, a structured tool call result and a
non-string error.type counted when the view read the Map; the stamps now
count them too, stringified the way the row encoder writes the Map. Only
Google ADK writes the confirmation request, so only its tool results are
searched for it.
@maple-review-bot

maple-review-bot Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Maple review

Confidence 4/5 · likely safe to merge
A deliberate behavior change: the span-name/model fallback for a model call now applies only to unknown: vendors, so known-vendor spans with an unrecognized op stop counting as LLM calls.
quality 100/100 · no findings · tests covered · risk medium

The ingest gateway now decides every Agent Sessions fact once per span and stamps it as maple_ai.*; ai_trace_index_mv becomes a projection of those stamps and the detail page reads the same verdicts. Contained and well tested; safe to merge once the ingest deploy ships first.

  • facts::stamps writes the gateway's verdicts, names and usage buckets on every stamped AI span
  • Migration 0035 recreates ai_trace_index_mv as a projection of maple_ai.* plus error.type and the fingerprint
  • classifyAiSpan/spanFailed/spanTokenBuckets take the gateway's verdicts for stamped spans
  • gen-ai.test.ts pins MAPLE_AI_STAMP_ATTRS against the Rust key literals

Fixed since the last review

  • F1 · Nothing pins the Rust stamp keys to MAPLE_AI_STAMP_ATTRS
What was checked
  • F1 resolved: every MAPLE_AI_STAMP_ATTRS key appears as a &str = "maple_ai..." literal in facts.rs/usage.rs, and turbo.json adds those sources as test inputs
  • The gateway's error, tool_call and model/agent/tool key lists reproduce the old view expressions, so IsError/IsToolCall/Model do not shift for stamped spans
  • No new entrypoint, outbound call or background loop: the change is a transform inside the existing ingest path plus a view recreate, so no span or metric is owed

38f256c · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple to ask about one.

@JeremyFunk
JeremyFunk merged commit 42e6357 into feat/ingest-usage-buckets Sep 29, 2026
36 checks passed
@JeremyFunk
JeremyFunk deleted the feat/ingest-ai-span-stamps branch September 29, 2026 21:00
JeremyFunk added a commit that referenced this pull request Sep 30, 2026
…1143)

* feat(ingest): stamp model-call usage as disjoint maple_ai.usage buckets

The gateway now restates a model call's token usage as five disjoint
buckets (uncached input, cache read, cache write, visible output,
reasoning) plus cost under maple_ai.usage.*, leaving the customer's
gen_ai.usage.* as sent. The convention is chosen by emitter, not by
gen_ai.provider.name, so inclusive emitters labelled anthropic
(OpenRouter Broadcast, OTel genai anthropic, Pydantic AI) no longer
double count their cache. Only the span that is the model call carries
buckets; agent, step and workflow wrappers get none.

Guards: a prompt smaller than its cache is read cache-exclusive, and
reasoning is clamped to the completion so a span's total matches the
provider's total_tokens.

* feat(ingest): mark every stamped span with maple_ai.llm_call

The gateway already decides which span is the model call to own the
usage buckets; it now says so on every stamped span: 1 on the model call
(with or without usage, so a failed call still counts), 0 on the rest.
The key's presence tells readers the gateway classified the span, so
rows without it keep the op/name heuristics until they age out.

An unknown dialect's server span (a proxy's POST /chat/completions) is
never the call. Tests cover the heuristic false positives: Spring AI
chat_client (op framework), LangSmith ChatPromptTemplate (op chain),
DSPy ChatAdapter.__call__ (no op) and the LiteLLM proxy server span.

* fix(ingest): treat the GenAI memory operations as known ops in the model-call rule

A memory operation naming its embedding model is agent bookkeeping, not
a model call; #1142 adds the same ops to the index's known-op list.

* feat(agent-sessions): decide every aggregate/filter fact at ingest, the index projects the stamps (#1160)

* feat(ingest): stamp every Agent Sessions aggregate and filter fact on the span

The gateway now decides, per stamped span, whether it is a tool call,
whether it failed, whether it is a tool call's paused copy, and its model,
agent, tool, tool call id, response id, tool description and a failed tool
call's result, and writes each as a maple_ai.* stamp beside the llm-call
marker and usage buckets. One pass over the span's attributes feeds every
fact, usage included, in place of a scan per key.

OpenAI Agents' OpenInference instrumentor names an agent only in
graph.node.id on its AGENT span; that id is the agent name where it equals
the span name.

* feat(agent-sessions): ai_trace_index_mv projects the gateway's maple_ai.* stamps

Migration 0039 (local schema v26; both placeholders, renumbered at merge)
recreates ai_trace_index_mv so Model, AgentName, ToolName, ResponseId,
IsLlmCall, IsToolCall, IsError, the five token buckets, Tokens, Cost,
ToolDescription and FailedToolCallResult each read the fact the ingest
gateway stamped on the span. The operation lists, span-name needles,
dialect key lists and per-provider usage conventions leave the view; what
stays is generic: the environment, error.type, the status message and the
failure fingerprint's redaction chain.

* feat(agent-sessions): the detail page and /summary read the gateway's stamps

On a span the ingest gateway stamped (maple_ai.llm_call present),
classifyAiSpan, isLlmCall and spanFailed take its verdicts, spanTokenBuckets
and the new spanCost its buckets and cost, and the /summary SQL its
verdicts, names and buckets: the facts ai_trace_index sums, so the list and
the page cannot disagree. Spans ingested before keep the op/name rules and
usage conventions until the 30-day TTL. The materialization e2e seeds carry
the stamps the gateway would have written.

* chore(ingest): name the stamp key table's entry type, point the stamps at MAPLE_AI_STAMP_ATTRS

* docs(agent-sessions): point comments at the gateway's stamps instead of the deleted SQL builders

* fix(agent-sessions): number the gateway-stamps view migration 0035

Migration versions must be contiguous, and a BYO-ClickHouse instance at 0039
would skip a lower number landing later. The in-flight view changes rebase
onto this one and drop their own migrations of the view.

* fix(agent-sessions): the prompt-cache check counts the gateway's cache buckets

* fix(agent-sessions): the detail page names the agent the gateway named

The span mapper prefers maple_ai.agent.name over the decoded dialect keys,
so every reader of genAi.agentName (header, turns, waterfall, filters)
shows the agent the list and its facets show, OpenAI Agents' graph node
included.

* chore(agent-sessions): drop unread stamp keys, unexport in-file constants, fix stale comments

* fix(agent-sessions): build the e2e gateway stamps without an open dictionary binding

* test(agent-sessions): pin the stamp keys against the ingest gateway's sources

* fix(ingest): read a text fact of any OTLP type, as the warehouse Map holds it

An integer tool call id or response id, a structured tool call result and a
non-string error.type counted when the view read the Map; the stamps now
count them too, stringified the way the row encoder writes the Map. Only
Google ADK writes the confirmation request, so only its tool results are
searched for it.

* fix(ingest): stamp a reported $0 cost, so a free call is not read as unpriced

* fix(ingest): saturate the cache sum, so an absurd customer figure cannot overflow it

* test(agent-sessions): hash the gateway's Rust sources into the stamp-key pin, match only their constants

* test(agent-sessions): prove the gateway's failure verdict overrides a span's own error.type

* chore(agent-sessions): cut process notes and a duplicate test, rewrap comments

* feat(agent-sessions): project the tool call id and the paused-copy stamp onto ai_trace_index

Migration 0035 and local schema v26 add ToolCallId (maple_ai.tool.call_id)
and IsPausedToolCall (maple_ai.tool.paused) to ai_trace_index, so the list
can count a call paused for approval and executed in a later trace once.
Only the columns and their projection; the counting stays with the list.

* test(agent-sessions): seed the ai-tools e2e spans with the gateway's stamps

ai_trace_index_mv projects only the maple_ai.* stamps since migration 0035,
so the ai-tools seeds, which carried none, would materialize as neither a
call nor a tool. The stamp helper moves to clickhouse-e2e-support so both
suites share it.

* feat(agent-sessions): a tool call paused for approval is no call, by its framework's explicit mark

Drops ToolCallId and IsPausedToolCall from migration 0035 and local schema
v26, and the gateway's maple_ai.tool.call_id and maple_ai.tool.paused
stamps. Instead the gateway stamps maple_ai.tool_call = 0 on the copy a
call paused for a human's approval leaves, so list, summary and detail all
count maple_ai.tool_call = 1; a call paused and then rejected counts none.
Such a copy is no failure either, though some frameworks end it in error.

A pause is read only from a framework's explicit mark, never from a
missing result (content capture off, or a tool that returns nothing):

- Google ADK: gcp.vertex.agent.tool_response carries the confirmation
  request (every ADK version writes the key; 2.6 writes no
  gen_ai.tool.call.result).
- OpenAI Agents SDK up to 0.22.0: output.value is a ToolApprovalItem's repr.
- pydantic-ai: pydantic_ai.tool.deferral.name = ApprovalRequired.
- LlamaIndex's own tracer: the step's status message "Waiting for event".

Strands, OpenAI Agents 0.22.1+ and a Mastra tool that suspends itself mark
nothing on the paused copy, so it still counts.

* fix(agent-sessions): the detail page counts and fails a stamped span by the gateway's verdict (folds #1141)

The gateway stamps maple_ai.tool_call = 0 on the copy a call paused for a
human's approval leaves, and no maple_ai.error on it even where its
framework ends it in error. The read path follows:

- isCountedToolCall: on a stamped span maple_ai.tool_call = 1. The detail
  count, the tool histogram, the checks' coverage and the repetition
  finding use it, so a paused copy is in none of them. The display kind
  (classifyAiSpan) still reads a paused copy as a tool: the waterfall, span
  kind and icons show it as one.
- countedToolCalls: the result-based pause merge (drop a no-result copy
  when a later copy under the same call id carries a result) serves only
  pre-stamp spans until the 30-day TTL, so a stamped call that recorded no
  result is no longer dropped.
- spanFailed and the totals read's failure count: on a stamped span the
  gateway's maple_ai.error alone, so a paused copy its framework ended in
  error (LlamaIndex, pydantic-ai before instrumentation v5) is no failure,
  as on the list.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant