feat(agent-sessions): decide every aggregate/filter fact at ingest, the index projects the stamps - #1160
Conversation
… the span The gateway now decides, per stamped span, whether it is a tool call, whether it failed, whether it is a tool call's paused copy, and its model, agent, tool, tool call id, response id, tool description and a failed tool call's result, and writes each as a maple_ai.* stamp beside the llm-call marker and usage buckets. One pass over the span's attributes feeds every fact, usage included, in place of a scan per key. OpenAI Agents' OpenInference instrumentor names an agent only in graph.node.id on its AGENT span; that id is the agent name where it equals the span name.
…ai.* stamps Migration 0039 (local schema v26; both placeholders, renumbered at merge) recreates ai_trace_index_mv so Model, AgentName, ToolName, ResponseId, IsLlmCall, IsToolCall, IsError, the five token buckets, Tokens, Cost, ToolDescription and FailedToolCallResult each read the fact the ingest gateway stamped on the span. The operation lists, span-name needles, dialect key lists and per-provider usage conventions leave the view; what stays is generic: the environment, error.type, the status message and the failure fingerprint's redaction chain.
… stamps On a span the ingest gateway stamped (maple_ai.llm_call present), classifyAiSpan, isLlmCall and spanFailed take its verdicts, spanTokenBuckets and the new spanCost its buckets and cost, and the /summary SQL its verdicts, names and buckets: the facts ai_trace_index sums, so the list and the page cannot disagree. Spans ingested before keep the op/name rules and usage conventions until the 30-day TTL. The materialization e2e seeds carry the stamps the gateway would have written.
…s at MAPLE_AI_STAMP_ATTRS
…of the deleted SQL builders
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Note A newer push replaced |
Migration versions must be contiguous, and a BYO-ClickHouse instance at 0039 would skip a lower number landing later. The in-flight view changes rebase onto this one and drop their own migrations of the view.
The span mapper prefers maple_ai.agent.name over the decoded dialect keys, so every reader of genAi.agentName (header, turns, waterfall, filters) shows the agent the list and its facets show, OpenAI Agents' graph node included.
…ants, fix stale comments
Maple reviewConfidence 3/5 · needs attention Warning This review ended early; what follows is what it established. Moves every AI-session fact the list and detail derive — call/tool/error kind, model, agent, usage buckets and cost — into ingest-gateway
FindingsNote · F1 · Nothing pins the Rust stamp keys to
|
Maple reviewConfidence 4/5 · likely safe to merge Since the last review the pull request only rewrites the e2e test's
Still open from earlier reviews
What was checkedObservability coverage: 1 of 1 changes observable
|
…holds it An integer tool call id or response id, a structured tool call result and a non-string error.type counted when the view read the Map; the stamps now count them too, stringified the way the row encoder writes the Map. Only Google ADK writes the confirmation request, so only its tool results are searched for it.
…key pin, match only their constants
… span's own error.type
Maple reviewConfidence 4/5 · likely safe to merge The ingest gateway now decides every Agent Sessions fact once per span and stamps it as
Fixed since the last review
What was checked
|
…1143) * feat(ingest): stamp model-call usage as disjoint maple_ai.usage buckets The gateway now restates a model call's token usage as five disjoint buckets (uncached input, cache read, cache write, visible output, reasoning) plus cost under maple_ai.usage.*, leaving the customer's gen_ai.usage.* as sent. The convention is chosen by emitter, not by gen_ai.provider.name, so inclusive emitters labelled anthropic (OpenRouter Broadcast, OTel genai anthropic, Pydantic AI) no longer double count their cache. Only the span that is the model call carries buckets; agent, step and workflow wrappers get none. Guards: a prompt smaller than its cache is read cache-exclusive, and reasoning is clamped to the completion so a span's total matches the provider's total_tokens. * feat(ingest): mark every stamped span with maple_ai.llm_call The gateway already decides which span is the model call to own the usage buckets; it now says so on every stamped span: 1 on the model call (with or without usage, so a failed call still counts), 0 on the rest. The key's presence tells readers the gateway classified the span, so rows without it keep the op/name heuristics until they age out. An unknown dialect's server span (a proxy's POST /chat/completions) is never the call. Tests cover the heuristic false positives: Spring AI chat_client (op framework), LangSmith ChatPromptTemplate (op chain), DSPy ChatAdapter.__call__ (no op) and the LiteLLM proxy server span. * fix(ingest): treat the GenAI memory operations as known ops in the model-call rule A memory operation naming its embedding model is agent bookkeeping, not a model call; #1142 adds the same ops to the index's known-op list. * feat(agent-sessions): decide every aggregate/filter fact at ingest, the index projects the stamps (#1160) * feat(ingest): stamp every Agent Sessions aggregate and filter fact on the span The gateway now decides, per stamped span, whether it is a tool call, whether it failed, whether it is a tool call's paused copy, and its model, agent, tool, tool call id, response id, tool description and a failed tool call's result, and writes each as a maple_ai.* stamp beside the llm-call marker and usage buckets. One pass over the span's attributes feeds every fact, usage included, in place of a scan per key. OpenAI Agents' OpenInference instrumentor names an agent only in graph.node.id on its AGENT span; that id is the agent name where it equals the span name. * feat(agent-sessions): ai_trace_index_mv projects the gateway's maple_ai.* stamps Migration 0039 (local schema v26; both placeholders, renumbered at merge) recreates ai_trace_index_mv so Model, AgentName, ToolName, ResponseId, IsLlmCall, IsToolCall, IsError, the five token buckets, Tokens, Cost, ToolDescription and FailedToolCallResult each read the fact the ingest gateway stamped on the span. The operation lists, span-name needles, dialect key lists and per-provider usage conventions leave the view; what stays is generic: the environment, error.type, the status message and the failure fingerprint's redaction chain. * feat(agent-sessions): the detail page and /summary read the gateway's stamps On a span the ingest gateway stamped (maple_ai.llm_call present), classifyAiSpan, isLlmCall and spanFailed take its verdicts, spanTokenBuckets and the new spanCost its buckets and cost, and the /summary SQL its verdicts, names and buckets: the facts ai_trace_index sums, so the list and the page cannot disagree. Spans ingested before keep the op/name rules and usage conventions until the 30-day TTL. The materialization e2e seeds carry the stamps the gateway would have written. * chore(ingest): name the stamp key table's entry type, point the stamps at MAPLE_AI_STAMP_ATTRS * docs(agent-sessions): point comments at the gateway's stamps instead of the deleted SQL builders * fix(agent-sessions): number the gateway-stamps view migration 0035 Migration versions must be contiguous, and a BYO-ClickHouse instance at 0039 would skip a lower number landing later. The in-flight view changes rebase onto this one and drop their own migrations of the view. * fix(agent-sessions): the prompt-cache check counts the gateway's cache buckets * fix(agent-sessions): the detail page names the agent the gateway named The span mapper prefers maple_ai.agent.name over the decoded dialect keys, so every reader of genAi.agentName (header, turns, waterfall, filters) shows the agent the list and its facets show, OpenAI Agents' graph node included. * chore(agent-sessions): drop unread stamp keys, unexport in-file constants, fix stale comments * fix(agent-sessions): build the e2e gateway stamps without an open dictionary binding * test(agent-sessions): pin the stamp keys against the ingest gateway's sources * fix(ingest): read a text fact of any OTLP type, as the warehouse Map holds it An integer tool call id or response id, a structured tool call result and a non-string error.type counted when the view read the Map; the stamps now count them too, stringified the way the row encoder writes the Map. Only Google ADK writes the confirmation request, so only its tool results are searched for it. * fix(ingest): stamp a reported $0 cost, so a free call is not read as unpriced * fix(ingest): saturate the cache sum, so an absurd customer figure cannot overflow it * test(agent-sessions): hash the gateway's Rust sources into the stamp-key pin, match only their constants * test(agent-sessions): prove the gateway's failure verdict overrides a span's own error.type * chore(agent-sessions): cut process notes and a duplicate test, rewrap comments * feat(agent-sessions): project the tool call id and the paused-copy stamp onto ai_trace_index Migration 0035 and local schema v26 add ToolCallId (maple_ai.tool.call_id) and IsPausedToolCall (maple_ai.tool.paused) to ai_trace_index, so the list can count a call paused for approval and executed in a later trace once. Only the columns and their projection; the counting stays with the list. * test(agent-sessions): seed the ai-tools e2e spans with the gateway's stamps ai_trace_index_mv projects only the maple_ai.* stamps since migration 0035, so the ai-tools seeds, which carried none, would materialize as neither a call nor a tool. The stamp helper moves to clickhouse-e2e-support so both suites share it. * feat(agent-sessions): a tool call paused for approval is no call, by its framework's explicit mark Drops ToolCallId and IsPausedToolCall from migration 0035 and local schema v26, and the gateway's maple_ai.tool.call_id and maple_ai.tool.paused stamps. Instead the gateway stamps maple_ai.tool_call = 0 on the copy a call paused for a human's approval leaves, so list, summary and detail all count maple_ai.tool_call = 1; a call paused and then rejected counts none. Such a copy is no failure either, though some frameworks end it in error. A pause is read only from a framework's explicit mark, never from a missing result (content capture off, or a tool that returns nothing): - Google ADK: gcp.vertex.agent.tool_response carries the confirmation request (every ADK version writes the key; 2.6 writes no gen_ai.tool.call.result). - OpenAI Agents SDK up to 0.22.0: output.value is a ToolApprovalItem's repr. - pydantic-ai: pydantic_ai.tool.deferral.name = ApprovalRequired. - LlamaIndex's own tracer: the step's status message "Waiting for event". Strands, OpenAI Agents 0.22.1+ and a Mastra tool that suspends itself mark nothing on the paused copy, so it still counts. * fix(agent-sessions): the detail page counts and fails a stamped span by the gateway's verdict (folds #1141) The gateway stamps maple_ai.tool_call = 0 on the copy a call paused for a human's approval leaves, and no maple_ai.error on it even where its framework ends it in error. The read path follows: - isCountedToolCall: on a stamped span maple_ai.tool_call = 1. The detail count, the tool histogram, the checks' coverage and the repetition finding use it, so a paused copy is in none of them. The display kind (classifyAiSpan) still reads a paused copy as a tool: the waterfall, span kind and icons show it as one. - countedToolCalls: the result-based pause merge (drop a no-result copy when a later copy under the same call id carries a result) serves only pre-stamp spans until the 30-day TTL, so a stamped call that recorded no result is no longer dropped. - spanFailed and the totals read's failure count: on a stamped span the gateway's maple_ai.error alone, so a paused copy its framework ended in error (LlamaIndex, pydantic-ai before instrumentation v5) is no failure, as on the list.
Stacked on #1143
The base is
feat/ingest-usage-buckets(#1143), so the diff shows this change only. CI for a stacked PR runs once #1143 merges and this PR's base is retargeted tomain.This is the pattern PR: integration rules live in the ingest gateway, the warehouse projects the gateway's
maple_ai.*stamps, and the read path keeps only display-level decoding. If it is approved, #1140, #1141, #1142 and #1149 get reworked onto it (see the last section).What
Ingest (
apps/ingest/src/ai_session/facts.rs, new). For every span the gateway already stamps, it decides each fact the list aggregates or filters on and writes it next to #1143'smaple_ai.llm_callandmaple_ai.usage.*:maple_ai.tool_call1on a tool call: opexecute_tool, or an op outside the convention with a tool name or a span name saying "tool"maple_ai.error1when the span failed: statusError,error.type, orgen_ai.response.statusfailed/errormaple_ai.tool.paused1on a tool call's copy that recorded no outcome (no result, or, from Google ADK, its confirmation request)maple_ai.modelai.response.model,ai.model.id,llm.model_namemaple_ai.agent.namegen_ai.agent.name,ai.telemetry.functionId; OpenAI Agents via OpenInference:graph.node.idon an AGENT span when it equals the span namemaple_ai.tool.namegen_ai.tool.name,ai.toolCall.name,tool.namemaple_ai.tool.call_idgen_ai.tool.call.id,ai.toolCall.idmaple_ai.response.idgen_ai.response.id,ai.response.id, on model callsmaple_ai.tool.descriptiongen_ai.tool.description,tool.description, cut to 2000 chars, on tool callsmaple_ai.tool.error_resultgen_ai.tool.call.result,ai.toolCall.result, cut to 1000 chars, on failed tool callsmaple_ai.llm_call(on every stamped span) is the marker that the rest were decided.fold_tool_failures) also setsmaple_ai.errorandmaple_ai.tool.error_resultand dropsmaple_ai.tool.paused.maple_ai.*, re-derived on re-ingest.Warehouse (migration 0035, local schema v26).
ai_trace_index_mvprojects the stamps and nothing else, plus generic columns: the environment,error.type, the status message (cut to 400) and the failure fingerprint's redaction chain, which is the oneerror_eventsuses and holds no vendor knowledge. The DDL goes from ~25k to ~2.9k characters. No column changes, no backfill,requiredForIngest: false.gen-ai-columns.tsloses every key list, the op/name rules and the per-provider usagemultiIf(-500 lines).Read path
get_agent_session(@maple/agent-sessions): on a stamped span,classifyAiSpan,isLlmCallandspanFailedtake the gateway's verdicts,spanTokenBucketsand the newspanCostits buckets and cost. These are the same facts the index sums, so list and detail cannot disagree for new rows. New catalog fields:mapleLlmCall,mapleToolCall,mapleError,maple*Tokens,mapleCost./summarySQL (oversized sessions, MCP): reads the verdicts, model, agent and buckets on stamped spans.maple_ai.agent.name, so every detail-page reader of the agent name shows the agent the list shows (OpenAI Agents' graph node included); the prompt-cache check counts the gateway's cache buckets.Inventory: integration rules on the DB layer before this PR
"Moved" means decided at ingest and stamped; the SQL now projects the stamp.
WHERE maple_ai.vendor.id != ''mapAiSpanisAiSpanGENAI_MODEL_KEYScoalesce (MV)Modelcolumn, model facet/filterai.*, OpenInferencellm.*spanModel,aiFieldSourceKeys,/summarymaple_ai.modelGENAI_AGENT_NAME_KEYS(MV)AgentName, agent facet, row headingfunctionIdagentNamesources,/summarymaple_ai.agent.name, plus the OpenAI Agentsgraph.node.idruleGENAI_TOOL_NAME_KEYS(MV)ToolName, tool facet, tools pagestoolNamesourcesmaple_ai.tool.nameGENAI_RESPONSE_ID_KEYS(MV)ResponseId, dedupe of mirrored callsresponseIdsourcesmaple_ai.response.idgenAiOperationExpr(MV)OPENINFERENCE_SPAN_KIND_OPERATIONSrefineFacts::operation)genAiIsLlmCallCond:INFERENCE_OPS,KNOWN_OPS,nameLooksneedlestool/agent/workflow/chat/completion, model presence (MV)IsLlmCall, call countsclassifyAiSpan/isLlmCall; a third variant in/summarymaple_ai.llm_call(with its LiteLLM, Semantic Kernel, legacy Vercel, ADKcall_llmand server-span rules)genAiIsToolCallCond(MV)IsToolCall, tool counts, tools pagesclassifyAiSpantool branch; variant in/summarymaple_ai.tool_callgenAiIsErrorCond(MV)IsError, error badges, failing filtererror.type,gen_ai.response.statusfailed/error dialectspanFailed(case-insensitive),/summarymaple_ai.errorGENAI_USAGE_KEYS+byConventionmultiIf by vendor (vercel_ai_sdk,maple) and provider (anthropic,openai,gcp.gemini/gemini,gcp.vertex_ai/vertex_ai,openrouter),greatest()guards (MV)Tokensand the five bucketsgen_ai.provider.namespanTokenBuckets+genAiUsageConvention; raw sums in/summarymaple_ai.usage.*, projectedGENAI_COST_KEYS(MV)Costtotal_cost, OpenInferenceusageCostsourcesmaple_ai.usage.costGENAI_TOOL_DESCRIPTION_KEYS, cut 2000 (MV)ToolDescription(tool page header)toolDescriptionsourcesmaple_ai.tool.descriptiongenAiFailedToolCallResultExpr, cut 1000 (MV)FailedToolCallResult, fingerprint inputtoolCallResultsourcesmaple_ai.tool.error_resulterror.type(MVErrorType)errorTypecityHash64of theMSG_TEXT_REDACTIONSchain (MV)StatusMessage,ErrorFingerprintgenAiErrorFingerprintTextmirrormaple_ai.tool.error_resultandmaple_ai.errorDEPLOYMENT_ENV_SQL(MV)/summarysummaryMeasures_(ai-sessions.ts):aiFieldSourceKeyscoalesce, op-only llm/tool rules, failure rule, raw usage sums, models, agentsclassifyAiSpan,spanTokenBucketsai-span-columns.ts: reporters, links, netted claims, response-idmaxMap)countableUsageSpans,countedLlmCalls/summaryturn key:maple_ai.turn.id, conversation id sources,eve.turn.idbuildSessionTurnsmaple_ai.turn.idstamp)aiToolErrorPayloadsQuery): args/result viaaiFieldSourceKeysaiSpanAttributeKeys)mapAiSpanIn-flight SQL rules this PR makes unnecessary: #1140's
toolneedle change togenAiIsToolCallCond, #1141'sgenAiToolCallIdExprandgenAiIsPausedToolCallCond, #1142's memory ops inKNOWN_OPS(the ingestKNOWN_OPSalready has them, and the tool rule uses it), #1149's MV read of the usage buckets.Stays on the read path by design: message and transcript decoding (
ai-messages.ts, OpenInferenceinput.value/output.valueand flattenedllm.*_messages.N), vendor refine hooks, turn segmentation, TTFT (gen_ai.response.time_to_first_chunkand the Vercel v7 key), provider name and its legacy spellings (display only now), tool arguments and results, span attribute labels.Performance
In short: +21% CPU per AI span in the stamping pass (286 → 346 ns), zero for non-AI spans. GenAI spans are ~0.01% of production rows, so the average per ingested span rises by well under 1 ns.
Criterion,
apps/ingest/benches/ai_session_bench.rs, M-series laptop,before= #1143 HEAD,after= this branch:ai_stamp_ai_batch/ai_1200_spans_10_vendors(new: 1200 agent spans, 10 vendors, wrapper/call/tool shapes)ai_stamp_trace/trace_20_spans_2_aiai_stamp/mixed_100_spans_95pct_non_aiai_stamp_trace/trace_20_spans_non_aiStrings each); the review fixes (text facts of any OTLP type, a $0 cost stamp) added about 5%.ai_trace_indexdatasource doc), so the average cost per ingested span rises by well under 1 ns.cargo bench --bench ai_session_bench -- ai_stamp --save-baseline beforeon feat(ingest): stamp model-call usage buckets and the llm-call marker #1143, then--baseline beforehere.Deploy
mainfirst).maple_ai.model,maple_ai.tool_call,maple_ai.error(inspect_span). If the view goes first, spans ingested in between materialize as neither a call nor a tool, with no usage.requiredForIngest: false). Tinybird is not deployed by CI:tinybird:deployper workspace (staging, prod EU,maple_us).Cleanup after the 30-day TTL (30 days after step 3): delete
reportedTokenBuckets,genAiUsageConventionand the convention tables, the unstamped branches ofclassifyAiSpan/isLlmCall/spanFailed/spanCostand of/summary, and the list's wrapper roll-up netting.Migration numbering
This PR takes 0035 and local v26: migration versions must be contiguous, and a BYO-ClickHouse instance past a higher number would skip a lower one landing later. #1140, #1141, #1142 and #1149 rebase onto this PR and drop their own
ai_trace_index_mvmigrations and local schema bumps; a PR that still needs a schema change (#1141's two columns) takes the next number after this one and re-renders the view from the latest snapshot.What the other PRs change to follow the pattern
toolneedle change goes intofacts::is_tool_callin Rust (plus a TS fallback for old rows if wanted); drop itsgen-ai-columns.tsedit, its view migration and its local schema bump. The Mastra scorer ingest commit and the sessionless-traceHAVING(reads index columns only) stay as they are.genAiToolCallIdExpr/genAiIsPausedToolCallCond; addToolCallId/IsPausedToolCallas projections ofmaple_ai.tool.call_idandmaple_ai.tool.paused = '1'; the detail page'scountedToolCallsreads the same stamps (two catalog fields) on stamped spans.KNOWN_OPShas the memory ops). What is left is the TS fallback for old rows and the Claude Codecache_writekey rename.How verified
cargo test --lib ai_session(apps/ingest): 71 pass, 11 new (incl. non-string values): facts per vendor (Vercel v7 and legacy, OpenAI Agents graph node, agno hash, ADK confirmation, failures by status/error.type/response status, unknown-op tools, memory ops, cuts by character), the Claude Code folded failure. feat(ingest): stamp model-call usage buckets and the llm-call marker #1143's 26 usage tests pass unchanged on the single pass.cargo clippy --all-targets: no warnings in the touched files.packages/domainmigrations (40),gen-ai-columns, projection order, manifest, datasource contract, ingest schema compatibility;packages/agent-sessionssession-turns(50),session-summary(77);packages/query-engine-integrationsai-sessions(100),ai-span-columns,ai-integrations,ai-vendors,ai-tools, SQL catalog baseline (updated);packages/query-enginecatalog baseline.apps/cli:local-store-migrations,store-version,inserts,store-marker-schema.clickhouse:schema:checkandtinybird:manifest:checkup to date.ai-trace-index-materializationseeds now carry the stamps the gateway would write; its assertions are unchanged except that the turn span's roll-up now reads no usage.WarehouseQueryServicee2e (migration replay) and the SQL catalog sweep run in CI.