Repository navigation
feat(ai): migrate support responses to Luna and OpenAI Agents SDK - #283
jerelvelarde wants to merge 87 commits into
Conversation
R1-DELIVERY-A1: join only the formatted summary and optional details when saving suggestedResponse, so no-adapter web sources retain technical details and source links. Keep private response/handoff metadata outside this publishable field. Responses without details keep their existing text. Affected symbol and callsites: - handleAiResponse: registered by apps/worker/src/index.ts:85, imported at line 27; re-exported by packages/outpost/queue/src/index.ts:4. - Test-only calls are in packages/outpost/queue/src/__tests__/ai-response.test.ts, imported at line 95; the new source-matrix call is at line 737. - The single suggestedResponse assignment is in handleAiResponse; audit of the full handler and queue module found no other durable suggestion writer. The separate shadow SYSTEM diagnostic uses the formatted summary, and adapter delivery already receives the full formatted response object. Validation: - Baseline: 98 handler tests pass. - Red: all five new no-adapter source cases fail (actual Summary only). - Green: 103 handler tests; 74 package files / 1457 tests pass. - Targeted Prettier and ESLint pass; queue tsc --noEmit and emitted build pass. - git diff --check passes. Review convergence remains owned by the parent task.
Reject non-positive, non-finite, fractional, malformed, and unsafe integer PATHFINDER_MAX_QUERY_CHARS values before direct Pathfinder clients can run. Keep the default at 1000 and preserve positive safe-integer overrides. AI-A01 regression: 10 new rejection cases failed before the fix because config import succeeded (including maxQueryChars: NaN); all now pass. Validation: 642 AI tests passed, including 70 config/Pathfinder tests; targeted Prettier, ESLint, and AI TypeScript noEmit passed. Affected-symbol audit: config.pathfinder.maxQueryChars is consumed only by PathfinderClient.searchEvidence, private search, and searchDocs via capQuery. All three now receive eagerly validated config. capQuery itself and its explicit non-positive helper semantics are unchanged. Full config.ts scan found no other numeric environment parsers with the same failure.
Require Node 24+ and use node:24-alpine in all seven app Dockerfiles, matching the Node 24 CI runtime. The locked OpenAI SDK requires Node 22+ and all 585 locked Node engine declarations accept Node 24.0.0. Runtime-reference audit covers the root manifest and README, all 21 Docker stages, deployment Markdown (including its base-image example), and the deployment/getting-started HTML pages. No Node 20 support claims remain outside locked third-party dependency metadata; the sample ticket Node Version remains historical mock data. pnpm, Turbo, and healthchecks are unchanged. Verified JSON and all Docker FROM declarations, scoped Prettier checks, HTML parse/format round-trip checks, dependency-engine compatibility, and git diff --check. Container verification uses this committed tree.
Recognize GPT and known o1/o3/o4 model families in the shared provider validation, including aliases and dated snapshots. Keep unknown custom provider deployment names allowed, and preserve Claude rejection under OpenAI. Callsite audit: - validateConfig routes response, confidence, classifier, and sentiment model settings through validateModelProvider before pipeline startup. - AuxiliaryModel uses the same guard for explicit per-call overrides; ConfidenceScorer, TicketClassifier, and analyzeSentiment use this class. - Provider-specific SupportAgent and ResponseGenerator remain unchanged; pipeline configuration validation guards their selected response model. - No other duplicated provider-family checks exist in the AI package. Regressions: all four config stages and direct auxiliary construction. 31 pre-fix failures become 91 passing focused cases; 169 related tests pass. Targeted Prettier, ESLint, and AI no-emit typecheck pass. Addresses AI-A02.
Addresses AI-A03. Distinguish explicit security vulnerabilities, data loss, and production outages from ordinary HIGH incidents. Preserve heuristic CRITICAL as an urgency floor without lowering model CRITICAL. Affected callsites: TicketClassifier.heuristicClassify, TicketClassifier.classify (model merge and degraded fallback), and AIPipeline.classifyTicket through the shared classifier. The pipeline inherits the corrected fallback without a separate change. Validation: regression cases failed in 14 assertions before the fix; all 61 classifier and auxiliary OpenAI tests now pass. Targeted Prettier, ESLint, and AI TypeScript noEmit checks pass.
Fix R1-DELIVERY-A2: read escalationRequiredReason alongside responseState after a no-op escalation claim. DELIVERED with a retained marker now fails for manual attention without claiming a handoff, clearing its reason, or regenerating/reposting. Explicit null markers preserve clean success. Apply the same invariant to the initial prior-response gate so retries cannot erase the failure, and to pending-response sweep no-op outcomes. The sweep still isolates failures per response and retains its scan policy. Symbol callsites: - reportSkippedRecoveryEscalation: recoverRequiredEscalation and recoverPendingResponse, ai-response.ts:312,360. - handleAiResponse: queue/src/index.ts:4 exports it; worker/src/index.ts:27 imports it and :85 registers AI_RESPONSE. Its handler regression suite exercises the initial gate and both no-op recovery entry points. - settleStrandedResponse: handlePendingResponseSweep at pending-response-sweep.ts:231; the public handler is exported by queue/src/index.ts:15 and registered by worker/src/index.ts:94. - stageRow test fixture: recovery no-op describe block in queue/src/__tests__/ai-response.test.ts, including the initial-gate tests. - settleAfterRead test fixture: pending-response-sweep.test.ts:367,379,398 covers concurrent ESCALATED, conflicting DELIVERED, and clean DELIVERED. Validation: red/green demonstrated for both no-op recovery routes, the initial-gate retry, and sweep no-op; 130 tests pass across both affected suites. Prettier and queue tsc pass. ESLint has zero errors and five pre-existing warnings in untouched sweep fake methods (baseline six).
Forward nonempty query.version from the shared search builder and the separate searchDocs builder, matching searchEvidence. Keep unversioned arguments and existing error handling unchanged. Callsite audit: searchCode, searchAgUiCode, and searchAgUiDocs delegate to the shared builder; searchDocs builds its own arguments; searchEvidence already forwards version. Pipeline searchDocs/searchCode callers omit version and retain their behavior. Support-agent searchEvidence callers are unchanged. exploreDocs, queryKnowledgeBase, and fallbackSearch do not accept PathfinderQuery. Regression evidence: serialized JSON-RPC assertions fail for v1/v2 across all four wrappers before the fix (8 failures; omitted-version cases pass). After the fix, 55 Pathfinder tests pass. Prettier, targeted ESLint, AI typecheck, and AI build pass.
Derive labels from the final rounded score using the documented thresholds. Use the existing 25/NEUTRAL empty-input contract for degraded fallbacks and empty trend periods, while retaining degraded flags and billed token usage. Callsite audit: account-scoring consumes the label for its Prisma mapping and still skips sentiment writes whenever degraded is true; no handler changes were needed. getSentimentTrend forwards analyzed scores and labels and now emits a consistent neutral pair for empty periods, which remain excluded from trend deltas. Both APIs keep their existing exports and shape. Validation: genuine red-green regression (19 expected failures before the fix, 59 passing tests after) across sentiment, OpenAI auxiliary transport, sentiment trend, and account-scoring. Changed-file Prettier, ESLint with zero warnings, AI tsc --noEmit --incremental false, and git diff --check pass.
Follow up AI-A03: avoid CRITICAL escalation for prevention questions and negated incident mentions. Inspect nearby phrase context within individual sentences so a separate affirmative incident still escalates. Preserve the original data-loss HIGH baseline and all model priority floors. Affected callsites: TicketClassifier.heuristicClassify and TicketClassifier.classify; AIPipeline.classifyTicket inherits the guarded heuristic without its own change. Guards intentionally cover common phrasing rather than full natural-language inference. Validation: 11 new assertions fail before the follow-up; all 77 classifier and auxiliary transport tests pass afterward. Targeted formatting, lint, AI noEmit typecheck, and diff checks pass.
Apply the requested API generation filter before deciding whether search_evidence needs its unfiltered fallback. Use the same case-insensitive v1-deprecated marker as final reply validation, and filter fallback hits before remembering them. Mixed usable results retain requested_version scope. Callsite audit: both support-agent searchEvidence calls share this filter across CopilotKit/AG-UI docs/code. Pathfinder strict parsing/version forwarding and final reply validation remain intact. The single spend(), shared abort signal, request/source limits, v1/unknown semantics, scope labels, and thrown retrieval failures are preserved; no other production searchEvidence callers exist. Validation: genuine SDK/aimock red-green roundtrips cover deprecated-only and uppercase-title first hits, useful fallback evidence, mixed results without fallback, v1/unknown retention, and fallback error propagation. Full AI suite: 27 files, 744 tests passed. Prettier, ESLint, AI typecheck, and isolated AI emit build passed. Temporary dependency/declaration links removed before commit.
Flatten every publishable Discord part in the support-agent string stream, without duplicating the first part. Append separate details only for web output and retain the publication gate. Add an aimock regression through the real SupportAgent, reply validation, and formatter. The pre-fix stream loses steps 12-24, citations, and the footer; short replies on all platforms and routed-draft confidentiality are also covered. Callsite audit: generateStreamingResponse has only test callers in this repo. The Discord adapter already consumes every formatted part; the web QA route consumes separate details metadata. The queue saved-suggestion continuation issue is separate from this stream API. Legacy streamed chunks remain unchanged. Validation: 770 AI tests pass, including 125 targeted pipeline/formatter tests; Prettier and ESLint pass for both changed files; AI tsc noEmit and isolated production build pass on Node 24.
Return a failed AI_RESPONSE result when both the DELIVERED write and its deliveryConfirmed fallback fail after successful delivery. Include both errors for manual attention and stop before progress 100 while preserving pipeline teardown and the one-post response claim. R1-DELIVERY-A4: the regression failed against a9db1d8 with success true, then passed with the fix. The fallback-marker-success control remains green. A retry schedules recovery without regenerating or reposting. Callsite audit: read the complete ai-response handler, its worker registration and JobResult failure consumer, and the pending-response sweep delivery marker consumer. Existing primary-response recovery prevents repeat posts; successful fallback markers still repair state. No owed-escalation reason persistence changes (R1-DELIVERY-A3). Validation: 108 handler tests and 256 queue tests pass; scoped Prettier and ESLint, queue tsc --noEmit, queue emitted build, and diff check pass.
R1-DELIVERY-A3: insert escalationRequiredReason with the primary response, so failed later writes or an interrupted enqueue cannot replace the exact low-confidence/suppression reason with generic delivery recovery. An insert failure prevents publication. Keep a failed platform post's reason ahead of suppression by replacing its marker and responseError together; report both write and queue failures when that replacement is not durable. Symbols and callsites: - handleAiResponse (ai-response.ts:480) is exported by queue/src/index.ts:4 and registered by apps/worker/src/index.ts:85. It computes the owed reason before its primary message insert, then publishes at most once. - recoverRequiredEscalation (ai-response.ts:294), called by the prior-response gate at :569, preserves deliveryFailed metadata for replaced reasons. - enqueueEscalationAtomically is unchanged; its callers at ai-response.ts:306, :354, :1049 and pending-response-sweep.ts:142 consume the durable reason. - trackResponsePersistence (ai-response.test.ts:264) provides real write-derived retry/sweep snapshots with rollback for the new failure-path regressions. Validation: genuine red/green retry and sweep regressions; 137 tests pass; changed-file Prettier and ESLint pass; queue no-emit typecheck and isolated queue emit build pass on Node 24. A4 delivery-confirmation failure handling is unchanged. Parent owns the CR loop; no push is part of this fix.
R1-DELIVERY-A3 follow-up: a concurrent retry can enqueue the persisted handoff while the original postResponse is pending. Guard the later delivery-reason replacement by PRIMARY_AI_RESPONSE and PENDING so it cannot recreate an owed marker on ESCALATED or DELIVERED. When the guard matches no row, retain responseError without changing lifecycle metadata. Symbols and callsites: - handleAiResponse (ai-response.ts:480), exported at queue/src/index.ts:4 and registered at apps/worker/src/index.ts:85, now conditions the marker replacement on the still-pending primary response. - recoverRequiredEscalation (ai-response.ts:294), called by the gate at :569, and enqueueEscalationAtomically (:120; calls :306, :354, :1064 and pending-response-sweep.ts:142) retain their existing transition behavior. - holdPlatformPost (ai-response.test.ts:341) and trackResponsePersistence (:264) drive the real overlapping retry regression. The analogous DELIVERED/interrupted-owner case proves no transient marker is stranded. Validation: both race cases fail before the guard and pass after it; 139 handler/sweep tests pass. Changed-file Prettier and ESLint, queue typecheck, and isolated queue emit build pass. A4 remains unchanged. Root owns CR and integration; no push or main-checkout mutation.
Follow up R1-DELIVERY-A1 by preferring nonempty formatted.parts over the aliased first-part text when saving the complete publishable suggestion. Join parts with paragraph separators, append optional details, and fall back to text when parts is absent or empty. Never append private draft or handoff metadata. Affected symbol and callsites: - handleAiResponse: imported by apps/worker/src/index.ts:27 and registered as the AI_RESPONSE handler at line 85; re-exported by packages/outpost/queue/src/index.ts:4. - Test calls remain in packages/outpost/queue/src/__tests__/ai-response.test.ts (import at line 95; new multipart/fallback call at line 741). - Full handler and queue-module audit found one suggestedResponse writer. Adapters already receive the full formatted object; shadow diagnostics remain separate. The analogous streaming pipeline remains unchanged. Validation: the multipart regression genuinely failed with Summary only before the fix; afterward all 105 handler tests pass, including existing web/no-details coverage and the empty-parts fallback. Targeted Prettier, ESLint, queue typecheck, and queue emitted build pass. Private fields and duplicate first parts are excluded by the exact expected suggestion.
Force OpenAI and Luna in the shared auxiliary call options. Override the pipeline test config's provider and all four model fields, then validate the same config object consumed by the default pipeline clients. Test callsite audit: every ConfidenceScorer, TicketClassifier, and analyzeSentiment call in auxiliary-openai.test.ts uses the pinned options. All three AIPipeline construction sites use the file-local config mock; the injected SupportAgent already defaults to Luna on OpenAI Responses. No production behavior or process-wide provider/model env mutation changed. AI-A10 red/green: Anthropic provider plus Claude model overrides produced 23 failures before the fix and 46 passes afterward. Both files also pass with default env and OpenAI provider plus Claude model overrides. Full AI suite: 776 tests pass. Targeted Prettier/ESLint and AI typecheck pass.
R2-AI-A02 keeps legacy Pathfinder parsing tolerant while making the strict searchEvidence boundary fail when a code result cannot be cited. Call sites audited: PathfinderClient.searchEvidence production use is via SupportAgent in packages/outpost/ai/src/support-agent.ts for the initial docs lookup and dynamic follow-up tool calls; direct strict-boundary tests live in packages/outpost/ai/src/pathfinder-evidence.test.ts. SearchResult kind/sourceUrl consumers in support-agent.ts and support-reply.ts were checked; no exported shape changed. Checks: red pathfinder-evidence regression failed before implementation; green pathfinder-evidence and pathfinder parser tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-DELIVERY-A1 keeps the SHADOW_MODE SYSTEM row aligned with the complete publishable response stored as suggestedResponse. Changed-symbol callsites: added private completePublishableResponse in packages/outpost/queue/src/handlers/ai-response.ts; it is called once in handleAiResponse to derive publishableResponse, which is then used by the suggestedResponse ticket update and the outpost-shadow SYSTEM message create. Added findShadowMessageCreateCall test helper used by the existing shadow assertion plus the new web-details and multipart Discord shadow regressions.
R2-AI-A05 prevents a caller-cancelled MCP setup from leaving the initialized session id cached for the next strict evidence request. The cancellation cleanup is limited to the caller AbortSignal path during notifications/initialized; non-cancellation notification failures remain best-effort. Call sites audited: connect is used directly by pathfinder tests and internally by callTool; production strict evidence flows through SupportAgent's Pick<PathfinderClient, 'searchEvidence'> dependency in support-agent.ts. reset remains private except disconnect, and no exported type or method shape changed. Checks: red pathfinder-evidence sequential cancellation regression failed before implementation; green pathfinder-evidence and pathfinder parser tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-AI-A04: recognize did/has/was/were/have/had/will question starters in the existing sentence-local CRITICAL guard. Keep the old HIGH baseline, healthy model CRITICAL, affirmative incidents and mixed sentences intact. Add 21 contrastive heuristic rows and real Responses LOW/MEDIUM and CRITICAL coverage. Genuine red: 20 cases initially, plus 15 additional same-pattern cases. Green: 120 classifier/auxiliary tests; scoped Prettier and ESLint, inclusive AI typecheck and isolated AI build all pass. Changed-symbol call-site audit: TicketClassifier.heuristicClassify is called by TicketClassifier.classify (classifier.ts:54), AIPipeline.classifyTicket fallback (pipeline.ts:432), and classifier.test.ts. The model-backed caller is pipeline.ts:426, constructed at pipeline.ts:110. Tests exercising the class also include auxiliary-openai.test.ts; pipeline.test.ts and pipeline-groundedness.test.ts retain their injected classifier mocks. No signature, exports, tag/category heuristics or priority merge changed.
R2-AI-A01 separates internal handoff diagnostics from public support prose validation while preserving strict public summary/details/appliesTo checks and the nonempty route reason rule. Changed-symbol callsites: validateSupportReply still parses supportReplySchema, checks route handoffReason nonempty, validates evidence, and now calls validateProse only for public summary/details/appliesTo. supportReplyText and supportReplyDetails were not changed; tests assert private handoffReason does not appear in either renderer. No queue, provider, or formatter callsites changed.
R2-AI-A06 fixes code citation URL construction for GitHub repository URLs that end in .git/. The normalization now removes a trailing slash before stripping a trailing .git suffix, preserving the existing provider allowlist, branch, ref, and source path behavior. Call sites audited: blobUrl is private and called only by parseSnippets for code snippets. Its sourceUrl output flows through tolerant searchCode/searchAgUiCode, strict searchEvidence, SupportAgent source filtering/deduplication, and support-reply citation validation. No exported interface shape changed. Checks: red pathfinder parser and strict evidence suffix regressions failed before implementation; green pathfinder and pathfinder-evidence tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-AI-A03: keep the message from known validation and investigation-budget errors for PipelineResult.handoffReason. Retain the existing 2000-character return bound, generic public handoff and empty-diagnosis fallback. Other provider/transport failures still propagate without replacement. Tests first: eight genuine failures for diagnosis loss, then 88 passing pipeline tests covering both classes, GitHub/Discord/web public and stream privacy, cause exclusion, low-confidence suppression and verifier skipping. Scoped Prettier/ESLint, inclusive AI typecheck and isolated AI build pass. Changed-symbol call-site audit: AIPipeline.generateSupportResponse is called by queue/handlers/ai-response.ts:735, apps/web/src/app/api/qa/route.ts:65, and generateStreamingResponse at ai/src/pipeline.ts:466, plus owning tests. Private handoffReason is consumed by queue escalation/shadow diagnostics at ai-response.ts:786,875; web/streaming select formatted output. No signature, export, queue state machine or provider error conversion was changed.
R3-AI-A02 requires strict searchEvidence calls for search-code and search-ag-ui-code to reject legacy JSON-array results that cannot be cited. The legacy parser can still return URL-less entries, but the strict boundary now also uses the requested tool name when enforcing code citation provenance. Changed symbol call sites audited: PathfinderClient.searchEvidence is called by SupportAgent primary retrieval in packages/outpost/ai/src/support-agent.ts, by SupportAgent follow-up tool retrieval in packages/outpost/ai/src/support-agent.ts, and by strict evidence tests in packages/outpost/ai/src/pathfinder-evidence.test.ts. Public tolerant wrappers searchCode/searchDocs/searchAgUiCode/searchAgUiDocs still route through search/parseSearchResults unchanged, with a regression control for URL-less legacy searchCode JSON. Validation: red pathfinder-evidence regression failed at 233baee with both strict code tools resolving URL-less legacy JSON; green pathfinder-evidence passed 28 tests; related pathfinder parser/wrapper test passed 41 tests; scoped Prettier, scoped ESLint, explicit AI typecheck, explicit AI build, and git diff --check passed.
R3-AI-A03 encodes read_source blob citation paths with the same per-segment encoding used for GitHub contents requests. Changed symbol/callsites: - Added encodeSourcePath(path) in packages/outpost/ai/src/support-agent.ts. - Updated read_source githubJson contents callsite to use encodedPath. - Updated read_source remembered sourceUrl callsite to use encodedPath in the GitHub blob URL. - Added support-agent regression rows for docs/My Guide.md and docs/100% ready (setup).md. - Added support-reply validation regression proving encoded blob citations validate while raw whitespace URLs remain rejected.
R3-AI-A05 extends the sentence-local critical incident suffix guard so passive absence reports such as security vulnerabilities were not found and production outage was not reported do not create a CRITICAL floor. Call-site enumeration for the changed classifier surface: TicketClassifier.heuristicClassify is called by TicketClassifier.classify in packages/outpost/ai/src/classifier.ts; AIPipeline.classifyTicket uses the classifier fallback in packages/outpost/ai/src/pipeline.ts; AIPipeline constructs TicketClassifier in packages/outpost/ai/src/pipeline.ts; queue AI response handling reaches it through pipeline.classifyTicket in packages/outpost/queue/src/handlers/ai-response.ts; classifier, auxiliary-openai, pipeline, and queue tests cover the concrete and mocked callers. Tests: node ../../node_modules/vitest/vitest.mjs run ai/src/classifier.test.ts --config vitest.config.ts --no-cache --reporter=dot; node node_modules/prettier/bin/prettier.cjs --check packages/outpost/ai/src/classifier.ts packages/outpost/ai/src/classifier.test.ts; node node_modules/eslint/bin/eslint.js packages/outpost/db/src packages/outpost/ai/src packages/outpost/queue/src packages/outpost/shared/src; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.json --noEmit --incremental false; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.build.json --outDir /tmp/outpost-r3-passive-ai-build-1789930006 --incremental false
Follow-up for R3-AI-A05: evaluate every critical incident mention in a sentence so a first passive absence mention cannot hide a later affirmative mention of the same phrase. Regression rows cover: Security vulnerabilities were not found in staging, but security vulnerabilities were found in production; Data loss was not found in staging, but data loss was found in production. Red log showed both returned HIGH before the fix; green log shows 84/84 classifier tests passing after the scanner evaluates all mentions. Call-site context remains the same as the prior A05 commit: TicketClassifier.heuristicClassify is called by TicketClassifier.classify; AIPipeline.classifyTicket uses classifier fallback logic; queue AI response handling reaches it through pipeline.classifyTicket. No public symbols or signatures changed. Tests: node ../../node_modules/vitest/vitest.mjs run ai/src/classifier.test.ts --config vitest.config.ts --no-cache --reporter=dot; node node_modules/prettier/bin/prettier.cjs --check packages/outpost/ai/src/classifier.ts packages/outpost/ai/src/classifier.test.ts; node node_modules/eslint/bin/eslint.js packages/outpost/db/src packages/outpost/ai/src packages/outpost/queue/src packages/outpost/shared/src; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.json --noEmit --incremental false; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.build.json --outDir /tmp/outpost-r3-passive-followup-ai-build-final-1789930263 --incremental false
R3-AI-A06: suppressed legacy responses now put deterministic groundedness/enforced-lint reasons before generic legacy generator reasoning, while explicit OpenAI support-agent route/error diagnostics still take precedence. Changed-symbol call sites: AIPipeline.generateSupportResponse is called by the queue AI response handler, web QA route, and generateStreamingResponse; PipelineResult.handoffReason is consumed by queue escalation persistence and by pipeline tests; public web/streaming paths publish formatted output. lintDraft, assessGroundedness, SUPPRESSED_RESPONSE_TEXT, and exported type shapes keep their existing call sites and signatures. Tests: red pipeline/pipeline-openai suite failed 2/91 as expected before the fix; final related Vitest suite passed 110/110; scoped Prettier, ESLint, AI typecheck, AI build, and git diff --check passed. Logs and round-3-fix-R3-AI-A06.md saved in the CR ledger.
A support answer puts its code inside a numbered step or a quoted example.
proseOutsideFences recognized a fence only at columns 0-3, so every one of
those was scanned as prose and the reply discarded:
- Example: -> Raw HTML is only allowed inside code
```tsx in a support reply
<CopilotKit runtimeUrl="…" />
```
> ```tsx -> same
> <CopilotKit runtimeUrl="…" />
> ```
10. Example: -> Support reply link URL must belong to
```text validated source evidence
https://example.invalid/…
```
The renderer publishes all three as <pre><code>: inert code that mounts no
element and resolves no link. A correct answer was thrown away and the ticket
escalated. The docstring already claimed this contract; the column test was a
different question.
A fence opens wherever its container's content starts, and four columns past
that is an indented code block with no fence at all — block structure, not a
line pattern. proseOutsideFences now asks the parser the chat renderer runs
(mdast-util-from-markdown + micromark-extension-gfm, already imported here for
the destination and definition scans) which nodes are code, and blanks their
lines. That removes the column from the question for every container nesting
and for indented code at once, and it closes the opening and closing match
together — the same regex carried the assumption on both, so fixing only the
opener would have left a container fence open and failed the reply with
"Unclosed code fence in support reply" instead.
Both refusals in this function are kept, not inherited by the new shapes:
- "Invalid code fence" still fires for a backtick fence whose info string holds
a backtick — CommonMark says that is not a fence, and refusing rather than
guessing is deliberate (R13-AI-A03 owns whether it should be). It now skips
lines the parser places inside code, which the column scan could not.
- "Unclosed code fence" exists because supportReplyDetails appends the
applicability, version and sources footer to `details` and an open fence
swallows it. Only a top-level fence can: a blank line closes a block
container before anything inside it. Asked directly, by parsing the field
with a probe paragraph appended, so the refusal follows the reason instead of
hand-tracked fence state. All four top-level spellings still reject; a
container's fence does not.
Verification: the three shapes above flip to accepted in
round13-code-literals-supplied-claim-probe; its A03 and A04 rows do not move,
and neither does any row of round13-gfm-supplied-claim-probe,
round13-ai-residual-claim-probe or -probe2 measured against a build of
1252e2b. Regressions are red against the unchanged implementation (17 failed /
241 passed) and green after (258 passed): 14 accept rows across list items,
block quotes, nesting, tilde fences and indented code; 8 reject rows proving a
container carrying a paragraph is still validated and so is anything after its
fence closes; the open-fence rows in both directions. The existing block
container row in "keeps prose that resolves to no reference definition" was
spelled //example.invalid/steal, which the raw-URL scan cannot match — it now
carries https:// so it fails if the fence is read as prose. The renderer half
is pinned in apps/web with literal fixtures through the app's real ChatMessage
-> ReactMarkdown + remark-gfm: each shape publishes one <pre><code>, no anchor
(not even for the bare address and www. host GFM linkifies elsewhere), no
mounted element, and the appended footer survives a container's fence while a
top-level one swallows it.
Call sites: proseOutsideFences is module-private with one caller, validateProse
(support-reply.ts:327), which validateSupportReply (:408, applying it at :466) applies to summary,
details and appliesTo — so all three fields are covered, and must be, since
escapeMarkdown escapes neither backticks' block role nor container markers for
appliesTo. validateSupportReply reaches production through support-agent.ts:347
and the ai barrel (index.ts:11). publishedDestinations is unchanged in
behaviour; it now shares the new parseMarkdown helper. supportReplyDetails,
whose footer the unclosed-fence refusal protects, is consumed by formatter.ts:47.
A run of three or more backticks that opens and closes on one line is a code span. The renderer publishes `<p><code>literal code</code></p>` and keeps the body inert: the JSX arrives as text, and neither a bare address nor a `www.` host GFM linkifies in prose becomes a link. The validator refused the same bytes with "Invalid code fence in support reply", discarding a valid reply, because any line opening with three backticks whose remainder held another backtick was rejected outright. Three backticks is exactly how an answer quotes a literal that already holds one — up to and including a fence, ```` ```tsx ````. Ask the parser instead of the line. The walk that already collects code blocks now also collects each code span's opening position, and the refusal fires only where no span starts at the run. Its purpose is unchanged and still covered: a run left open is not a span, and a fence whose info string holds a backtick is not a fence, so the renderer commits to neither and publishes the marker as literal paragraph text with every following line read as prose. Refusing rather than guessing there is deliberate. A span does not exempt the rest of its line: `proseOutsideInlineCode` already strips runs of any length, which is why a mid-line triple span was accepted before this change and a line-initial one was not — the guard ran first. Those two are now one code path. Call sites — all inside packages/outpost/ai/src/support-reply.ts: collectCodeBlocks -> collectCodeNodes (+ CodeNodes): sole caller proseOutsideFences:151; no other reference in the repo. proseOutsideFences: sole caller validateProse:344. proseOutsideInlineCode: sole caller validateProse:356. validateProse: sole caller validateSupportReply:483. validateSupportReply is the only exported symbol reached, via ai/src/index.ts:11 and ai/src/support-agent.ts:347. No signature changed, so no call site needed an edit. Renderer behavior pinned in apps/web/src/__tests__/qa-components.test.tsx against the app's real ReactMarkdown + remark-gfm: nine spans published inline and inert, three open runs published as literal text. (cherry picked from commit df2c06b12c950cfc9560eb2655c490fabb7112e6)
`appliesTo: 'Runtimes on <v2 releases'` was refused as raw HTML. The renderer publishes `<p>Runtimes on <v2 releases</p>` — a literal '<', markup-free. `appliesTo` is the field the schema dedicates to version applicability, so a '<vN' range is that field's own vocabulary, and the guard's own boundary sat between two spellings of one sentence: `<1.9` passed because a digit is not a tag name, `<v1.9` did not. A '<' opens a tag only where the grammar can close one. `publishesRawHtml` asks the parser the chat renderer runs — `mdast-util-from-markdown` plus the GFM extension, already used for code blocks, definitions and destinations — whether any `html` node resolves, and refuses only then. The comment, processing-instruction, declaration, CDATA and closing-tag forms the pattern covered are `html` nodes too, so all of them stay refused; so does a '<vN' opening an unquoted attribute value closes into a real tag. The autolink pre-strip goes with the pattern it existed for: an autolink is a `link` to the grammar, never `html`, and the destination checks above still hold it to the evidence set. Same assumption, second site: `proseOutsideInlineCode` skipped from '<' to the end of the line whenever no '>' followed, so every code span after an inert '<vN' range went unmasked and its contents — example URLs included — were read as prose. No '>' closes no tag, so there is nothing to skip. The guard runs over the masked prose, not the original text: the masking is what implements "only inside code", and it stays deliberately stricter than the grammar for a span crossing a line. That control is unchanged. Call sites: - publishesRawHtml (new) — support-reply.ts:446, the only caller, inside validateProse; collectHtml is called only by it and itself. - proseOutsideInlineCode — support-reply.ts:378, the only caller. - validateProse — support-reply.ts:510, the only caller, applied to summary, details and appliesTo alike, so one change covers all three fields. - parseMarkdown — now five callers (:136, :156, :240, :265), all read-only. (cherry picked from commit ebfd0ab8358693e21f5e6bf624f787a951d372e0)
The inline destination scan matched every literal `](` in prose, with no requirement of a preceding link label or of a CommonMark-legal destination, so `The literal punctuation ](not a link) is part of this sentence.` was read as a citation of `not` and the whole reply was discarded. The renderer publishes no anchor for that sentence at all. Ask the grammar which links exist instead. inlineDestinations parses the prose with the parser and GFM extension the chat surface runs, walks link and image nodes, and recovers each destination's own offset from the node's range by counting the label's matching bracket. The check and the mask that follow are unchanged, so a destination is still read as the reply spells it rather than as the parser decodes it, escaped and angle-bracket forms still resolve, and relative, protocol-relative and non-HTTP destinations of real links are still held to the evidence. A real link that looks like punctuation keeps being checked: this renderer publishes `arr[i](x)` as an anchor, and the node walk sees it where a label-shaped pattern would not. Second site, same assumption. `proseOutsideInlineCode` decides which same-line code spans are masked from the raw-URL and raw-HTML scans, and it skipped a destination-shaped run after every literal `](` because a backtick inside a real destination is a literal rather than a code delimiter. So `See ](`https://example.invalid/x`) here.` — which the renderer prints as inert code and publishes no anchor for — had its code span stepped over and the address in it read as an invented citation. `inlineDestination` now reports the offset of the `]` alongside the destination's own start, and `inlineDestinationColumns` turns those into the per-line columns the scan consults. Where a label really did open the destination the skip is unchanged, so `[guide](`https://…`)` stays a link with a backticked destination and stays rejected. Renderer-pinned in the qa-components fixture table: the two spellings differ only in the label, and only the labelled one emits an anchor. Regression evidence: 6 failures against the unchanged implementation for the literal-punctuation rows, including one alongside a genuine evidence link. Seven must-reject rows (subscript, image, image nested in a link label, nested/escaped/multiline labels, second invented link) and three masking rows passed before and after. Call sites — all inside packages/outpost/ai/src/support-reply.ts: inlineDestination:278 — sole caller inlineDestinations:330. collectInlineLinks:302 — sole caller inlineDestinations:322, plus itself. inlineDestinations:320 — callers inlineDestinationColumns:350 and validateProse:533. inlineDestinationColumns:342 — sole caller validateProse:495. proseOutsideInlineCode:425 — sole caller validateProse:501; the added parameter is the only signature change and has no other reference. validateProse:484 — sole caller validateSupportReply:635, applied to summary, details and appliesTo alike. validateSupportReply is the only exported symbol reached, via ai/src/index.ts and ai/src/support-agent.ts. Its signature is unchanged, so no call site outside this file needed an edit. Integrates R13 A04: cherry-picked from cd8f81bc1712aab845b4f683a55e833730cafbe1 and its corrective 8c9ef62eb3e8f3d2f8d5559009da5c4bd0e6c521, squashed into one concern.
An inline destination was read as the reply spelled it; a reference definition arrived from the parser already decoded. So one published href had two spellings on opposite sides of the evidence check: See [docs](https://docs.copilotkit.ai/search?a=1&b=2). rejected See [docs][d] accepted [d]: https://docs.copilotkit.ai/search?a=1&b=2 Both publish href="…/search?a=1&b=2" — the cited evidence — and the first was discarded with "Support reply link URL must belong to validated source evidence". Recorded through the app's real ReactMarkdown + remark-gfm and through this parser (logs 00, 01). `inlineDestinations` now returns the destination the link or image node already carries, decoded, alongside the span the raw-URL scans mask. `checkUrl` compares that value, so backslash escapes and HTML character references resolve the same way here as in a definition. This is per syntax because the renderer is. The same probe shows a CommonMark autolink and a GFM autolink literal publishing their address exactly as written — `<…?a=1&b=2>` links to `…?a=1&b=2` — so the two autolink scans keep comparing the spelling. Normalizing all four alike would ground two of them on a URL no reader reaches. `unescapeMarkdownDestination` and `checkUrl`'s `markdownDestination` parameter lose their only caller and are gone. `referenceDefinitions` moves onto `parseMarkdown`, leaving one parser configuration in the file, so the definitions it masks cannot drift from the destinations `publishedDestinations` checks; log 02 records zero difference across 14 definition shapes, GFM-specific ones included. Call sites, all module-private to support-reply.ts: unescapeMarkdownDestination removed; had one caller, checkUrl checkUrl :537 :543 :554 :557 :564 inlineDestination(s) :349 (columns) :536 (validateProse) inlineDestinationEnd :300 :453 referenceDefinitions :492 :542 parseMarkdown :147 :169 :253 :324 :378 :399 validateProse :636 — one caller, all three fields No export changed; nothing outside this file references any of them. Tests: the entity-encoded spellings of one evidence URL across inline, angle-inline, image, definition, `&` and `&`; the same in summary, details and appliesTo; a five-syntax table checked in both directions so the relation holds rather than the values; and four encoded destinations no evidence decodes to, which stay refused. The renderer-side pin in qa-components gains the four rows that record the decode/no-decode split it rests on. (cherry picked from commit a36728a3ac3ac74d92fc95ff05e82d863789e4e5)
…ion class A GFM autolink literal written inside an emphasis or strikethrough run reached the evidence check with the closing run attached. The raw-URL scan finds an address by pattern and runs it to the next space, and the punctuation class it then trims against holds no `*`, `_` or `~`, so `**Read <evidence url>**` was compared as `<evidence url>**` and refused. The renderer publishes the run outside the anchor: the reader clicks exactly the cited evidence URL. A reply that was correct and fully grounded was discarded, wrapped as InvalidSupportReplyError, and escalated to a human over a reason naming neither the field nor the URL. Ask the parser where the literal ends instead. `autolinkLiterals` reports the span the grammar gives each bare address; each one is held to the evidence set in the spelling the reader clicks and then masked from the pattern scan. The check runs before the mask, so the scan is nowhere looser than it was: a literal the parser finds in a region that scan deliberately still reaches — the contents of a code span crossing a line — stays refused. Only the literal form is taken; `[label](…)`, a reference definition and a CommonMark `<…>` autolink carry their own delimiters and are already checked in the form each syntax publishes. Same-pattern audit over the finite delimiter dimension — every run this grammar forms (`*`, `**`, `***`, `_`, `__`, `~`, `~~`) in each position around an evidence, a non-evidence and a scheme-less address, 126 rows, oracle = hrefs the app's actual ReactMarkdown + remark-gfm publishes: 31 rows diverged from the renderer before, 0 after, every one of them in the reject direction. The other two scans do not carry the pattern: the `<…>` scan is bounded by its own `>`, and publishedDestinations is grammar-derived already. Host-case canonicalization (R14-PKG-A03) is a separate contract and is verified unmoved — `www.` accepted, `WWW.`/`Www.` refused, as recorded at d8c932f. Call sites of the changed symbols: - AutolinkLiteral (new, module-local): support-reply.ts:334, :360. - autolinkLiterals (new, module-local): one call site, support-reply.ts:589, inside validateProse. - validateProse (body only, signature unchanged): one caller, validateSupportReply at support-reply.ts:683. - checkUrl / maskMarkdownDestination (module-local closures, unchanged): one new call each at :590 / :591, alongside the existing :575/:581, :576/:582. - Exported surface of support-reply.ts unchanged. validateSupportReply consumers are unaffected: support-agent.ts:347 and the AI barrel index.ts:11. Red before the fix: 10 of the new AI rows failed at checkUrl (support-reply.ts :524 via :557). Green after: 347/347 support-reply, 3995/3995 packages/outpost, 816/816 apps/web. No existing expectation weakened or removed.
…me from `validateSupportReply` checked `summary`, `details` and `appliesTo` as the model wrote them, and then `supportReplyDetails` moved both of the fields it composes before anyone saw them. Trimming `details` deleted the four leading spaces that made it an indented code block, so an example address the evidence check had credited as inert republished as a live link to a host no evidence mentions. Normalizing `appliesTo` onto one line deleted the same indentation, and the escape it then ran through covered Markdown structure only, never the '@', '.' and '/' a GFM autolink literal is built from, so a `www.` host and a bare email published as links the same way. The corruption ran in the other direction too. The escape rewrote the characters of a cited URL — `…/reference/my-guide` reached the reader as `…/reference/my%5C-guide` — and escaping a `[label](…)` citation into plain URL text had GFM relinkify it onto that same rewritten address, so escaping a link neither removed it nor kept it. Auditing the rest of the helper found the third instance: an inline destination is decoded, so the source list published a cited `…?a=1&b=2` as the address that reference decodes to. `details` now keeps its first line where the model put it and loses only blank edges; `appliesTo` copies through the spans the grammar already resolves to a link or image and escapes between them, over the inline constructs only; the source list writes its destination so it decodes back. `validateSupportReply` then checks the composed string against the same evidence set, which is what holds whichever transform is added next. Call sites of the changed symbols: - supportReplyDetails (exported): formatter.ts:47, support-reply.ts:759 (supportReplyText), support-reply.ts:645 (new: the composed check), barrel re-export index.ts:13, and support-reply.test.ts. Its output is unchanged for every reply whose fields carry no leading indentation, no address in the applicability and no character reference in an evidence URL. - escapeMarkdown (module-private): sole caller is now publishedAppliesTo (support-reply.ts:707, :710); it no longer collapses whitespace, which publishedAppliesTo does instead. The unrelated same-named helper in apps/teams-bot/src/cards/ticket-created-card.ts is untouched. - validateProse (module-private): callers are validateSupportReply:636 (the three fields, unchanged) and :645 (new). - collectInlineLinks / parseMarkdown (module-private): gain one caller each, resolvedLinkSpans:672. - publishedAppliesTo, resolvedLinkSpans, trimBlankEdges, escapeDestination are new and module-private, one caller each. Red against the unchanged implementation at d8c932f: 9 of the new rows fail (4 indented-details, 2 applicability-refusal, 2 applicability-fidelity, 1 source list). Published strings are pinned against the app's real ReactMarkdown + remark-gfm in apps/web/src/__tests__/qa-components.test.tsx, including the two spellings publication used to emit.
…an guesses `proseOutsideInlineCode` paired backtick runs by hand, which meant hand-parsing everything else on the line that can hold a backtick without opening a span. Two of those skips decided the outcome, and both answered a different question than the renderer's. The tag skip tested `/^<(?:!|\?|\/?[a-z])/i` and then jumped to the next '>' on the line, whether or not the grammar closes a tag between them. Where it jumped too far it swallowed a code span: `Compare <b, ` + span + `, and c> here.` is published as `Compare <b, <code>…</code>, and c> here.`, no tag and no anchor, yet the example address inside that span was checked as a citation and the reply discarded and escalated. Where the jump landed past a backtick it left the run after it to pair with a later one, and the span that mispairing invented covered raw HTML the renderer publishes as prose — `<b\`> x \`<script>…</script>\` y` passed a guard whose rule is that HTML is allowed only inside code. This renderer escapes that HTML rather than mounting it, so what got through is the policy boundary, not markup that runs. The parser already reports each code span and its exact range, and a span it reports can be neither an attribute nor a destination: a backtick inside either opens nothing it reports. Masking those ranges retires both skips at once. Only what a span encloses is blanked, never its delimiters. The raw-HTML check re-reads the masked view, and `Use <b ` + span + ` > carefully.` — which the renderer escapes whole — would become `Use <b > carefully.`, a tag the grammar does close, if the backticks went with the contents. Unchanged: spans crossing a line are left whole and stay subject to the prose checks, deliberately stricter than the renderer; reference definitions stay exempt; a backtick inside a real destination still opens no span; and the invalid/unclosed fence refusals are untouched. Call sites of the changed symbols, all module-private to support-reply.ts: - proseOutsideInlineCode: second parameter is now the line's code-span ranges rather than its destination columns. One call site, validateProse (:527). - inlineDestinationColumns: removed. Its only call site was validateProse. - InlineDestinationSpan.openerStart: removed with it. Produced by inlineDestination (:297), read by nothing else. - sameLineCodeSpanContents, collectInlineCode, CodeSpanContents: new, called from validateProse (:523) and from each other (:378, :403). - inlineDestinationEnd and inlineDestinations lost a caller each and keep their remaining ones (:297, :560). Tests, oracle = the app's own ReactMarkdown + remark-gfm: four rows red at b21d810 (two inert-'<' shapes wrongly refused, two mispaired spans wrongly accepted), plus the controls that stay green — masking must not manufacture a tag, an address outside the span is still refused, a '<' the grammar does close is still raw HTML, and a span closed on a later line is still read as prose. The published form of every shape is recorded in apps/web.
Planned production takedowns and security-scanning setup no longer force CRITICAL priority. Recognized incident reports keep their priority floor. The infinitival exemption requires a named planning/prospective frame; unrecognized completed-action phrases retain their original behavior. Validation: genuine red/green for the three original failures plus completed- action regressions, 1,781 classifier tests, 4,110 core tests, scoped formatter/ lint, AI noEmit typecheck and isolated AI build. Existing expectations retained. Call sites: TicketClassifier.heuristicClassify is called by classify in classifier.ts and the fallback in pipeline.ts; the public signature and exports are unchanged. Pipeline.classifyTicket is consumed by the queue AI-response handler. New guards are local to heuristicClassify. Combines isolated R14-AI-A05 commits 903ac8c, d6e8404 and 0b6e993 after root verification; source and tests exactly match the final isolated tree.
proseOutsideFences refuses a line whose leading backtick run is followed by
another backtick, unless the parser reports a code span opening at exactly that
line and column. collectCodeNodes recorded openings only, so every later line of
a span the grammar had already closed carried a run the guard could not account
for. `Use `` a\n```b `` here.` was refused as an ambiguous fence while the chat
renderer publishes it as one paragraph holding one <code> element and no fence
at all, and the reply was discarded and the ticket escalated.
Record the lines each span already covers where they begin — every line after
the one that opened it, through the line its closing run is on — and skip the
refusal on those. The docstring at :137 already states the rationale the guard
now follows: only a run left open is the ambiguity it exists for.
Scope is the refusal alone. A span crossing a line is still not masked, so its
contents stay conservatively read as prose: an ungrounded address inside one is
still refused, now by the link check that owns that question rather than by the
fence guard. A run genuinely left open still refuses, including on a line that
also carries a closed span. The unclosed-fence footer probe, the column-free
block detection from R13-AI-A02 and the same-line masking from R14-AI-A02 are
untouched.
Call sites of every symbol changed, all module-private to support-reply.ts:
- CodeNodes (interface, new `spanContinuations` field): declared :96,
constructed :160, read :174 and :178. No other file names the type.
- collectCodeNodes: defined :108, recursive call :118, called :161.
- proseOutsideFences: defined :157, sole caller validateProse :529.
No export surface changes; nothing outside this module references any of them.
Red before the fix: the three accepted rows threw 'Invalid code fence in support
reply' at support-reply.ts:163; two refusal rows threw the fence error instead of
the link error that should own them. The renderer fixtures in
apps/web/src/__tests__/qa-components.test.tsx pin the published form — one <p>,
one <code>, no <pre>, no anchor — against the app's real ReactMarkdown +
remark-gfm.
R14-AI-A04. Owns L14-3's multiline-structure rows.
A citation of a scheme-less `www.` host was held to the evidence set only when the host was written in lowercase. GFM linkifies that prefix in any case, and the scan that finds these addresses matches any case, but the scheme rebuilt before the comparison was rebuilt only for `www.` — so `WWW.copilotkit.ai/reference/provider` reached `canonicalSourceUrl` with no scheme, parsed as nothing, and was refused, while the grammar-derived destination for that same sentence carried `http://` and grounded. The two halves of one check disagreed about which addresses exist, and a reply that cited its evidence exactly was discarded and escalated to a human over a link whose reader would have landed on the validated source. A sentence-initial capital is the ordinary trigger: the schema permits it and the instructions do not forbid it. The renderer is the oracle for why accepting these is correct, not a case-folding assumption. Its anchor for each spelling carries the address in the case it was written, over http:// — recorded against the app's real ReactMarkdown + remark-gfm in apps/web/src/__tests__/qa-components.test.tsx — and a URL host is case-insensitive, so every one of those hrefs resolves to the cited evidence URL. Each accept row asserts that resolution. Only the prefix test changes, and only its case sensitivity. The host is still folded by the URL parser rather than here, so nothing compares a path or a query case-insensitively: `Www.copilotkit.ai/reference/Provider` is still refused. So are an ungrounded host in every case, a different path on the evidence host, and the bare evidence host with its path dropped. The recorded reason for choosing http:// over https:// is preserved and extended rather than replaced. Red before the change, against this same source otherwise unmodified: the three capitalized accept rows threw 'Support reply link URL must belong to validated source evidence' from checkUrl, reached through the grammar-derived autolink-literal loop — the structural autolink fix in b21d810 moved where the refusal surfaces but not its cause. The eight refusal rows passed before and after. Call sites of the changed closure, all seven inside `validateProse`, which is the only reader of this scheme reconstruction: the inline-destination, reference-definition and autolink-literal loops, the two raw-URL scans (the second the only caller passing `allowProsePunctuation`), and the grammar-derived `publishedDestinations` loop. `validateProse` has one caller, `validateSupportReply`, which runs it over summary, details and appliesTo; that function is called in production only at ai/src/support-agent.ts:347 and is re-exported from ai/src/index.ts. No app package imports it.
…tails Compose applicability from parser-resolved spans and bound literal URLs only when escaping changes their destinations. Escape nested images and escaped-bang adjacency so metadata cannot introduce remote images. Keep final composed-output grounding validation intact. Call sites checked: resolvedSpans and composeAppliesTo are private to publishedAppliesTo; boundedAutolink is called only from composeAppliesTo; supportReplyDetails remains consumed by formatter, serialization and barrel exports. Public signatures unchanged. Verified isolated red-green regressions, 4036 core and 830 web tests, typechecks and build. Combined branch targeted support-reply and formatter suites: 445 passing. Full combined verification follows.
…nsumer `negativeObservationPrefix`'s trailing object phrase repeated three arms whose whitespace overlapped: `\s*[,/]\s*` carried a run on both sides and the word and coordinator arms both opened with `\s+`. Every gap in a separator run could therefore be split two ways, so a prefix the group cannot accept made the engine enumerate 2^gaps partitions before failing. `heuristicClassify` runs synchronously in the queue worker's AI_RESPONSE handler on the raw, uncapped body — only the model call is truncated — so a merely comma-heavy ticket (a pasted log, a quoted CSV) pinned a worker CPU. Measured on the emitted method: 10 separators 2.4ms, 20 separators 90.8ms, 24 separators 1.45s, 26 separators unfinished in 3s. Each arm now consumes only the whitespace preceding its own token. A token is never whitespace, so every leading run is pinned to a whole gap. The `and`/`or` arm must still assert a following gap, so it takes a single `\s` and leaves the remainder to the next arm — `\s` plus the next arm's leading run accepts exactly what `\s+…\s+` did. The accepted language is unchanged; the exponent is gone rather than bounded, and the body is neither capped nor truncated. Verified equivalent, not assumed: 1,050,256 exhaustive regex inputs (token sequences to depth 3 over gap widths 0-3) plus 333,453 separator-only inputs to depth 4, 0 divergences; the 608 distinct bodies in classifier.test.ts and 500,000 randomised negated-observation sentences replayed through both builds, 0 divergences. Pumping every other repeated group in the file — the `without` and `no` filler lists, the negated-predicate opener, the remaining-incident list, the takedown determiners, the outage adverbs, `\w+n't`, `\w+ly` — found no second occurrence: all stay flat to 20,000 repetitions before and after. The regression runs in a hard-bounded child process. In-process it could not fail, only hang: Vitest's testTimeout is a timer and the timer needs the event loop the blocked regex is holding. It asserts completion within a generous liveness budget rather than a duration, so it does not depend on machine speed. Against the unchanged guard it fails in 60s with "did not finish"; the eleven semantic controls — four negated separator-joined object phrases, each paired with the same sentence minus its negation — pass identically on both sides. Call-site verification before the original commit is recorded in the raw 13-symbol-call-sites.txt report. negativeObservationPrefix has one unchanged caller in hasNonIncidentPrefix. separatedObjectToken has one caller in that regex; incidentMention gains that one reader. New bounded-process helpers and types are referenced only by classifier-backtracking.test.ts and its harness. No public runtime exports changed. Root adds this omitted record to the commit message only; the tested source tree is unchanged.
A reference definition can write its destination on the line after `]:`, where the block container it sits in re-states the markers the parser already consumed. Those markers are not part of the node the parser reports, so they are still inside the raw range it gives, and skipping whitespace alone stopped on the '>' and reported the marker as the destination. The marker was masked and the destination was not, so a URL the definition check had already approved decoded reached the raw scans in the spelling they cannot accept: an entity-encoded query and a destination ending in an escaped ')' were both refused, and the reply was escalated to a human over the link its own evidence backed. A literal destination passed by accident, which is why the suite could not see it. The padding is now stepped over by the grammar's own bound: indentation, one line ending, then one '>' per open block quote with its optional space, and a '>' only ever at the start of a line. An angle destination may hold no line ending, so its scan stops at one rather than closing on a later line's marker. The span still covers the destination alone, so a raw URL written into a label or a title stays subject to the scans, and the canonical, escape and entity comparisons are untouched. Call sites checked: `destinationSpan` has one caller, `referenceDefinitions`, whose signature and `ReferenceDefinition` shape are unchanged; `continuationPadding` is new and private to `destinationSpan`. No exported symbol changed. Verified red before green: 12 of the 27 new matrix rows failed on the previous locator, all 415 prior support-reply controls held. 4224 core and 852 web tests pass, with cold AI and web `--noEmit`, an isolated AI build, and scoped Prettier and ESLint clean.
`no` and `without` cancelled an incident mention unconditionally, so "No
data loss root cause has been identified yet.", "No production outage
postmortem has been written." and "Without a data loss postmortem we
cannot close the incident." all landed at HIGH, indistinguishable from
the genuine absence report "No data loss has been reported.".
The file already decides this reading twice: `negatedIncidentObjectHead`
requires the mention to head the negated object rather than sit as a
modifier inside it, and the `incidentToolingHeads` rationale names
"postmortem", "root cause" and "mitigation" as heads that presuppose an
instance and "must leave it standing". The two determiner arms were the
ones that never consulted the suffix.
`negatedIncidentObjectHead` is not reusable here: it is a verb's object,
where the phrase runs to the clause end, so any further bare noun
continues it. `no`/`without` take a determiner phrase in subject or
complement position, where a predicate legitimately follows — "No data
loss has been reported." must stay cancelled. So the test is inverted:
enumerate the heads that end the cancellation instead of the predicates
that keep it. `incidentPresupposingHeads` is that enumeration, the
complement of `incidentToolingHeads`, and an unlisted continuation keeps
the absence reading.
"report" and "incident" are named as presupposing in the same comment but
stay out of the set: `objectPhraseHandoff` already reads those two words
as heading an absence report about the incident ("no data loss reports"),
and splitting those readings is a separate decision, recorded as unsettled
rather than taken here.
Callsite enumeration (performed before this commit):
incidentPresupposingHeadSuffix -> hasNonIncidentPrefix (classifier.ts:247), sole reference
hasNonIncidentPrefix -> classifier.ts:449, consumed at :468, sole callsite
heuristicClassify -> classify() floor (classifier.ts:54, :72-78)
-> AIPipeline.classifyTicket (pipeline.ts:454)
-> queue/src/handlers/ai-response.ts:779, persisting
classification.priority onto Ticket.priority
test-only: ai/src/test-utils/bounded-heuristic-classify.child.ts:37,41
no other module references the new symbols
Tests: the three reported spellings join
observedAbsenceVersusStandingIncidentContrasts paired with absence
controls, plus a closed grid —
cancellingDeterminerVersusPresupposingHeadContrasts — crossing the two
determiners with the three heads the suite already pins and the three
incident terms. 12 rows / 24 sentences, each pair also asserted not to
collapse. Red on the base source: 48 failures (12 standing sentences x 3
boundaries + 12 relations), every absence control already green. Green
after the fix.
Brings the Slack ticket mirror (#150), the worker health/boot rework (#262) and the worker poll-loop claim fencing (#224) under this branch's support-agent and AI-response-lifecycle work. One content conflict, resolved by union: packages/outpost/queue/src/__tests__/ai-response.test.ts Both sides edited the file's import preamble and nothing else collided. Ours imported `Message` (@prisma/client) and `EscalationPayload` for the migration / support tests; upstream widened the vitest import to include beforeAll/afterAll and added a suite-level SHADOW_MODE guard that clears an inherited SHADOW_MODE=true and restores it afterwards, because seven of its claim-fencing and mirror tests assert the non-shadow path. Resolution keeps every element of both: the 8-name vitest import, both of our type imports, and upstream's ambient-shadow guard. The only deviation from either side is that `JobHandlerContext`'s type import is hoisted to sit with the other imports rather than below the guard block — same semantics, no statement between imports. Both `Message` and `EscalationPayload` were confirmed still referenced in the merged body (PersistedResponse, escalationPayload()). No test was lost. Block counts reconcile exactly: base 98, ours +10, upstream +8, merged 116, and a title-set comparison shows no block from either side dropped. The single base title upstream still carries and the merge does not ("succeeds without claiming a handoff when the response is already DELIVERED") is a base test upstream never touched and this branch deliberately rewrote into two stricter cases, so taking our deletion is correct. packages/outpost/queue/src/handlers/ai-response.ts auto-merged; reviewed against both contracts and both hold: - Upstream's mirror contract is intact. `delivery: SlackMirrorDelivery` is declared before the delivery branches and claimed by every branch upstream claimed it in — 'shadow' before the shadow try, 'withheld'/'delivered' on a successful post, 'post-failed' in the catch — and the SLACK_MIRROR enqueue still runs after reportProgress(85), gated on isMirrorableSource + isSlackMirrorEnabled. - This branch's contract is intact. nonDeliveryEscalationReason is still computed before the message.create that persists it, escalationReason still gives delivery failure precedence, the PENDING-guarded updateMany and escalationReasonPersistenceError threading survive, and the both-writes-failed early return is unchanged. - The two do not interfere: the mirror renders aiMessage.content (pipelineResult.response), which this branch never retargeted — its publishableResponse change only touched ticket.suggestedResponse and the shadow SYSTEM row. Symbol audit before committing: updateJobProgress gained a required third claimToken argument upstream; every callsite is upstream's own and already passes it, and this branch only ever uses context.reportProgress. JobHandlerContext is unchanged. JobType carries both SLACK_MIRROR (upstream) and PENDING_RESPONSE_SWEEP (ours). The new Job.claimToken column and its two migrations do not collide with this branch, which adds no migration. The test file's platforms mock already stubs readSlackMirrorConfig, isSlackMirrorEnabled and isMirrorableSource. .env.example and docs/deployment.md auto-merged as disjoint additive sections; both sides' edits verified present. Emergent behavior worth noting, not altered here: when both the DELIVERED state write and the deliveryConfirmed marker write fail, this branch's early return now also skips the mirror enqueue, which upstream documents as running "regardless of whether the draft was delivered". That path is a double Prisma failure, the enqueue is itself a DB write that would fail and only log, and the job returns failure for retry — so it is benign, but it is a real interaction between the two changes rather than something either side chose.
R15-AI-A04. `parseSourceUrl` accepts `'` and '`' in an evidence URL and this renderer publishes both in an href — `…/provider's` and `…/provider%60name`, each the `new URL(evidence)` canonicalization of the cited address. The raw-URL scan in `validateProse` runs an address to the next character outside a class that excludes both, and the CommonMark `<…>` autolink was the one link syntax with no grammar-derived span masking it from that scan: an inline destination, a reference definition and a bare GFM literal all have one. So the scan read a truncated prefix of the cited address, failed to find that prefix in the evidence set, and escalated to a human a reply whose only citation was its own evidence and whose reader would have clicked straight through to it. `uriAutolinks` asks the parser the renderer runs where each autolink's address begins and ends, and `validateProse` holds it to the evidence in the spelling this syntax publishes before masking that span — the same check-then-mask order the three syntaxes above it use, which is what keeps the pattern scans no looser than they were. The class is not widened, the delimiters are not stripped, nothing else is masked, and canonicalization is untouched: an ungrounded address, a non-HTTP or credentialed URI, an address the autolink production does not close around, and an address in a code span crossing a line all stay refused. Private-helper audit, done before committing: - `uriAutolinks` is new and file-local; one callsite, support-reply.ts:698, inside `validateProse`. - `collectInlineLinks` is reused unchanged; callsites are now support-reply.ts:377, :415, :457 and :853. - No other helper in the file is touched. Every helper here is file-local — the module exports only `supportReplySchema`, `SupportReply`, `parseSourceUrl`, `validateSupportReply`, `supportReplyDetails` and `supportReplyText`, none of which changed signature or behaviour outside the refusal above. Red on b9f889d: 8 failed / 454 passed in ai/src/support-reply.test.ts. Green here: 462 passed in that file; core 77 files / 4475 tests; web 53 files / 856 tests; both cold typechecks; fresh non-incremental AI emit at 23 modules / 69 files with no test material in dist.
NathanTarbert
left a comment
There was a problem hiding this comment.
Thanks Jerel, this is a huge amount of careful work. The writeup is honest about what the live runs do and don't show, which makes it much easier to review. The paragraph-plus-details format reads really well, and the MCP headers answer is exactly the kind of reply we want to be sending.
The first thing I'd like to look at is what happens when read_source is pointed at something that isn't a file. If the model asks for a directory like packages/react-core/src/hooks, GitHub returns a list instead of a single file. The parse in support-agent.ts around line 253 then throws. That error isn't one of the two the pipeline turns into a handoff, so it comes back up as a failure. In the worker, the job retries and fails the same way each time. In web QA, the user sees the raw validation error. Symlinks and submodules hit the same path, and so does a bad path around line 230. If these came back to the model as tool errors, it could correct itself and keep going.
The second is that the GitHub calls in support-agent.ts go out without a token. GitHub caps unauthenticated requests at a low hourly limit per IP, and each read_source makes two calls. A busy worker will hit that cap, and when it does, the investigation fails outright instead of telling the model the source is unavailable. Passing the token the rest of the app already uses would fix the limit. Turning a 403 into a tool result would cover the rest.
The third is about escalation after an interrupted run. In ai-response.ts around line 839, the escalation reason is now written when the row is created, before anything is posted. If the worker stops between creating the row and posting, the retry escalates with only the low-confidence reason. The human who picks it up isn't told the reporter got no reply, and the delayed takeover that waits for an in-flight post is skipped. Main handles this case today by flagging the delivery as failed, so it would be good to keep that.
A smaller one: for web tickets, formatStructured puts the footer in text, and the details get appended after it. So the saved suggestion and the shadow record end with the footer, then the technical details. The dashboard chat isn't affected because it renders details separately.
I also had a question. The investigator runs with store: false and doesn't request reasoning.encrypted_content. My understanding is that multi-turn runs with reasoning sometimes need that to carry reasoning between turns. Your September 21 runs clearly made tool calls fine, so I may be wrong here. Did you see any "item not found" errors along the way?
Last, I'd like us to agree on the default before this merges. As it stands, merging moves every AI call to Luna unless AI_RESPONSE_PROVIDER is set to anthropic. Since the description says posting should stay off while evaluation continues, it might be safer to land this with Anthropic as the default and flip it once the matched old/new comparison is done. Happy to talk it through if you see it differently.
Thanks again for all of this!
read_source and read_release set errorFunction=null, so every non-404 GitHub failure escaped the model loop and aborted the whole investigation instead of letting the agent correct itself. Six predictable failures escaped: - 403/429/5xx and any other non-ok status threw `GitHub evidence request failed` - a transport failure (fetch reject) threw a raw TypeError - a non-JSON body threw SyntaxError out of response.json() - a directory path returned an entry array and failed `size` parsing (ZodError) - symlink/submodule/other non-file blobs failed the base64 parse (ZodError) - an unsafe path threw InvalidSupportReplyError, ending the run githubJson now reports these as a discriminated result, and both tools turn them into bounded tool output the model can act on: `invalid_path` (returned before any URL is built, so an unsafe path is never fetched), `not_a_file`, `unreadable` and `unavailable` with a sanitized reason — access_denied, rate_limited, upstream_error, invalid_response or transport_error. No GitHub response body, header or address reaches the model or the trace. Preserved deliberately: - repository/ref/path allowlists, and `not_found` + `too_large` as they were - the size check still precedes base64 decoding - only a validated single file resolves to a pinned sha and enters `remember` - a failed call still spends one of the six; the budget never resets - errorFunction stays null, so programmer errors still surface; AbortError and the 60s deadline still terminate the run rather than becoming tool output - InvalidSupportReplyError on final output and InvestigationBudgetError unchanged Callsite audit: - pipeline.ts:171-194 is the only consumer of investigate(). It rethrows anything that is not InvalidSupportReplyError/InvestigationBudgetError, which is why a 403 previously failed the ticket and triggered a worker retry; those cases now route through the agent instead. Its two handled error types are untouched, so pipeline-openai.test.ts is unaffected. - index.ts re-exports both error classes; neither signature changed. - docs/support-agent.md documented "API outages still raise errors" and is updated to the new status vocabulary. Tests: real SDK multi-turn aimock runs where a failed or wrong-resource read is followed by a corrective grounded source read that succeeds, plus budget exhaustion on six failed reads, sanitization assertions, and abort termination. All 16 new cases fail on the parent commit. Two tests that asserted the old escaping behavior (bad path throwing, malformed payload throwing) are converted to assert the structured result; no assertion was weakened. Verified: packages/outpost suite 4491 passed / 77 files (baseline 4475 / 77), cold typecheck clean, eslint clean, prettier clean.
`escalationRequiredReason` is stamped on the primary response when the row
is created, before any publication. The prior-response gate read it first
and went straight to `recoverRequiredEscalation`, so an attempt interrupted
anywhere around the platform post — PENDING row, owed reason, no
`deliveryConfirmed`, no `responseError` — was settled as if the answer had
already gone out. Two consequences on exactly the rows that most need care:
- the delayed takeover was bypassed. That delay exists so recovery never
races a post still in flight; the owed marker predates the post, so
acting on it summoned a human against an answer about to land.
- the handoff reported `deliveryFailed: false` and handed the human the
bare generation-time reason ("Low AI confidence (25%)"), which reads as
"a weak answer was sent" to someone whose reporter may have nothing.
Delivery was also genuinely unrecorded for these rows: a successful post
that owes a handoff keeps the row PENDING, so it took neither the DELIVERED
transition nor the `deliveryConfirmed` flag, and nothing distinguished it
from an attempt that died before posting.
- record `deliveryConfirmed` when publication succeeds on a row held
PENDING by its owed handoff (non-fatal; the escalation below it is what
the reporter is owed).
- route an owed handoff with no recorded delivery outcome, still owned by
this job, through the existing `schedulePendingResponseRecovery`. The
takeover re-enters the same branch after the delay and escalates.
- report `deliveryFailed` from proven delivery rather than from a recorded
error, so unknown reads as failed instead of as delivered.
- compose the reason in one shared helper, `recoveredHandoffReason`: the
stored promise verbatim when the outcome is known, and the promise plus
"the reporter may have received no response at all / a human must verify
the thread" when it is not.
- PENDING_RESPONSE_SWEEP: an owed handoff now outranks `deliveryConfirmed`.
Repairing such a row to DELIVERED dropped the promised human AND left the
marker on a settled row — the pair the gate's first branch fails on — and
it uses the same helper so a sweep and a takeover read identically.
Preserved: proven delivery never reports a failure, a recorded delivery
failure keeps its diagnostic and escalates with no delay, escalated rows stay
idempotent through the compare-and-set, no path reposts, and the claim
fencing (`responseJobId` = this job), database-clock recovery window, and
single CAS/transaction are untouched.
Callsite audit — `deliveryConfirmed` / `escalationRequiredReason` are written
and read only by these two handlers (repo-wide grep; no app, API or dashboard
reader). `recoverRequiredEscalation` has one caller, updated for its new
ticketSource argument. `hasConfirmedDelivery` widened to a Pick, all four
callers still type-check. `recoveredHandoffReason` is consumed by the handler
and the sweep. `schedulePendingResponseRecovery` gains a caller whose row
satisfies the same PRIMARY/PENDING/own-claim preconditions its CAS requires.
The schema change is comment only — no column, no migration.
Tests: nine added (interrupted low/medium/suppressed rows deferred then
escalated undelivered; a same-job retry waiting out an in-flight post;
delivery recorded on a PENDING owed-handoff row; a recorded failure escalating
at once; a stale duplicate not deferring a claim it lacks; sweep escalating a
confirmed delivery that owes a handoff; sweep keeping a failure diagnostic).
Seven of them fail on c602d6c. Two existing fixtures gained
`deliveryConfirmed: true` because the handler now writes it on the rows they
model, with their assertions unchanged.
…p installation
read_source and read_release were the only GitHub callers in this package that
sent no credential — githubJson set an Accept header and nothing else — so every
ref resolution, file read and release lookup ran on the anonymous 60 req/hour
budget shared by the whole host, while the worker already holds App installation
credentials for the same two repositories.
Reuse those credentials rather than introducing a new one. A new
github-evidence-auth module reads GITHUB_APP_ID / GITHUB_PRIVATE_KEY /
GITHUB_INSTALLATION_ID exactly as shared/platforms/registry.ts and
apps/github-app/src/lib/github-client.ts already do, and mints an installation
token through @octokit/auth-app (already a dependency of this package — no
dependency change). No gh CLI token, no personal access token, no key file.
Behaviour:
- Tokens are narrowed to `contents: read`; metadata:read is inherent. The
strategy is built on first use, so a run that never reads GitHub signs no JWT.
- auth() is invoked per request so @octokit/auth-app owns caching and pre-expiry
refresh; nothing here holds a bearer string that would die an hour into uptime.
- Anonymous public reads remain available, but only when none of the three
variables is set. A partial or malformed configuration rejects at the point of
use instead of degrading into a quietly rate-limited anonymous success.
- A signing or token-exchange failure is re-thrown carrying only the underlying
error's class name — not its message and not as `cause` — because that error
reaches worker logs and run reports and the underlying one can quote the PEM
or the freshly minted token.
- SupportAgent defaults githubAuth from the environment, so pipeline consumers
(pipeline.ts constructs `new SupportAgent({ pathfinder, model })`) authenticate
without opting in. The option exists as a test seam only.
Integrated with the recoverable-evidence-failure change already on this branch
(c6f4922), which this commit was authored against a sibling of. Credential
resolution now happens inside githubJson's recoverability boundary, so a
configured-but-unusable credential becomes bounded tool output rather than an
escaping throw that fails the ticket:
- githubEvidenceHeaders() is awaited in its own try, ahead of the fetch, and a
GitHubEvidenceAuthError becomes `unavailable` with the new sanitized reason
`auth_unavailable`. Only the bare reason crosses to the model; the message
naming the missing variable, and the underlying error, stay worker-side.
- The catch converts GitHubEvidenceAuthError and nothing else. Any other fault
in the resolver is a programmer error and is re-thrown, preserving the
branch's errorFunction=null contract; rethrowIfTerminal still runs first, so
AbortError and the 60s deadline terminate the run as before.
- The request is abandoned on a credential failure, never retried without the
credential, so a misconfigured host cannot silently fall back to a
rate-limited anonymous read.
- Everything else the recovery commit established is retained unchanged: the
sanitized GithubResult union, safeParse on every payload, not_found /
invalid_path / not_a_file / too_large / unreadable shaping, the six-call
budget that a failed call still spends, and terminal abort propagation.
The existing auth test that asserted a thrown GitHubEvidenceAuthError is
deliberately re-pointed at the integrated behaviour: it now asserts the model
receives a bounded `unavailable` / `auth_unavailable` tool result, that no
request reached api.github.com (proving no anonymous retry), and that the
missing-variable name does not appear in what the model sees. The other two auth
tests and all 18 github-evidence-auth unit tests are unchanged.
Callsite audit — every githubJson caller, all in support-agent.ts:
- L322 read_source ref -> commit resolution ...... passes githubAuth
- L341 read_source contents read ................. passes githubAuth
- L414 read_release releases/tags ................. passes githubAuth
No other call exists; grep for api.github.com / raw.githubusercontent across
packages/outpost/ai/src finds no second GitHub fetch. pipeline.ts:106 is the only
production SupportAgent construction and takes the env default. The origin stays
hardcoded and repositories stay enum-allowlisted, so the model can neither supply
a URL nor pass an Authorization header to an arbitrary source; headers are
assembled in one place and a token can only ride on a request to GitHub's API.
docs/support-agent.md gains the three App variables under Configuration and
`auth_unavailable` in the documented failure-reason vocabulary.
Tests use synthetic placeholders only. Verified: support-agent.test.ts 46 passed
(43 recovery cases preserved unchanged + 3 auth cases), github-evidence-auth.test.ts
18 passed, packages/outpost suite green, cold ai typecheck clean, eslint and
prettier clean on the touched files.
(cherry picked from commit 42e55fb5929c24a4149df2f5784e9ecbc22b6f56)
…e shaping The cherry-pick of the App-installation auth change (78d6f8c) onto the recoverable-evidence-failure change (c6f4922) created a boundary neither parent commit could test on its own: credential resolution now sits inside githubJson's recoverability try. Two properties of that seam were only implied by the resolved code, so they are asserted here rather than left to a future reader's reading. - A token-exchange failure must not leak the signing key into model-visible output. The auth module already strips the underlying error to its class name (github-evidence-auth.test.ts covers that in isolation), but nothing proved the sanitized value is what actually reaches the tool result after shaping. This drives a real githubEvidenceAuthFromEnv resolver — through the SupportAgent test seam — whose exchange throws an error quoting a synthetic PEM, then asserts the tool result the model receives carries `auth_unavailable` and neither the PEM header nor its body. - A non-credential fault in the resolver must still escape the model loop. The catch converts GitHubEvidenceAuthError and re-throws everything else, which is what keeps the branch's errorFunction=null contract intact; without a test, a later widening of that catch to `catch (error)` would silently launder programmer errors into a status the investigator routes past. A resolver that throws a TypeError is asserted to reject the investigation. Both also assert no request reached api.github.com, so neither path can regress into an anonymous retry of a request whose credential failed. Synthetic placeholders only; no credential here works anywhere. Mutation-checked against the resolved source: widening the catch to all errors fails the second test; falling back to an anonymous read or hoisting the headers call out of the recoverability try fails the first (and the re-pointed partial-config test from 78d6f8c). support-agent.test.ts 48 passed, packages/outpost 4523 passed / 78 files, cold ai typecheck clean, eslint and prettier clean.
`formatStructured(reply, 'web')` returns a two-pane result: `text` is the
summary pane and ALREADY ends in WEB_FOOTER, `details` is the disclosure pane.
That split is a UI contract, not a serialization — but the two sinks that hold
exactly one string reassembled it by appending `details` to `text`:
packages/outpost/queue/src/handlers/ai-response.ts completePublishableResponse
packages/outpost/ai/src/pipeline.ts streaming fallback composer
so the durable `suggestedResponse`, the SHADOW_MODE SYSTEM record and the web
string stream all read summary → footer → details, with the footer and the
disclaimer stranded mid-response and the sources printed after the sign-off.
Fixed at the formatter, the only place that knows where the trailing matter
goes. `FormattedResponse` gains an optional `completeText`, composed by the web
branch from the same pieces in reading order (summary, details/sources,
disclaimer, footer last) and left absent wherever `text`/`parts` is already the
whole response. Both composers now read it through one shared `publishableText`
helper instead of re-deriving the join in two places.
Deliberately NOT done: splitting the footer back out of text the formatter did
not write, and smuggling details through `parts` — the latter would have
duplicated them in the streaming composer, which joins parts and then appends
details.
Callsite audit — every reader of a FormattedResponse:
queue handlers/ai-response.ts:841 suggestedResponse -> completeText (FIXED)
queue handlers/ai-response.ts:906 shadow SYSTEM row -> completeText (FIXED)
ai pipeline.ts:489 generateStreamingResponse -> completeText (FIXED)
apps/web api/qa/route.ts:78,91 streams `text`, sends `details` as metadata
-> UNCHANGED; the native QA disclosure still renders the two panes
apps/github-app github-poster.ts:53 `text` + feedback -> UNCHANGED
shared platforms/{discord,slack,teams,github}.ts postResponse
-> UNCHANGED; Discord keeps `parts` precedence and ordering, GitHub keeps
its single collapsed <details>, neither carries a `details` field
shared platforms/types.ts FormattedResponse (adapter-side copy)
-> untouched; adapters never read `details`, so nothing to mirror
A FormattedResponse without `completeText` — a hand-built fixture, or a value
from an older run — still serializes exactly as before: `publishableText` keeps
the historical text/parts-then-details join as its fallback branch. Suppressed
runs go through `format()`, which sets neither `details` nor `completeText`, so
no withheld draft or internal handoff reason gains a new way out.
Tests: red first on all three suites (footer at index 50 ahead of the sources at
77 in the web stream; the shadow record asserting the stranded-footer string),
green after. The pre-existing queue fixtures that assert the old join are
correct as written — they carry no footer in `text`, so they pin the
compatibility branch rather than the bug, and are left alone. The queue module
mock now spreads the real `@copilotkit/outpost/ai` so these assertions exercise
the real serialization rather than a stub.
Verified: 77 files / 4493 tests pass, `pnpm typecheck` clean across db+ai+queue+
shared, apps/{web,github-app,worker} `tsc --noEmit` clean, eslint and prettier
clean on every changed file.
(cherry picked from commit fd0a4238a46f6e5715f048a851b2bb8e2ec9e609)
…dline
githubJson awaited githubEvidenceHeaders(auth) before the fetch that carries the
investigation's 60-second AbortSignal, and authorization() took no signal at all.
A stalled App token exchange therefore held the investigation open past its own
deadline: the bound only ever covered the request, never the credential the
request was waiting for.
Thread the investigation signal through the headers seam and race the
authorization against it. The pending exchange is detached rather than cancelled:
@octokit/auth-app serves every concurrent caller with the same installation and
permissions from one shared in-flight request (verified against 7.2.2 — two
concurrent auth() calls issue a single HTTP request), so cancelling on behalf of
one abandoned investigation would reject the others. Detaching lets it finish and
fill the SDK cache for whoever is still waiting, and leaves SDK-owned caching and
pre-expiry refresh untouched.
The investigation's signal never reaches the SDK. The exchange gets a fixed
10-second HTTP deadline of its own instead, applied per attempt through the
`request` option — the only seam the strategy exposes, since auth() forwards no
per-call request. Without it a hung exchange is permanent: the SDK hands the same
stalled promise to every later caller.
Cancellation is terminal. It rejects with the signal's own reason, so a cancelled
authorization is indistinguishable from the cancelled fetch it was about to
authorize, and can never be mistaken for a retryable GitHubEvidenceAuthError.
An already-aborted signal does no credential work at all — no strategy
construction, no JWT, no exchange — and abort listeners are dropped whether the
authorization succeeds or fails.
Also:
- Replace `createAppAuth as unknown as InstallationTokenFactory` and the
redundant Promise cast with a typed adapter. The overloaded AuthInterface is
not structurally assignable to the one narrow signature this module needs;
adapting lets both sides keep their real types. Test mocks drop `as never`
along with it.
- Lazy strategy construction now runs inside the same sanitized boundary as the
exchange, so a synchronous createAppAuth validation failure cannot echo the
PEM. Still no cause and no message forwarding — only the error's class name.
Mutation-checked: eight mutations, all killed, no survivors. Reverting the single
wiring line to its previous form fails `abandons a cancelled investigation
without issuing the evidence request` by timeout; reverting both production files
to 42e55fb fails 11 tests behaviourally, none by module resolution.
Default provider and reasoning settings are unchanged. Tool failure shaping is
untouched and still belongs to the concurrent tool-recovery finding, whose
terminal error guard continues to receive cancellation unwrapped.
(cherry picked from commit a67858c4790803f8f18b8a7d82f77b1409432085)
`formatDiscord` split at `DISCORD_MAX_LENGTH - 50`, then appended a 109-UTF-16-unit `STANDARD_FOOTER` to the last part and cut the result back to fit with `slice(0, 1997) + '...'`. The reserve was less than half the thing it reserved for, so the cut was the normal path for a band of response sizes rather than an overflow guard — and the end of the last message is the footer, so what it removed was always published copy. Verified against the real formatter, not read off the source. A 1892-char body lost the 👎 half of the feedback prompt; 1949 and 1950 lost the Docs URL mid-link; 3898-3900 took the same damage one part further along. At 1896 the cut landed between the two code units of 👍 and published a bare \uD83D. Sweeping body lengths 1-6000 across six shapes and eleven adversarial inputs (36,011 bodies), the previous code broke the posting contract on 3,229 of them: 1,068 trailing ellipses, 1,068 mangled footers, 25 unpaired surrogates. It is now 0. The splitter is told what the footer costs and returns a last part with room for it, so there is nothing left to truncate and the truncation branch is gone. Closing a fence at a split point is charged against the budget too, rather than fitting by accident in the 46 characters that happened to be spare. The fallback hard cut steps back a unit rather than halving a surrogate pair. Same splitter, second defect: the continuation was built with `slice(splitAt).trimStart()`, which stripped all leading whitespace and not just the newline the split consumed — inside a fenced block that is the snippet's own indentation, so code a reader was told to copy came out rewritten. `splitPoint` now reports how much of the break it consumed and exactly that much is dropped. Call sites audited. `shared/platforms/discord.ts:177` posts `parts` and attaches the buttons to the last message, which now also carries a complete footer; its own `splitMessage` only runs when `parts` is absent, which is the <=2000 single-message case. `publishableText` and its two consumers — `queue/handlers/ai-response.ts:947` and `pipeline.ts:491` — keep their shape and stop recording a truncated response durably. `github-poster.ts` is on the GitHub path and untouched. `truncated: true` is deliberately left on a split response: no runtime code branches on it (only fixtures and the type do), and changing it would alter a public field for every split Discord response, outside this fix. It now means "split", which is worth renaming separately. Not fixed, same class, different file: `splitMessage` in `shared/platforms/discord.ts:244` hard-splits at 2000 with no surrogate guard on the `postMessage` path. Tests first: 7 failed / 41 passed before, 48 / 48 after. The one new test that was green at baseline is labelled as a regression guard rather than evidence — plain prose was never the part that got cut.
…aders `githubJson` mapped every 403 to `access_denied`. GitHub does not give a rate limit a status of its own: an exhausted primary limit answers 403 *or* 429 with `x-ratelimit-remaining: 0`, and a secondary limit answers 403 or 429 and may add `retry-after`. https://docs.github.com/en/rest/using-the-rest-api/rate-limits-for-the-rest-api So a throttled `read_source`/`read_release` reached the model as a permission denial. That is the one failure the model cannot reason about correctly: a denial invites switching to another repository or path, while a throttle means every GitHub read in this investigation is spent and the remaining budget belongs on search evidence. The reply built on top of that is wrong for a reason the model cannot see. A 403 is now `rate_limited` when the headers prove a throttle, and `access_denied` otherwise. 429 is unchanged and still unconditionally `rate_limited`. Bounded deliberately: - Headers only, read positively. An absent, empty or unparsable header leaves a 403 a permission denial, so a malformed header can never manufacture a throttle. `x-ratelimit-remaining` must equal `0`; `retry-after` must parse as digits. A duplicated header (`Headers.get` joins to "0, 0") fails both and falls through to `access_denied`. - The response body is never inspected. It is untrusted and already excluded from the model and the trace; classifying on its text would both reintroduce that coupling and misread an ordinary denial whose message happens to mention a limit. - No sleeping, backoff, retry or transport framework. The tool still returns one recoverable result and the model still decides what to do next. - No new vocabulary. `rate_limited` was already in `GithubFailureReason` and already documented in docs/support-agent.md; only 403 now reaches it. Caller and pattern audit: - `githubJson` has exactly three callers — `read_source` (commits, contents) and `read_release`. All three forward a non-`not_found` failure to `unavailableEvidence`, unchanged. - The only branch on `reason` anywhere is `=== 'not_found'`. `access_denied` and `rate_limited` are pure pass-through, so no control flow changes: the fix alters only the word the model is told. - No other rate-limit, 403 or 429 handling exists in production source. The `apps/web` 403 checks are Next.js session auth, and `apps/web/src/lib/auth.ts` hits a different GitHub endpoint for org membership — neither shares this path. - Terminal behaviour untouched: `rethrowIfTerminal`, the investigation deadline, cancellation, the six-call budget and `errorFunction: null` propagation all unchanged. A failed call still spends one of the six. - docs/support-agent.md describes the reason vocabulary without asserting a status mapping, so it stays accurate. Tests: 7 added, table-driven over the same stub helpers. Primary-exhausted 403, secondary 403 with `retry-after`, 429 with and without headers, 403 with budget remaining, and 403 with unparsable and with empty headers. Each asserts the model receives a recoverable failure, then corrects the ref and still synthesizes an answer, and that no header value or body text reaches the model. Red before green: at parent 1239b8b the two header-signalled 403 rows fail (2 failed | 54 passed); the five guard rows pass, so the fix had to be narrow. Mutation sweep: 7 mutations, 0 survivors.
An evidence-auth failure gave the operator nothing to act on: the model-facing recoverable reason stayed auth_unavailable, so an operator could not tell a missing credential from a rejected one without reading raw credential state. Classify the failure into an allowlisted operator category and surface that category only. The diagnostic is built from a snapshot read once per call rather than once per branch, so a credential that changes mid-call cannot produce a self-contradictory category, and no raw credential value, token or key material reaches the message. The recoverable model-facing auth_unavailable reason is unchanged.
The web response sanitized validated code on its way out, so a reply whose code had already passed validation could reach the requester altered and no longer runnable. Both the single-string response and the summary pane took this path, and each serialized independently, so fixing one still left the other able to corrupt the same snippet. Preserve validated code verbatim in both validated fields and in both serializations. Sanitization now applies only to the unvalidated disclaimer, which is the untrusted prose it was meant for. The legacy formatter path and the existing completeText output are unchanged.
…knows
Two ways a reason the pipeline derived itself failed to reach the human who
picks up the escalation.
An explicit investigator diagnosis short-circuited the deterministic reason
instead of joining it. On the support-agent path `explicitHandoffReason ||
legacySuppressionReason` meant that whenever the model supplied its own
handoffReason, the groundedness and lint findings the pipeline had already
computed about that same draft were dropped. The two explain different things
— why the investigator gave up, versus what the gate independently caught in
what it produced — so neither substitutes for the other. They are now joined,
deterministic first, which is also what survives the 2000-character bound when
a diagnosis runs long. The legacy path already combined them; this makes the
two paths agree.
A forced escalation carried no reason at all. `groundedness.forcesEscalation`
clamps the score below the escalation gate, so a human is committed — but the
reason was gated on `suppressed`, which that outcome deliberately does not set
(the draft still publishes). The result reached the reviewer with the
escalation and no statement of which assertion to check.
Reasons are attached to those two deterministic outcomes only. A reply that
merely scored under the gate stays reason-free: its score is not a finding, and
restating it would bury the real ones. Suppression, published body, disclaimer
selection and confidence arithmetic are untouched.
The reason now crosses into a durable consumer.
Producing a reason on the forced-escalation outcome is only half of it. That
outcome publishes, so its handoff comes down the queue handler's
low-confidence arm, and `nonDeliveryEscalationReason` consulted `handoffReason`
on the suppressed arm alone. The newly-produced reason would have been dropped
at that seam — and dropped for good, because that expression is what the
durable `escalationRequiredReason` column is written from, before any
publication. Neither the owning job's retry via `recoveredHandoffReason` nor
the PENDING_RESPONSE_SWEEP backstop could recover a reason never written;
both read the stored row. So the handler now carries it too, and the rules
that govern it there match the ones above:
- Appended to the generic sentence, never substituted. Existing readers key
on the percentage — including the sweep, whose recovered text asserts the
stored reason leads it — so the numeric prefix stays first and literal.
- An absent, empty, or whitespace-only reason contributes nothing, so a
merely low-scoring reply reads exactly as it always has. Pinned by three
negative controls.
- Only the text changes, never the null-ness. The condition deciding who
queues a human is bit-for-bit identical, as are suppression, publication,
delivery, the compare-and-set, idempotency, phase faults and shadow mode.
No new routing, no new escalation, no new public export.
- The delivery-failure arm still replaces the owed reason with its more
actionable diagnostic, as documented. That precedence is deliberate and
unchanged.
- The 2000-character bound is restated at the handler seam. It duplicates
the producer's own slice on purpose: past this point the value lives in a
durable column and a queued job payload, neither of which should be able
to grow because a future producer skipped its bound.
- The reason is internal. It reaches the durable row and the queue payload,
and is asserted absent from the arguments the platform post is built from.
Handler tests are on the handler, not the pipeline: four were red against the
unchanged handler for exactly this reason, covering the seeded durable marker,
the actual ESCALATION job payload, a failed enqueue recovered on retry with
proven delivery, and the interrupted takeover that proves the reason survives
to a recovered escalation.
Review status
Marked ready for maintainer review at the user’s request. That is a request for review; it is not a claim that this is ready to merge or to enable automatic posting.
Fifteen broad review rounds preceded this branch. The eleven commits after the previously reviewed head c602d6c answer a later, separate round of review updates and are not the output of those earlier rounds. Six of them carry the earlier concerns from that round:
The current amendment round adds five further concerns, each with its own red/green record:
auth_unavailableresult the model sees unchanged.These have accepted source-and-test review, not completed review confirmation: a scoped twenty-role confirmation cluster is still running, and the broader original-PR confirmation and the original 130-case / 90-disposition promotion audit are unfinished. The two original-PR reason-propagation defects found in this review are resolved by the producer and durable-consumer changes in these commits; the broader audit remains open. No convergence claim is made.
CI is green at this head. CI run 37701387192 and security run 37701387244 passed at
b09050e. Local checks are recorded below. Review confirmation and the broader audit remain unfinished.Summary
Support replies could repeat the report, guess API compatibility, or bury the useful answer in a long response. This change introduces a bounded support investigator using GPT-5.6 Luna and the OpenAI Agents SDK. It validates structured evidence, runs a separate verifier, and renders a concise human paragraph with expandable technical details and sources on GitHub and in web QA.
Luna also handles confidence verification, ticket classification, and sentiment. The default path requires only
OPENAI_API_KEY. The existing queue, one-primary-reply-per-ticket behavior, delivery adapters, feedback calibration, and durable escalation workflow remain in place. Delivery fixes preserve expanded details in saved suggestions and shadow records, and report unresolved handoff state instead of silently accepting it.Runtime requirement: Node 24+ for local and container execution, aligned with CI and the OpenAI dependency stack.
Old vs. new responses
The table compares historical failure examples with actual saved output from the new pipeline, evaluated September 21, 2026, on build
d8c932f. It shows the public body; AI disclosure and attribution footers are omitted here for readability.Evidence scope: D is a verified excerpt from an original Outpost reply. A, B, C, and E describe reconstructed regression fixtures, not original transcripts. All five new outputs are real live runs. No replies were posted.
headersmap, while the transport requires those headers underrequestInit. The current main branch already contains the proposed mapping, so this appears fixed in source, but the published version containing it cannot be determined without the exact installed@copilotkit/runtimeversion.The earlier September 17–18 pipeline run remains archived separately; it is not relabeled as this build.
The original D response had 281 whitespace-delimited body words/tokens, including Markdown heading markers and excluding its opening disclosure and footers; it had no dropdown. The new handoff body has 12 words. That is a readability improvement, not a technical resolution.
See the complete new response format: Deep Agents subagents
The following is the saved formatted output. The inner disclosure is the actual technical-details dropdown.
Yes. Deep Agents supports subagents: the top-level agent can spawn specialized agents for delegated tasks, and CopilotKit’s integration can stream those subagent interactions to the frontend.
Technical details and sources
The Deep Agents documentation explicitly lists “spawn subagents for specialized tasks” as a capability. CopilotKit’s integration documentation also states that users can interact with subagents as they spawn. This support belongs to the Deep Agents backend; CopilotKit provides the frontend surface and AG-UI connection rather than defining the subagent execution model.
Applies to: Deep Agents integration; no specific package version or CopilotKit API generation was provided.
API version: unknown
Sources
🤖 Generated by CopilotKit AI Support · Documentation
See the complete new response format: MCP SSE authentication headers
The following is the saved formatted output. It cites the source retrieved during evaluation; it does not establish a minimum fixed release or an executed reproduction.
The report matches the current v2 source: the SSE config exposes a raw
headersmap, while the transport requires those headers underrequestInit. The current main branch already contains the proposed mapping, so this appears fixed in source, but the published version containing it cannot be determined without the exact installed@copilotkit/runtimeversion.Technical details and sources
The verified evidence supports the reported defect mechanism and correction:
MCPClientConfigSSEdocumentsheaders?: Record<string, string>.{ requestInit: { headers: serverConfig.headers } }, rather than passing the raw map as the second argument.For a checkout containing the current source, the relevant implementation is:
This finding is scoped to CopilotKit v2. The retrieved implementation is from
main; that proves the fix exists in source but does not prove which published package release includes it. Provide the exact@copilotkit/runtimeversion, including any prerelease tag, to verify whether upgrading resolves the issue for the affected installation.Applies to: CopilotKit v2 BuiltInAgent MCP SSE configuration; verified in current main source. Published-version applicability remains unresolved until the exact @copilotkit/runtime version is known.
API version: v2
Sources
🤖 Generated by CopilotKit AI Support · Documentation
Architecture and behavior
answer / partial / route, one summary paragraph up to 80 words, details, API generation, applicability, evidence quotes, and internal handoff reasonTicket classification preserves recognized critical incidents when the model fails or underestimates urgency. Regression tests distinguish incident reports from negation, prevention, and diagnostic questions.
The investigator and verifier receive available author/timestamp metadata. The thread tool reads supplied history; it does not fetch missing remote comments or claim the remote thread is complete. Version-aware search and evidence checks reduce definite v1/v2 mismatches. Source reads can pin commits, but search-result citations are not universally pinned. Main-branch code is not release evidence.
The style linter is connected in report mode by default;
AI_DRAFT_LINT_MODE=enforceenables its blocking rules. Exact quote matching checks provenance, not whether every conclusion follows. The verifier is a separate call using the same model family, so it is not independent ground truth.Validation and live evaluation
Validation of the current branch head
b09050e(tree7baae56), eleven commits and 17 files (+3,506/−207) after the previously reviewed headc602d6c:src, so it exercises source rather than the compiled artifact.discord-mcpanddocswere deliberately skipped — neither depends on it.--checkon all 16 changed Prettier-supported files. The seventeenth,schema.prisma, is outside the Prettier glob and was checked withprisma validateinstead, which passed against a placeholder connection string — static schema parsing, no database contacted. Its only change is a comment block, and the original migration is byte-unchanged, so the existing column is untouched.docs,discord-mcp) are outside it.c602d6c, with none removed. The new GitHub evidence-auth module is module-private — reached from the support agent, absent from the package index — and its only module specifiers are two Octokit packages already declared atc602d6c, so no dependency entered the runtime graph.The reasoning-continuity proof — the Agents SDK 0.18.0 converter reading and the two-response interaction showing encrypted reasoning carried across a stateless request — is a dated historical record from an earlier session, not a new live call on this head. The new App-installation authentication has been exercised only against synthetic credentials and offline SDK requests; live installation authentication is still unverified.
This amendment changes no Dockerfile and no runtime dependency, so no container build was run and no container runtime claim is made for it; the earlier
c602d6cimage proof is archived and does not carry forward. No provider, GitHub, or platform call was made during these checks, no service was deployed, and no support reply was posted. The local credentials and/qacheck remains tied to previewa4bf059, which is still the old build and does not contain any of these commits.The five-case output comparison remains the September 21 live run at
d8c932f. It is not a replay of this head, and tests and verifier scores do not establish real-world support accuracy.Regression coverage includes the SDK tool loop, malformed/segmented structured output, evidence validation, investigation budgets, provider failures, rendering, stream boundaries, and delivery/escalation invariants.
The five-case evaluation was rerun September 21, 2026, against a freshly emitted AI build of
d8c932f. The raw artifact records the commit and compiled-module hashes:This is not a controlled Sonnet-versus-Luna comparison. Historical failures predate prompt improvements already present in the old code baseline. A/B use reconstructed questions, C uses a synthetic resolved conversation, and D/E use real issue openers without later comments. Retrieval uses current evidence, not snapshots from the original reply dates. Model, prompts, tools, validation, and rendering changed together; these results do not isolate model quality or estimate production accuracy.
Rollout and remaining work
Keep posting disabled during further evaluation. This change does not deploy a service or enable external replies. Before enabling automatic posting:
Luna is the default here because that was an explicit user direction to route all AI calls through it. That direction is still being reconciled against the proposal to stay on Anthropic until a matched evaluation exists, so the default is a rollout decision rather than a settled benchmark result. For explicit rollback, set
AI_RESPONSE_PROVIDER=anthropic, configure the Anthropic key, and clear or replace OpenAI-specific model overrides. Provider failures never silently switch providers.The paragraph-plus-dropdown presentation follows the Orca response on #3434; that comment is a layout reference, not a correctness benchmark.