Skip to content

feat(ai): migrate support responses to Luna and OpenAI Agents SDK - #283

Open
jerelvelarde wants to merge 87 commits into
mainfrom
jerel/luna-support-responses
Open

jerelvelarde wants to merge 87 commits into
mainfrom
jerel/luna-support-responses

Conversation

@jerelvelarde

@jerelvelarde jerelvelarde commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Review status

Marked ready for maintainer review at the user’s request. That is a request for review; it is not a claim that this is ready to merge or to enable automatic posting.

Fifteen broad review rounds preceded this branch. The eleven commits after the previously reviewed head c602d6c answer a later, separate round of review updates and are not the output of those earlier rounds. Six of them carry the earlier concerns from that round:

  • Expected tool failures — directories, symlinks and submodules, invalid paths, and GitHub transport and API errors — return bounded tool results to the investigator instead of ending the run, so it can correct the lookup or fall back to other evidence. Cancellation and the investigation budget still terminate the run.
  • GitHub evidence reads authenticate with the existing App installation at read-only contents permission, bounded by the investigation deadline and a ten-second installation request deadline. One investigation can stop waiting without canceling another caller’s shared token request, and incomplete configuration cannot silently fall back to anonymous access.
  • An interrupted delivery keeps the owed handoff reason, waits for a possibly in-flight post, and warns the human when delivery is unknown. A successful post records confirmation even when the row stays pending for escalation.
  • The whole web response serializes with its footer last, keeping technical details and sources ahead of it in saved suggestions, shadow records and streaming output, while the native separate details pane is preserved.

The current amendment round adds five further concerns, each with its own red/green record:

  • The Discord footer is counted inside the message budget, with UTF-16 boundary and indentation handling corrected.
  • An exhausted GitHub rate limit is classified from its response headers rather than inferred.
  • Evidence-auth failure is reported as a safe operator diagnostic — an allowlisted category name, never a credential value — leaving the recoverable auth_unavailable result the model sees unchanged.
  • Validated code stays literal across both web serializations and in both validated fields; only the unvalidated disclaimer is sanitized, and the legacy formatter path is preserved.
  • Deterministic handoff reasons the pipeline already knows are preserved and carried across into a durable consumer, so a known reason reaches the human-facing ticket note, the system message and the recorded-delivery retry instead of stopping at pipeline metadata.

These have accepted source-and-test review, not completed review confirmation: a scoped twenty-role confirmation cluster is still running, and the broader original-PR confirmation and the original 130-case / 90-disposition promotion audit are unfinished. The two original-PR reason-propagation defects found in this review are resolved by the producer and durable-consumer changes in these commits; the broader audit remains open. No convergence claim is made.

CI is green at this head. CI run 37701387192 and security run 37701387244 passed at b09050e. Local checks are recorded below. Review confirmation and the broader audit remain unfinished.

Summary

Support replies could repeat the report, guess API compatibility, or bury the useful answer in a long response. This change introduces a bounded support investigator using GPT-5.6 Luna and the OpenAI Agents SDK. It validates structured evidence, runs a separate verifier, and renders a concise human paragraph with expandable technical details and sources on GitHub and in web QA.

Luna also handles confidence verification, ticket classification, and sentiment. The default path requires only OPENAI_API_KEY. The existing queue, one-primary-reply-per-ticket behavior, delivery adapters, feedback calibration, and durable escalation workflow remain in place. Delivery fixes preserve expanded details in saved suggestions and shadow records, and report unresolved handoff state instead of silently accepting it.

Runtime requirement: Node 24+ for local and container execution, aligned with CI and the OpenAI dependency stack.

Old vs. new responses

The table compares historical failure examples with actual saved output from the new pipeline, evaluated September 21, 2026, on build d8c932f. It shows the public body; AI disclosure and attribution footers are omitted here for readability.

Evidence scope: D is a verified excerpt from an original Outpost reply. A, B, C, and E describe reconstructed regression fixtures, not original transcripts. All five new outputs are real live runs. No replies were posted.

Case Previous response New pipeline's actual public body Assessment
A · Deep Agents subagents Reconstructed fixture: treated support as uncertain, suggested a speculative workaround, and sent the developer to a GitHub discussion. Yes. Deep Agents supports subagents: the top-level agent can spawn specialized agents for delegated tasks, and CopilotKit’s integration can stream those subagent interactions to the frontend. Direct, 26-word answer. Version scope and evidence are collapsed below it.
B · Rendering delegation Reconstructed fixture: mixed API generations and hedged an unverified render-hook name. This needs a maintainer review to give you a reliable next step. The draft failed the exact evidence-quote check. The handoff avoids publishing it, but gives no useful clarification.
C · Maintainer follow-up Reconstructed fixture: praised a maintainer and claimed it could not read other replies. Yes—the reporter confirmed that aligning the package versions restored the chat, so the rendering issue appears resolved. The exact framework, API generation, and package versions were not provided, so the underlying compatibility mismatch cannot be identified from the available thread. Reads the supplied resolution correctly, but adds little and discusses missing context. A no-reply policy is still needed.
D · Dependency audit, #6423 Verified original excerpt: “I haven't read the source files or run resolution myself, so I won't confirm causation or prescribe the specific fix” This needs a maintainer review to give you a reliable next step. Less noise, but no technical resolution. The investigator could not verify historical manifests and lockfile within its tool budget.
E · MCP SSE headers, #6927 Reconstructed fixture: praised the reproduction, explained its own limitations, and deferred to engineering. The report matches the current v2 source: the SSE config exposes a raw headers map, while the transport requires those headers under requestInit. The current main branch already contains the proposed mapping, so this appears fixed in source, but the published version containing it cannot be determined without the exact installed @copilotkit/runtime version. A concrete finding in 53 words, with installed-version uncertainty visible. Code and sources are collapsed.

The earlier September 17–18 pipeline run remains archived separately; it is not relabeled as this build.

The original D response had 281 whitespace-delimited body words/tokens, including Markdown heading markers and excluding its opening disclosure and footers; it had no dropdown. The new handoff body has 12 words. That is a readability improvement, not a technical resolution.

See the complete new response format: Deep Agents subagents

The following is the saved formatted output. The inner disclosure is the actual technical-details dropdown.

Yes. Deep Agents supports subagents: the top-level agent can spawn specialized agents for delegated tasks, and CopilotKit’s integration can stream those subagent interactions to the frontend.

Technical details and sources

The Deep Agents documentation explicitly lists “spawn subagents for specialized tasks” as a capability. CopilotKit’s integration documentation also states that users can interact with subagents as they spawn. This support belongs to the Deep Agents backend; CopilotKit provides the frontend surface and AG-UI connection rather than defining the subagent execution model.

Applies to: Deep Agents integration; no specific package version or CopilotKit API generation was provided.

API version: unknown

Sources


🤖 Generated by CopilotKit AI Support · Documentation

See the complete new response format: MCP SSE authentication headers

The following is the saved formatted output. It cites the source retrieved during evaluation; it does not establish a minimum fixed release or an executed reproduction.

The report matches the current v2 source: the SSE config exposes a raw headers map, while the transport requires those headers under requestInit. The current main branch already contains the proposed mapping, so this appears fixed in source, but the published version containing it cannot be determined without the exact installed @copilotkit/runtime version.

Technical details and sources

The verified evidence supports the reported defect mechanism and correction:

  • MCPClientConfigSSE documents headers?: Record<string, string>.
  • The current v2 runtime constructs SSE transport options as { requestInit: { headers: serverConfig.headers } }, rather than passing the raw map as the second argument.
  • The same source catches MCP connection failures, logs that the server is being skipped, and continues, which explains why the agent can appear without those tools.

For a checkout containing the current source, the relevant implementation is:

transport = new SSEClientTransport(
  new URL(serverConfig.url),
  serverConfig.headers
    ? { requestInit: { headers: serverConfig.headers } }
    : undefined,
);

This finding is scoped to CopilotKit v2. The retrieved implementation is from main; that proves the fix exists in source but does not prove which published package release includes it. Provide the exact @copilotkit/runtime version, including any prerelease tag, to verify whether upgrading resolves the issue for the affected installation.

Applies to: CopilotKit v2 BuiltInAgent MCP SSE configuration; verified in current main source. Published-version applicability remains unresolved until the exact @copilotkit/runtime version is known.

API version: v2

Sources


🤖 Generated by CopilotKit AI Support · Documentation

Architecture and behavior

Area Old baseline New implementation
Generation Fixed docs/code retrieval followed by one Claude Sonnet 4.6 call Luna investigation through the OpenAI Agents SDK and Responses API
Investigation No tools on the generator Read-only, adaptive CopilotKit/AG-UI docs/code search, supplied-thread reading, source-at-ref lookup, and tagged-release lookup
Bounds Single generation call Six tool calls, eight turns, 60-second investigation deadline; tools are disabled after the sixth call to leave room for synthesis
Response contract Free-form Markdown, mainly prompt-controlled answer / partial / route, one summary paragraph up to 80 words, details, API generation, applicability, evidence quotes, and internal handoff reason
Rendering Collapsed long code blocks on GitHub One GitHub technical-details dropdown; native disclosure in web QA; Discord splits to ordinary messages, and Slack and Teams still fall through to the legacy format
Verification Haiku judgment using truncated draft/source context Separate Luna judgment over the complete bounded draft/evidence and supplied conversation, followed by grounding checks
Publication controls Grounding suppression; low confidence did not generally withhold the draft Invalid output, unusable verification, intentional routing, or a verifier score below 0.40 withholds the technical draft and produces a concise handoff
Auxiliary models Haiku defaults for confidence, classification and sentiment Luna defaults throughout; strict structured results and preserved degraded states

Ticket classification preserves recognized critical incidents when the model fails or underestimates urgency. Regression tests distinguish incident reports from negation, prevention, and diagnostic questions.

The investigator and verifier receive available author/timestamp metadata. The thread tool reads supplied history; it does not fetch missing remote comments or claim the remote thread is complete. Version-aware search and evidence checks reduce definite v1/v2 mismatches. Source reads can pin commits, but search-result citations are not universally pinned. Main-branch code is not release evidence.

The style linter is connected in report mode by default; AI_DRAFT_LINT_MODE=enforce enables its blocking rules. Exact quote matching checks provenance, not whether every conclusion follows. The verifier is a separate call using the same model family, so it is not independent ground truth.

Validation and live evaluation

Validation of the current branch head b09050e (tree 7baae56), eleven commits and 17 files (+3,506/−207) after the previously reviewed head c602d6c:

  • 5,901 tests passed at this exact head: 4,599 core tests across 78 files, and 1,302 downstream tests across 94 files — 856 web, 84 worker, 73 Teams, 66 Discord, 64 root scripts, 60 GitHub App, 52 Slack, 47 Linear sync. The downstream suites were run fresh against the current build output. The core suite was measured twice at this same head and tree and reused rather than repeated; the web suite aliases core to src, so it exercises source rather than the compiled artifact.
  • Eleven cold typechecks passed: four core projects with their build info deleted first, plus each of the seven workspaces that depend on the core package. discord-mcp and docs were deliberately skipped — neither depends on it.
  • Lint exits clean with no errors and the same 40 pre-existing warnings across 27 files. The five warnings inside a changed test file sit at identical lines on both sides of the diff; nothing was repaired or suppressed.
  • Formatting passed --check on all 16 changed Prettier-supported files. The seventeenth, schema.prisma, is outside the Prettier glob and was checked with prisma validate instead, which passed against a placeholder connection string — static schema parsing, no database contacted. Its only change is a comment block, and the original migration is byte-unchanged, so the existing column is untouched.
  • The full production build passed: 10 tasks, 10 successful. All 8 packages and apps within the change's dependency reach were cache misses and executed fresh; the 2 cached tasks (docs, discord-mcp) are outside it.
  • Compiled modules were loaded and inspected after that build: 54 AI runtime exports and 21 queue exports. The public surface moves by exactly one addition versus c602d6c, with none removed. The new GitHub evidence-auth module is module-private — reached from the support agent, absent from the package index — and its only module specifiers are two Octokit packages already declared at c602d6c, so no dependency entered the runtime graph.

The reasoning-continuity proof — the Agents SDK 0.18.0 converter reading and the two-response interaction showing encrypted reasoning carried across a stateless request — is a dated historical record from an earlier session, not a new live call on this head. The new App-installation authentication has been exercised only against synthetic credentials and offline SDK requests; live installation authentication is still unverified.

This amendment changes no Dockerfile and no runtime dependency, so no container build was run and no container runtime claim is made for it; the earlier c602d6c image proof is archived and does not carry forward. No provider, GitHub, or platform call was made during these checks, no service was deployed, and no support reply was posted. The local credentials and /qa check remains tied to preview a4bf059, which is still the old build and does not contain any of these commits.

The five-case output comparison remains the September 21 live run at d8c932f. It is not a replay of this head, and tests and verifier scores do not establish real-world support accuracy.

Regression coverage includes the SDK tool loop, malformed/segmented structured output, evidence validation, investigation budgets, provider failures, rendering, stream boundaries, and delivery/escalation invariants.

The five-case evaluation was rerun September 21, 2026, against a freshly emitted AI build of d8c932f. The raw artifact records the commit and compiled-module hashes:

  • Five live Luna pipeline replays produced one answer, two partial answers, and two handoffs. Classification and sentiment returned valid, non-degraded results in all five; this establishes call/schema health, not semantic accuracy.
  • In an earlier, separate September 17–18 live test, the Luna verifier withheld 5/5 reconstructed historical bad replies, scoring them 0.05–0.18 against the 0.40 threshold.
  • Local replay times were 21.4–35.6 seconds. Total input ranged from 11,844 to 74,892 tokens, including cached tokens and discarded drafts. There is no matched old-system latency or cost measurement.

This is not a controlled Sonnet-versus-Luna comparison. Historical failures predate prompt improvements already present in the old code baseline. A/B use reconstructed questions, C uses a synthetic resolved conversation, and D/E use real issue openers without later comments. Retrieval uses current evidence, not snapshots from the original reply dates. Model, prompts, tools, validation, and rendering changed together; these results do not isolate model quality or estimate production accuracy.

Rollout and remaining work

Keep posting disabled during further evaluation. This change does not deploy a service or enable external replies. Before enabling automatic posting:

  • Give Discord a compact summary with an on-demand details interaction; details currently remain ordinary visible messages.
  • Carry validated-code rendering into Slack and Teams, which still fall through to the legacy format. The collapsed GitHub section and the web disclosure are the only two surfaces this change actually gives a distinct presentation.
  • Preserve useful clarification questions instead of replacing invalid or routed drafts with generic handoffs (B).
  • Add a no-reply decision for resolved conversations and maintainer-owned follow-ups (C).
  • Improve historical repository-layout and published-package evidence lookup (D).
  • Deduplicate retrieved context, improve citation pinning, and measure the token/coverage tradeoff.
  • Run matched old/new cases with frozen evidence and blind human review, grading added value, factual/version correctness, unnecessary handoffs, and appropriate silence.

Luna is the default here because that was an explicit user direction to route all AI calls through it. That direction is still being reconciled against the proposal to stay on Anthropic until a matched evaluation exists, so the default is a rollout decision rather than a settled benchmark result. For explicit rollback, set AI_RESPONSE_PROVIDER=anthropic, configure the Anthropic key, and clear or replace OpenAI-specific model overrides. Provider failures never silently switch providers.

The paragraph-plus-dropdown presentation follows the Orca response on #3434; that comment is a layout reference, not a correctness benchmark.

R1-DELIVERY-A1: join only the formatted summary and optional details when
saving suggestedResponse, so no-adapter web sources retain technical
details and source links. Keep private response/handoff metadata outside
this publishable field. Responses without details keep their existing text.

Affected symbol and callsites:
- handleAiResponse: registered by apps/worker/src/index.ts:85, imported at
  line 27; re-exported by packages/outpost/queue/src/index.ts:4.
- Test-only calls are in packages/outpost/queue/src/__tests__/ai-response.test.ts,
  imported at line 95; the new source-matrix call is at line 737.
- The single suggestedResponse assignment is in handleAiResponse; audit of
  the full handler and queue module found no other durable suggestion writer.
  The separate shadow SYSTEM diagnostic uses the formatted summary, and
  adapter delivery already receives the full formatted response object.

Validation:
- Baseline: 98 handler tests pass.
- Red: all five new no-adapter source cases fail (actual Summary only).
- Green: 103 handler tests; 74 package files / 1457 tests pass.
- Targeted Prettier and ESLint pass; queue tsc --noEmit and emitted build pass.
- git diff --check passes. Review convergence remains owned by the parent task.
Reject non-positive, non-finite, fractional, malformed, and unsafe integer
PATHFINDER_MAX_QUERY_CHARS values before direct Pathfinder clients can run.
Keep the default at 1000 and preserve positive safe-integer overrides.

AI-A01 regression: 10 new rejection cases failed before the fix because
config import succeeded (including maxQueryChars: NaN); all now pass.
Validation: 642 AI tests passed, including 70 config/Pathfinder tests;
targeted Prettier, ESLint, and AI TypeScript noEmit passed.

Affected-symbol audit: config.pathfinder.maxQueryChars is consumed only by
PathfinderClient.searchEvidence, private search, and searchDocs via
capQuery. All three now receive eagerly validated config. capQuery itself
and its explicit non-positive helper semantics are unchanged. Full config.ts
scan found no other numeric environment parsers with the same failure.
Require Node 24+ and use node:24-alpine in all seven app Dockerfiles,
matching the Node 24 CI runtime. The locked OpenAI SDK requires Node 22+
and all 585 locked Node engine declarations accept Node 24.0.0.

Runtime-reference audit covers the root manifest and README, all 21
Docker stages, deployment Markdown (including its base-image example),
and the deployment/getting-started HTML pages. No Node 20 support claims
remain outside locked third-party dependency metadata; the sample ticket
Node Version remains historical mock data. pnpm, Turbo, and healthchecks
are unchanged.

Verified JSON and all Docker FROM declarations, scoped Prettier checks,
HTML parse/format round-trip checks, dependency-engine compatibility,
and git diff --check. Container verification uses this committed tree.
Recognize GPT and known o1/o3/o4 model families in the shared provider
validation, including aliases and dated snapshots. Keep unknown custom
provider deployment names allowed, and preserve Claude rejection under
OpenAI.

Callsite audit:
- validateConfig routes response, confidence, classifier, and sentiment
  model settings through validateModelProvider before pipeline startup.
- AuxiliaryModel uses the same guard for explicit per-call overrides;
  ConfidenceScorer, TicketClassifier, and analyzeSentiment use this class.
- Provider-specific SupportAgent and ResponseGenerator remain unchanged;
  pipeline configuration validation guards their selected response model.
- No other duplicated provider-family checks exist in the AI package.

Regressions: all four config stages and direct auxiliary construction.
31 pre-fix failures become 91 passing focused cases; 169 related tests pass.
Targeted Prettier, ESLint, and AI no-emit typecheck pass.

Addresses AI-A02.
Addresses AI-A03. Distinguish explicit security vulnerabilities, data loss, and production outages from ordinary HIGH incidents. Preserve heuristic CRITICAL as an urgency floor without lowering model CRITICAL.

Affected callsites: TicketClassifier.heuristicClassify, TicketClassifier.classify (model merge and degraded fallback), and AIPipeline.classifyTicket through the shared classifier. The pipeline inherits the corrected fallback without a separate change.

Validation: regression cases failed in 14 assertions before the fix; all 61 classifier and auxiliary OpenAI tests now pass. Targeted Prettier, ESLint, and AI TypeScript noEmit checks pass.
Fix R1-DELIVERY-A2: read escalationRequiredReason alongside responseState
after a no-op escalation claim. DELIVERED with a retained marker now fails
for manual attention without claiming a handoff, clearing its reason, or
regenerating/reposting. Explicit null markers preserve clean success.

Apply the same invariant to the initial prior-response gate so retries
cannot erase the failure, and to pending-response sweep no-op outcomes.
The sweep still isolates failures per response and retains its scan policy.

Symbol callsites:
- reportSkippedRecoveryEscalation: recoverRequiredEscalation and
  recoverPendingResponse, ai-response.ts:312,360.
- handleAiResponse: queue/src/index.ts:4 exports it; worker/src/index.ts:27
  imports it and :85 registers AI_RESPONSE. Its handler regression suite
  exercises the initial gate and both no-op recovery entry points.
- settleStrandedResponse: handlePendingResponseSweep at
  pending-response-sweep.ts:231; the public handler is exported by
  queue/src/index.ts:15 and registered by worker/src/index.ts:94.
- stageRow test fixture: recovery no-op describe block in
  queue/src/__tests__/ai-response.test.ts, including the initial-gate tests.
- settleAfterRead test fixture: pending-response-sweep.test.ts:367,379,398
  covers concurrent ESCALATED, conflicting DELIVERED, and clean DELIVERED.

Validation: red/green demonstrated for both no-op recovery routes, the
initial-gate retry, and sweep no-op; 130 tests pass across both affected
suites. Prettier and queue tsc pass. ESLint has zero errors and five
pre-existing warnings in untouched sweep fake methods (baseline six).
Forward nonempty query.version from the shared search builder and the separate searchDocs builder, matching searchEvidence. Keep unversioned arguments and existing error handling unchanged.

Callsite audit: searchCode, searchAgUiCode, and searchAgUiDocs delegate to the shared builder; searchDocs builds its own arguments; searchEvidence already forwards version. Pipeline searchDocs/searchCode callers omit version and retain their behavior. Support-agent searchEvidence callers are unchanged. exploreDocs, queryKnowledgeBase, and fallbackSearch do not accept PathfinderQuery.

Regression evidence: serialized JSON-RPC assertions fail for v1/v2 across all four wrappers before the fix (8 failures; omitted-version cases pass). After the fix, 55 Pathfinder tests pass. Prettier, targeted ESLint, AI typecheck, and AI build pass.
Derive labels from the final rounded score using the documented thresholds.
Use the existing 25/NEUTRAL empty-input contract for degraded fallbacks and
empty trend periods, while retaining degraded flags and billed token usage.

Callsite audit: account-scoring consumes the label for its Prisma mapping
and still skips sentiment writes whenever degraded is true; no handler
changes were needed. getSentimentTrend forwards analyzed scores and labels
and now emits a consistent neutral pair for empty periods, which remain
excluded from trend deltas. Both APIs keep their existing exports and shape.

Validation: genuine red-green regression (19 expected failures before the
fix, 59 passing tests after) across sentiment, OpenAI auxiliary transport,
sentiment trend, and account-scoring. Changed-file Prettier, ESLint with
zero warnings, AI tsc --noEmit --incremental false, and git diff --check pass.
Follow up AI-A03: avoid CRITICAL escalation for prevention questions and negated incident mentions. Inspect nearby phrase context within individual sentences so a separate affirmative incident still escalates. Preserve the original data-loss HIGH baseline and all model priority floors.

Affected callsites: TicketClassifier.heuristicClassify and TicketClassifier.classify; AIPipeline.classifyTicket inherits the guarded heuristic without its own change. Guards intentionally cover common phrasing rather than full natural-language inference.

Validation: 11 new assertions fail before the follow-up; all 77 classifier and auxiliary transport tests pass afterward. Targeted formatting, lint, AI noEmit typecheck, and diff checks pass.
Apply the requested API generation filter before deciding whether search_evidence needs its unfiltered fallback. Use the same case-insensitive v1-deprecated marker as final reply validation, and filter fallback hits before remembering them. Mixed usable results retain requested_version scope.

Callsite audit: both support-agent searchEvidence calls share this filter across CopilotKit/AG-UI docs/code. Pathfinder strict parsing/version forwarding and final reply validation remain intact. The single spend(), shared abort signal, request/source limits, v1/unknown semantics, scope labels, and thrown retrieval failures are preserved; no other production searchEvidence callers exist.

Validation: genuine SDK/aimock red-green roundtrips cover deprecated-only and uppercase-title first hits, useful fallback evidence, mixed results without fallback, v1/unknown retention, and fallback error propagation. Full AI suite: 27 files, 744 tests passed. Prettier, ESLint, AI typecheck, and isolated AI emit build passed. Temporary dependency/declaration links removed before commit.
Flatten every publishable Discord part in the support-agent string stream, without duplicating the first part. Append separate details only for web output and retain the publication gate.

Add an aimock regression through the real SupportAgent, reply validation, and formatter. The pre-fix stream loses steps 12-24, citations, and the footer; short replies on all platforms and routed-draft confidentiality are also covered.

Callsite audit: generateStreamingResponse has only test callers in this repo. The Discord adapter already consumes every formatted part; the web QA route consumes separate details metadata. The queue saved-suggestion continuation issue is separate from this stream API. Legacy streamed chunks remain unchanged.

Validation: 770 AI tests pass, including 125 targeted pipeline/formatter tests; Prettier and ESLint pass for both changed files; AI tsc noEmit and isolated production build pass on Node 24.
Return a failed AI_RESPONSE result when both the DELIVERED write and
its deliveryConfirmed fallback fail after successful delivery. Include
both errors for manual attention and stop before progress 100 while
preserving pipeline teardown and the one-post response claim.

R1-DELIVERY-A4: the regression failed against a9db1d8 with success true,
then passed with the fix. The fallback-marker-success control remains
green. A retry schedules recovery without regenerating or reposting.

Callsite audit: read the complete ai-response handler, its worker
registration and JobResult failure consumer, and the pending-response
sweep delivery marker consumer. Existing primary-response recovery
prevents repeat posts; successful fallback markers still repair state.
No owed-escalation reason persistence changes (R1-DELIVERY-A3).

Validation: 108 handler tests and 256 queue tests pass; scoped Prettier
and ESLint, queue tsc --noEmit, queue emitted build, and diff check pass.
R1-DELIVERY-A3: insert escalationRequiredReason with the primary response,
so failed later writes or an interrupted enqueue cannot replace the exact
low-confidence/suppression reason with generic delivery recovery. An insert
failure prevents publication. Keep a failed platform post's reason ahead
of suppression by replacing its marker and responseError together; report
both write and queue failures when that replacement is not durable.

Symbols and callsites:
- handleAiResponse (ai-response.ts:480) is exported by queue/src/index.ts:4
  and registered by apps/worker/src/index.ts:85. It computes the owed reason
  before its primary message insert, then publishes at most once.
- recoverRequiredEscalation (ai-response.ts:294), called by the prior-response
  gate at :569, preserves deliveryFailed metadata for replaced reasons.
- enqueueEscalationAtomically is unchanged; its callers at ai-response.ts:306,
  :354, :1049 and pending-response-sweep.ts:142 consume the durable reason.
- trackResponsePersistence (ai-response.test.ts:264) provides real write-derived
  retry/sweep snapshots with rollback for the new failure-path regressions.

Validation: genuine red/green retry and sweep regressions; 137 tests pass;
changed-file Prettier and ESLint pass; queue no-emit typecheck and isolated
queue emit build pass on Node 24. A4 delivery-confirmation failure handling
is unchanged. Parent owns the CR loop; no push is part of this fix.
R1-DELIVERY-A3 follow-up: a concurrent retry can enqueue the persisted
handoff while the original postResponse is pending. Guard the later
delivery-reason replacement by PRIMARY_AI_RESPONSE and PENDING so it
cannot recreate an owed marker on ESCALATED or DELIVERED. When the guard
matches no row, retain responseError without changing lifecycle metadata.

Symbols and callsites:
- handleAiResponse (ai-response.ts:480), exported at queue/src/index.ts:4
  and registered at apps/worker/src/index.ts:85, now conditions the marker
  replacement on the still-pending primary response.
- recoverRequiredEscalation (ai-response.ts:294), called by the gate at
  :569, and enqueueEscalationAtomically (:120; calls :306, :354, :1064 and
  pending-response-sweep.ts:142) retain their existing transition behavior.
- holdPlatformPost (ai-response.test.ts:341) and trackResponsePersistence
  (:264) drive the real overlapping retry regression. The analogous
  DELIVERED/interrupted-owner case proves no transient marker is stranded.

Validation: both race cases fail before the guard and pass after it;
139 handler/sweep tests pass. Changed-file Prettier and ESLint, queue
typecheck, and isolated queue emit build pass. A4 remains unchanged.
Root owns CR and integration; no push or main-checkout mutation.
Follow up R1-DELIVERY-A1 by preferring nonempty formatted.parts over the
aliased first-part text when saving the complete publishable suggestion.
Join parts with paragraph separators, append optional details, and fall
back to text when parts is absent or empty. Never append private draft or
handoff metadata.

Affected symbol and callsites:
- handleAiResponse: imported by apps/worker/src/index.ts:27 and registered
  as the AI_RESPONSE handler at line 85; re-exported by
  packages/outpost/queue/src/index.ts:4.
- Test calls remain in packages/outpost/queue/src/__tests__/ai-response.test.ts
  (import at line 95; new multipart/fallback call at line 741).
- Full handler and queue-module audit found one suggestedResponse writer.
  Adapters already receive the full formatted object; shadow diagnostics
  remain separate. The analogous streaming pipeline remains unchanged.

Validation: the multipart regression genuinely failed with Summary only
before the fix; afterward all 105 handler tests pass, including existing
web/no-details coverage and the empty-parts fallback. Targeted Prettier,
ESLint, queue typecheck, and queue emitted build pass. Private fields and
duplicate first parts are excluded by the exact expected suggestion.
Force OpenAI and Luna in the shared auxiliary call options. Override the
pipeline test config's provider and all four model fields, then validate
the same config object consumed by the default pipeline clients.

Test callsite audit: every ConfidenceScorer, TicketClassifier, and
analyzeSentiment call in auxiliary-openai.test.ts uses the pinned options.
All three AIPipeline construction sites use the file-local config mock;
the injected SupportAgent already defaults to Luna on OpenAI Responses.
No production behavior or process-wide provider/model env mutation changed.

AI-A10 red/green: Anthropic provider plus Claude model overrides produced
23 failures before the fix and 46 passes afterward. Both files also pass
with default env and OpenAI provider plus Claude model overrides. Full AI
suite: 776 tests pass. Targeted Prettier/ESLint and AI typecheck pass.
R2-AI-A02 keeps legacy Pathfinder parsing tolerant while making the strict searchEvidence boundary fail when a code result cannot be cited.

Call sites audited: PathfinderClient.searchEvidence production use is via SupportAgent in packages/outpost/ai/src/support-agent.ts for the initial docs lookup and dynamic follow-up tool calls; direct strict-boundary tests live in packages/outpost/ai/src/pathfinder-evidence.test.ts. SearchResult kind/sourceUrl consumers in support-agent.ts and support-reply.ts were checked; no exported shape changed.

Checks: red pathfinder-evidence regression failed before implementation; green pathfinder-evidence and pathfinder parser tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-DELIVERY-A1 keeps the SHADOW_MODE SYSTEM row aligned with the complete publishable response stored as suggestedResponse.

Changed-symbol callsites: added private completePublishableResponse in packages/outpost/queue/src/handlers/ai-response.ts; it is called once in handleAiResponse to derive publishableResponse, which is then used by the suggestedResponse ticket update and the outpost-shadow SYSTEM message create. Added findShadowMessageCreateCall test helper used by the existing shadow assertion plus the new web-details and multipart Discord shadow regressions.
R2-AI-A05 prevents a caller-cancelled MCP setup from leaving the initialized session id cached for the next strict evidence request. The cancellation cleanup is limited to the caller AbortSignal path during notifications/initialized; non-cancellation notification failures remain best-effort.

Call sites audited: connect is used directly by pathfinder tests and internally by callTool; production strict evidence flows through SupportAgent's Pick<PathfinderClient, 'searchEvidence'> dependency in support-agent.ts. reset remains private except disconnect, and no exported type or method shape changed.

Checks: red pathfinder-evidence sequential cancellation regression failed before implementation; green pathfinder-evidence and pathfinder parser tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-AI-A04: recognize did/has/was/were/have/had/will question starters in
the existing sentence-local CRITICAL guard. Keep the old HIGH baseline,
healthy model CRITICAL, affirmative incidents and mixed sentences intact.

Add 21 contrastive heuristic rows and real Responses LOW/MEDIUM and
CRITICAL coverage. Genuine red: 20 cases initially, plus 15 additional
same-pattern cases. Green: 120 classifier/auxiliary tests; scoped Prettier
and ESLint, inclusive AI typecheck and isolated AI build all pass.

Changed-symbol call-site audit: TicketClassifier.heuristicClassify is called
by TicketClassifier.classify (classifier.ts:54), AIPipeline.classifyTicket
fallback (pipeline.ts:432), and classifier.test.ts. The model-backed caller
is pipeline.ts:426, constructed at pipeline.ts:110. Tests exercising the
class also include auxiliary-openai.test.ts; pipeline.test.ts and
pipeline-groundedness.test.ts retain their injected classifier mocks.
No signature, exports, tag/category heuristics or priority merge changed.
R2-AI-A01 separates internal handoff diagnostics from public support prose validation while preserving strict public summary/details/appliesTo checks and the nonempty route reason rule.

Changed-symbol callsites: validateSupportReply still parses supportReplySchema, checks route handoffReason nonempty, validates evidence, and now calls validateProse only for public summary/details/appliesTo. supportReplyText and supportReplyDetails were not changed; tests assert private handoffReason does not appear in either renderer. No queue, provider, or formatter callsites changed.
R2-AI-A06 fixes code citation URL construction for GitHub repository URLs that end in .git/. The normalization now removes a trailing slash before stripping a trailing .git suffix, preserving the existing provider allowlist, branch, ref, and source path behavior.

Call sites audited: blobUrl is private and called only by parseSnippets for code snippets. Its sourceUrl output flows through tolerant searchCode/searchAgUiCode, strict searchEvidence, SupportAgent source filtering/deduplication, and support-reply citation validation. No exported interface shape changed.

Checks: red pathfinder parser and strict evidence suffix regressions failed before implementation; green pathfinder and pathfinder-evidence tests passed; scoped Prettier, scoped ESLint, AI tsc --noEmit, AI tsc build with /tmp output, and git diff --check passed.
R2-AI-A03: keep the message from known validation and investigation-budget
errors for PipelineResult.handoffReason. Retain the existing 2000-character
return bound, generic public handoff and empty-diagnosis fallback. Other
provider/transport failures still propagate without replacement.

Tests first: eight genuine failures for diagnosis loss, then 88 passing
pipeline tests covering both classes, GitHub/Discord/web public and stream
privacy, cause exclusion, low-confidence suppression and verifier skipping.
Scoped Prettier/ESLint, inclusive AI typecheck and isolated AI build pass.

Changed-symbol call-site audit: AIPipeline.generateSupportResponse is called
by queue/handlers/ai-response.ts:735, apps/web/src/app/api/qa/route.ts:65,
and generateStreamingResponse at ai/src/pipeline.ts:466, plus owning tests.
Private handoffReason is consumed by queue escalation/shadow diagnostics
at ai-response.ts:786,875; web/streaming select formatted output. No signature,
export, queue state machine or provider error conversion was changed.
R3-AI-A02 requires strict searchEvidence calls for search-code and search-ag-ui-code to reject legacy JSON-array results that cannot be cited. The legacy parser can still return URL-less entries, but the strict boundary now also uses the requested tool name when enforcing code citation provenance.

Changed symbol call sites audited: PathfinderClient.searchEvidence is called by SupportAgent primary retrieval in packages/outpost/ai/src/support-agent.ts, by SupportAgent follow-up tool retrieval in packages/outpost/ai/src/support-agent.ts, and by strict evidence tests in packages/outpost/ai/src/pathfinder-evidence.test.ts. Public tolerant wrappers searchCode/searchDocs/searchAgUiCode/searchAgUiDocs still route through search/parseSearchResults unchanged, with a regression control for URL-less legacy searchCode JSON.

Validation: red pathfinder-evidence regression failed at 233baee with both strict code tools resolving URL-less legacy JSON; green pathfinder-evidence passed 28 tests; related pathfinder parser/wrapper test passed 41 tests; scoped Prettier, scoped ESLint, explicit AI typecheck, explicit AI build, and git diff --check passed.
R3-AI-A03 encodes read_source blob citation paths with the same per-segment encoding used for GitHub contents requests.

Changed symbol/callsites:

- Added encodeSourcePath(path) in packages/outpost/ai/src/support-agent.ts.

- Updated read_source githubJson contents callsite to use encodedPath.

- Updated read_source remembered sourceUrl callsite to use encodedPath in the GitHub blob URL.

- Added support-agent regression rows for docs/My Guide.md and docs/100% ready (setup).md.

- Added support-reply validation regression proving encoded blob citations validate while raw whitespace URLs remain rejected.
R3-AI-A05 extends the sentence-local critical incident suffix guard so passive absence reports such as security vulnerabilities were not found and production outage was not reported do not create a CRITICAL floor.

Call-site enumeration for the changed classifier surface: TicketClassifier.heuristicClassify is called by TicketClassifier.classify in packages/outpost/ai/src/classifier.ts; AIPipeline.classifyTicket uses the classifier fallback in packages/outpost/ai/src/pipeline.ts; AIPipeline constructs TicketClassifier in packages/outpost/ai/src/pipeline.ts; queue AI response handling reaches it through pipeline.classifyTicket in packages/outpost/queue/src/handlers/ai-response.ts; classifier, auxiliary-openai, pipeline, and queue tests cover the concrete and mocked callers.

Tests: node ../../node_modules/vitest/vitest.mjs run ai/src/classifier.test.ts --config vitest.config.ts --no-cache --reporter=dot; node node_modules/prettier/bin/prettier.cjs --check packages/outpost/ai/src/classifier.ts packages/outpost/ai/src/classifier.test.ts; node node_modules/eslint/bin/eslint.js packages/outpost/db/src packages/outpost/ai/src packages/outpost/queue/src packages/outpost/shared/src; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.json --noEmit --incremental false; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.build.json --outDir /tmp/outpost-r3-passive-ai-build-1789930006 --incremental false
Follow-up for R3-AI-A05: evaluate every critical incident mention in a sentence so a first passive absence mention cannot hide a later affirmative mention of the same phrase.

Regression rows cover: Security vulnerabilities were not found in staging, but security vulnerabilities were found in production; Data loss was not found in staging, but data loss was found in production. Red log showed both returned HIGH before the fix; green log shows 84/84 classifier tests passing after the scanner evaluates all mentions.

Call-site context remains the same as the prior A05 commit: TicketClassifier.heuristicClassify is called by TicketClassifier.classify; AIPipeline.classifyTicket uses classifier fallback logic; queue AI response handling reaches it through pipeline.classifyTicket. No public symbols or signatures changed.

Tests: node ../../node_modules/vitest/vitest.mjs run ai/src/classifier.test.ts --config vitest.config.ts --no-cache --reporter=dot; node node_modules/prettier/bin/prettier.cjs --check packages/outpost/ai/src/classifier.ts packages/outpost/ai/src/classifier.test.ts; node node_modules/eslint/bin/eslint.js packages/outpost/db/src packages/outpost/ai/src packages/outpost/queue/src packages/outpost/shared/src; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.json --noEmit --incremental false; node node_modules/typescript/bin/tsc --project packages/outpost/ai/tsconfig.build.json --outDir /tmp/outpost-r3-passive-followup-ai-build-final-1789930263 --incremental false
R3-AI-A06: suppressed legacy responses now put deterministic groundedness/enforced-lint reasons before generic legacy generator reasoning, while explicit OpenAI support-agent route/error diagnostics still take precedence.

Changed-symbol call sites: AIPipeline.generateSupportResponse is called by the queue AI response handler, web QA route, and generateStreamingResponse; PipelineResult.handoffReason is consumed by queue escalation persistence and by pipeline tests; public web/streaming paths publish formatted output. lintDraft, assessGroundedness, SUPPRESSED_RESPONSE_TEXT, and exported type shapes keep their existing call sites and signatures.

Tests: red pipeline/pipeline-openai suite failed 2/91 as expected before the fix; final related Vitest suite passed 110/110; scoped Prettier, ESLint, AI typecheck, AI build, and git diff --check passed. Logs and round-3-fix-R3-AI-A06.md saved in the CR ledger.
A support answer puts its code inside a numbered step or a quoted example.
proseOutsideFences recognized a fence only at columns 0-3, so every one of
those was scanned as prose and the reply discarded:

  - Example:                          -> Raw HTML is only allowed inside code
    ```tsx                               in a support reply
    <CopilotKit runtimeUrl="…" />
    ```

  > ```tsx                            -> same
  > <CopilotKit runtimeUrl="…" />
  > ```

  10. Example:                        -> Support reply link URL must belong to
      ```text                            validated source evidence
      https://example.invalid/…
      ```

The renderer publishes all three as <pre><code>: inert code that mounts no
element and resolves no link. A correct answer was thrown away and the ticket
escalated. The docstring already claimed this contract; the column test was a
different question.

A fence opens wherever its container's content starts, and four columns past
that is an indented code block with no fence at all — block structure, not a
line pattern. proseOutsideFences now asks the parser the chat renderer runs
(mdast-util-from-markdown + micromark-extension-gfm, already imported here for
the destination and definition scans) which nodes are code, and blanks their
lines. That removes the column from the question for every container nesting
and for indented code at once, and it closes the opening and closing match
together — the same regex carried the assumption on both, so fixing only the
opener would have left a container fence open and failed the reply with
"Unclosed code fence in support reply" instead.

Both refusals in this function are kept, not inherited by the new shapes:

- "Invalid code fence" still fires for a backtick fence whose info string holds
  a backtick — CommonMark says that is not a fence, and refusing rather than
  guessing is deliberate (R13-AI-A03 owns whether it should be). It now skips
  lines the parser places inside code, which the column scan could not.
- "Unclosed code fence" exists because supportReplyDetails appends the
  applicability, version and sources footer to `details` and an open fence
  swallows it. Only a top-level fence can: a blank line closes a block
  container before anything inside it. Asked directly, by parsing the field
  with a probe paragraph appended, so the refusal follows the reason instead of
  hand-tracked fence state. All four top-level spellings still reject; a
  container's fence does not.

Verification: the three shapes above flip to accepted in
round13-code-literals-supplied-claim-probe; its A03 and A04 rows do not move,
and neither does any row of round13-gfm-supplied-claim-probe,
round13-ai-residual-claim-probe or -probe2 measured against a build of
1252e2b. Regressions are red against the unchanged implementation (17 failed /
241 passed) and green after (258 passed): 14 accept rows across list items,
block quotes, nesting, tilde fences and indented code; 8 reject rows proving a
container carrying a paragraph is still validated and so is anything after its
fence closes; the open-fence rows in both directions. The existing block
container row in "keeps prose that resolves to no reference definition" was
spelled //example.invalid/steal, which the raw-URL scan cannot match — it now
carries https:// so it fails if the fence is read as prose. The renderer half
is pinned in apps/web with literal fixtures through the app's real ChatMessage
-> ReactMarkdown + remark-gfm: each shape publishes one <pre><code>, no anchor
(not even for the bare address and www. host GFM linkifies elsewhere), no
mounted element, and the appended footer survives a container's fence while a
top-level one swallows it.

Call sites: proseOutsideFences is module-private with one caller, validateProse
(support-reply.ts:327), which validateSupportReply (:408, applying it at :466) applies to summary,
details and appliesTo — so all three fields are covered, and must be, since
escapeMarkdown escapes neither backticks' block role nor container markers for
appliesTo. validateSupportReply reaches production through support-agent.ts:347
and the ai barrel (index.ts:11). publishedDestinations is unchanged in
behaviour; it now shares the new parseMarkdown helper. supportReplyDetails,
whose footer the unclosed-fence refusal protects, is consumed by formatter.ts:47.
A run of three or more backticks that opens and closes on one line is a
code span. The renderer publishes `<p><code>literal code</code></p>` and
keeps the body inert: the JSX arrives as text, and neither a bare address
nor a `www.` host GFM linkifies in prose becomes a link. The validator
refused the same bytes with "Invalid code fence in support reply",
discarding a valid reply, because any line opening with three backticks
whose remainder held another backtick was rejected outright. Three
backticks is exactly how an answer quotes a literal that already holds
one — up to and including a fence, ```` ```tsx ````.

Ask the parser instead of the line. The walk that already collects code
blocks now also collects each code span's opening position, and the
refusal fires only where no span starts at the run. Its purpose is
unchanged and still covered: a run left open is not a span, and a fence
whose info string holds a backtick is not a fence, so the renderer
commits to neither and publishes the marker as literal paragraph text
with every following line read as prose. Refusing rather than guessing
there is deliberate.

A span does not exempt the rest of its line: `proseOutsideInlineCode`
already strips runs of any length, which is why a mid-line triple span
was accepted before this change and a line-initial one was not — the
guard ran first. Those two are now one code path.

Call sites — all inside packages/outpost/ai/src/support-reply.ts:
  collectCodeBlocks -> collectCodeNodes (+ CodeNodes): sole caller
    proseOutsideFences:151; no other reference in the repo.
  proseOutsideFences: sole caller validateProse:344.
  proseOutsideInlineCode: sole caller validateProse:356.
  validateProse: sole caller validateSupportReply:483.
  validateSupportReply is the only exported symbol reached, via
    ai/src/index.ts:11 and ai/src/support-agent.ts:347. No signature
    changed, so no call site needed an edit.

Renderer behavior pinned in apps/web/src/__tests__/qa-components.test.tsx
against the app's real ReactMarkdown + remark-gfm: nine spans published
inline and inert, three open runs published as literal text.

(cherry picked from commit df2c06b12c950cfc9560eb2655c490fabb7112e6)
`appliesTo: 'Runtimes on <v2 releases'` was refused as raw HTML. The renderer
publishes `<p>Runtimes on &lt;v2 releases</p>` — a literal '<', markup-free.
`appliesTo` is the field the schema dedicates to version applicability, so a
'<vN' range is that field's own vocabulary, and the guard's own boundary sat
between two spellings of one sentence: `<1.9` passed because a digit is not a
tag name, `<v1.9` did not.

A '<' opens a tag only where the grammar can close one. `publishesRawHtml`
asks the parser the chat renderer runs — `mdast-util-from-markdown` plus the
GFM extension, already used for code blocks, definitions and destinations —
whether any `html` node resolves, and refuses only then. The comment,
processing-instruction, declaration, CDATA and closing-tag forms the pattern
covered are `html` nodes too, so all of them stay refused; so does a '<vN'
opening an unquoted attribute value closes into a real tag.

The autolink pre-strip goes with the pattern it existed for: an autolink is a
`link` to the grammar, never `html`, and the destination checks above still
hold it to the evidence set.

Same assumption, second site: `proseOutsideInlineCode` skipped from '<' to the
end of the line whenever no '>' followed, so every code span after an inert
'<vN' range went unmasked and its contents — example URLs included — were read
as prose. No '>' closes no tag, so there is nothing to skip.

The guard runs over the masked prose, not the original text: the masking is
what implements "only inside code", and it stays deliberately stricter than
the grammar for a span crossing a line. That control is unchanged.

Call sites:
- publishesRawHtml (new) — support-reply.ts:446, the only caller, inside
  validateProse; collectHtml is called only by it and itself.
- proseOutsideInlineCode — support-reply.ts:378, the only caller.
- validateProse — support-reply.ts:510, the only caller, applied to summary,
  details and appliesTo alike, so one change covers all three fields.
- parseMarkdown — now five callers (:136, :156, :240, :265), all read-only.

(cherry picked from commit ebfd0ab8358693e21f5e6bf624f787a951d372e0)
The inline destination scan matched every literal `](` in prose, with no
requirement of a preceding link label or of a CommonMark-legal destination,
so `The literal punctuation ](not a link) is part of this sentence.` was
read as a citation of `not` and the whole reply was discarded. The renderer
publishes no anchor for that sentence at all.

Ask the grammar which links exist instead. inlineDestinations parses the
prose with the parser and GFM extension the chat surface runs, walks link
and image nodes, and recovers each destination's own offset from the node's
range by counting the label's matching bracket. The check and the mask that
follow are unchanged, so a destination is still read as the reply spells it
rather than as the parser decodes it, escaped and angle-bracket forms still
resolve, and relative, protocol-relative and non-HTTP destinations of real
links are still held to the evidence.

A real link that looks like punctuation keeps being checked: this renderer
publishes `arr[i](x)` as an anchor, and the node walk sees it where a
label-shaped pattern would not.

Second site, same assumption. `proseOutsideInlineCode` decides which
same-line code spans are masked from the raw-URL and raw-HTML scans, and it
skipped a destination-shaped run after every literal `](` because a backtick
inside a real destination is a literal rather than a code delimiter. So
`See ](`https://example.invalid/x`) here.` — which the renderer prints as
inert code and publishes no anchor for — had its code span stepped over and
the address in it read as an invented citation. `inlineDestination` now
reports the offset of the `]` alongside the destination's own start, and
`inlineDestinationColumns` turns those into the per-line columns the scan
consults. Where a label really did open the destination the skip is
unchanged, so `[guide](`https://…`)` stays a link with a backticked
destination and stays rejected.

Renderer-pinned in the qa-components fixture table: the two spellings differ
only in the label, and only the labelled one emits an anchor.

Regression evidence: 6 failures against the unchanged implementation for the
literal-punctuation rows, including one alongside a genuine evidence link.
Seven must-reject rows (subscript, image, image nested in a link label,
nested/escaped/multiline labels, second invented link) and three masking
rows passed before and after.

Call sites — all inside packages/outpost/ai/src/support-reply.ts:
  inlineDestination:278 — sole caller inlineDestinations:330.
  collectInlineLinks:302 — sole caller inlineDestinations:322, plus itself.
  inlineDestinations:320 — callers inlineDestinationColumns:350 and
    validateProse:533.
  inlineDestinationColumns:342 — sole caller validateProse:495.
  proseOutsideInlineCode:425 — sole caller validateProse:501; the added
    parameter is the only signature change and has no other reference.
  validateProse:484 — sole caller validateSupportReply:635, applied to
    summary, details and appliesTo alike.
  validateSupportReply is the only exported symbol reached, via
    ai/src/index.ts and ai/src/support-agent.ts. Its signature is unchanged,
    so no call site outside this file needed an edit.

Integrates R13 A04: cherry-picked from cd8f81bc1712aab845b4f683a55e833730cafbe1
and its corrective 8c9ef62eb3e8f3d2f8d5559009da5c4bd0e6c521, squashed into
one concern.
An inline destination was read as the reply spelled it; a reference
definition arrived from the parser already decoded. So one published
href had two spellings on opposite sides of the evidence check:

  See [docs](https://docs.copilotkit.ai/search?a=1&amp;b=2).   rejected
  See [docs][d]                                               accepted
  [d]: https://docs.copilotkit.ai/search?a=1&amp;b=2

Both publish href="…/search?a=1&b=2" — the cited evidence — and the
first was discarded with "Support reply link URL must belong to
validated source evidence". Recorded through the app's real
ReactMarkdown + remark-gfm and through this parser (logs 00, 01).

`inlineDestinations` now returns the destination the link or image node
already carries, decoded, alongside the span the raw-URL scans mask.
`checkUrl` compares that value, so backslash escapes and HTML character
references resolve the same way here as in a definition.

This is per syntax because the renderer is. The same probe shows a
CommonMark autolink and a GFM autolink literal publishing their address
exactly as written — `<…?a=1&amp;b=2>` links to `…?a=1&amp;b=2` — so
the two autolink scans keep comparing the spelling. Normalizing all
four alike would ground two of them on a URL no reader reaches.

`unescapeMarkdownDestination` and `checkUrl`'s `markdownDestination`
parameter lose their only caller and are gone. `referenceDefinitions`
moves onto `parseMarkdown`, leaving one parser configuration in the
file, so the definitions it masks cannot drift from the destinations
`publishedDestinations` checks; log 02 records zero difference across
14 definition shapes, GFM-specific ones included.

Call sites, all module-private to support-reply.ts:
  unescapeMarkdownDestination  removed; had one caller, checkUrl
  checkUrl                     :537 :543 :554 :557 :564
  inlineDestination(s)         :349 (columns) :536 (validateProse)
  inlineDestinationEnd         :300 :453
  referenceDefinitions         :492 :542
  parseMarkdown                :147 :169 :253 :324 :378 :399
  validateProse                :636 — one caller, all three fields
No export changed; nothing outside this file references any of them.

Tests: the entity-encoded spellings of one evidence URL across inline,
angle-inline, image, definition, `&#38;` and `&#x26;`; the same in
summary, details and appliesTo; a five-syntax table checked in both
directions so the relation holds rather than the values; and four
encoded destinations no evidence decodes to, which stay refused. The
renderer-side pin in qa-components gains the four rows that record the
decode/no-decode split it rests on.

(cherry picked from commit a36728a3ac3ac74d92fc95ff05e82d863789e4e5)
…ion class

A GFM autolink literal written inside an emphasis or strikethrough run reached
the evidence check with the closing run attached. The raw-URL scan finds an
address by pattern and runs it to the next space, and the punctuation class it
then trims against holds no `*`, `_` or `~`, so `**Read <evidence url>**` was
compared as `<evidence url>**` and refused. The renderer publishes the run
outside the anchor: the reader clicks exactly the cited evidence URL. A reply
that was correct and fully grounded was discarded, wrapped as
InvalidSupportReplyError, and escalated to a human over a reason naming neither
the field nor the URL.

Ask the parser where the literal ends instead. `autolinkLiterals` reports the
span the grammar gives each bare address; each one is held to the evidence set
in the spelling the reader clicks and then masked from the pattern scan. The
check runs before the mask, so the scan is nowhere looser than it was: a literal
the parser finds in a region that scan deliberately still reaches — the contents
of a code span crossing a line — stays refused. Only the literal form is taken;
`[label](…)`, a reference definition and a CommonMark `<…>` autolink carry their
own delimiters and are already checked in the form each syntax publishes.

Same-pattern audit over the finite delimiter dimension — every run this grammar
forms (`*`, `**`, `***`, `_`, `__`, `~`, `~~`) in each position around an
evidence, a non-evidence and a scheme-less address, 126 rows, oracle = hrefs the
app's actual ReactMarkdown + remark-gfm publishes: 31 rows diverged from the
renderer before, 0 after, every one of them in the reject direction. The other
two scans do not carry the pattern: the `<…>` scan is bounded by its own `>`,
and publishedDestinations is grammar-derived already. Host-case canonicalization
(R14-PKG-A03) is a separate contract and is verified unmoved — `www.` accepted,
`WWW.`/`Www.` refused, as recorded at d8c932f.

Call sites of the changed symbols:
- AutolinkLiteral (new, module-local): support-reply.ts:334, :360.
- autolinkLiterals (new, module-local): one call site, support-reply.ts:589,
  inside validateProse.
- validateProse (body only, signature unchanged): one caller,
  validateSupportReply at support-reply.ts:683.
- checkUrl / maskMarkdownDestination (module-local closures, unchanged): one new
  call each at :590 / :591, alongside the existing :575/:581, :576/:582.
- Exported surface of support-reply.ts unchanged. validateSupportReply consumers
  are unaffected: support-agent.ts:347 and the AI barrel index.ts:11.

Red before the fix: 10 of the new AI rows failed at checkUrl (support-reply.ts
:524 via :557). Green after: 347/347 support-reply, 3995/3995 packages/outpost,
816/816 apps/web. No existing expectation weakened or removed.
…me from

`validateSupportReply` checked `summary`, `details` and `appliesTo` as the model
wrote them, and then `supportReplyDetails` moved both of the fields it composes
before anyone saw them. Trimming `details` deleted the four leading spaces that
made it an indented code block, so an example address the evidence check had
credited as inert republished as a live link to a host no evidence mentions.
Normalizing `appliesTo` onto one line deleted the same indentation, and the
escape it then ran through covered Markdown structure only, never the '@', '.'
and '/' a GFM autolink literal is built from, so a `www.` host and a bare email
published as links the same way.

The corruption ran in the other direction too. The escape rewrote the characters
of a cited URL — `…/reference/my-guide` reached the reader as
`…/reference/my%5C-guide` — and escaping a `[label](…)` citation into plain URL
text had GFM relinkify it onto that same rewritten address, so escaping a link
neither removed it nor kept it. Auditing the rest of the helper found the third
instance: an inline destination is decoded, so the source list published a cited
`…?a=1&amp;b=2` as the address that reference decodes to.

`details` now keeps its first line where the model put it and loses only blank
edges; `appliesTo` copies through the spans the grammar already resolves to a
link or image and escapes between them, over the inline constructs only; the
source list writes its destination so it decodes back. `validateSupportReply`
then checks the composed string against the same evidence set, which is what
holds whichever transform is added next.

Call sites of the changed symbols:
- supportReplyDetails (exported): formatter.ts:47, support-reply.ts:759
  (supportReplyText), support-reply.ts:645 (new: the composed check), barrel
  re-export index.ts:13, and support-reply.test.ts. Its output is unchanged for
  every reply whose fields carry no leading indentation, no address in the
  applicability and no character reference in an evidence URL.
- escapeMarkdown (module-private): sole caller is now publishedAppliesTo
  (support-reply.ts:707, :710); it no longer collapses whitespace, which
  publishedAppliesTo does instead. The unrelated same-named helper in
  apps/teams-bot/src/cards/ticket-created-card.ts is untouched.
- validateProse (module-private): callers are validateSupportReply:636 (the
  three fields, unchanged) and :645 (new).
- collectInlineLinks / parseMarkdown (module-private): gain one caller each,
  resolvedLinkSpans:672.
- publishedAppliesTo, resolvedLinkSpans, trimBlankEdges, escapeDestination are
  new and module-private, one caller each.

Red against the unchanged implementation at d8c932f: 9 of the new rows fail
(4 indented-details, 2 applicability-refusal, 2 applicability-fidelity, 1 source
list). Published strings are pinned against the app's real ReactMarkdown +
remark-gfm in apps/web/src/__tests__/qa-components.test.tsx, including the two
spellings publication used to emit.
…an guesses

`proseOutsideInlineCode` paired backtick runs by hand, which meant hand-parsing
everything else on the line that can hold a backtick without opening a span. Two
of those skips decided the outcome, and both answered a different question than
the renderer's.

The tag skip tested `/^<(?:!|\?|\/?[a-z])/i` and then jumped to the next '>' on
the line, whether or not the grammar closes a tag between them. Where it jumped
too far it swallowed a code span: `Compare <b, ` + span + `, and c> here.` is
published as `Compare &lt;b, <code>…</code>, and c&gt; here.`, no tag and no
anchor, yet the example address inside that span was checked as a citation and
the reply discarded and escalated. Where the jump landed past a backtick it left
the run after it to pair with a later one, and the span that mispairing invented
covered raw HTML the renderer publishes as prose — `<b\`> x \`<script>…</script>\` y`
passed a guard whose rule is that HTML is allowed only inside code. This
renderer escapes that HTML rather than mounting it, so what got through is the
policy boundary, not markup that runs.

The parser already reports each code span and its exact range, and a span it
reports can be neither an attribute nor a destination: a backtick inside either
opens nothing it reports. Masking those ranges retires both skips at once.

Only what a span encloses is blanked, never its delimiters. The raw-HTML check
re-reads the masked view, and `Use <b ` + span + ` > carefully.` — which the
renderer escapes whole — would become `Use <b   > carefully.`, a tag the grammar
does close, if the backticks went with the contents.

Unchanged: spans crossing a line are left whole and stay subject to the prose
checks, deliberately stricter than the renderer; reference definitions stay
exempt; a backtick inside a real destination still opens no span; and the
invalid/unclosed fence refusals are untouched.

Call sites of the changed symbols, all module-private to support-reply.ts:
- proseOutsideInlineCode: second parameter is now the line's code-span ranges
  rather than its destination columns. One call site, validateProse (:527).
- inlineDestinationColumns: removed. Its only call site was validateProse.
- InlineDestinationSpan.openerStart: removed with it. Produced by
  inlineDestination (:297), read by nothing else.
- sameLineCodeSpanContents, collectInlineCode, CodeSpanContents: new, called
  from validateProse (:523) and from each other (:378, :403).
- inlineDestinationEnd and inlineDestinations lost a caller each and keep their
  remaining ones (:297, :560).

Tests, oracle = the app's own ReactMarkdown + remark-gfm: four rows red at
b21d810 (two inert-'<' shapes wrongly refused, two mispaired spans wrongly
accepted), plus the controls that stay green — masking must not manufacture a
tag, an address outside the span is still refused, a '<' the grammar does close
is still raw HTML, and a span closed on a later line is still read as prose. The
published form of every shape is recorded in apps/web.
Planned production takedowns and security-scanning setup no longer force
CRITICAL priority. Recognized incident reports keep their priority floor.
The infinitival exemption requires a named planning/prospective frame;
unrecognized completed-action phrases retain their original behavior.

Validation: genuine red/green for the three original failures plus completed-
action regressions, 1,781 classifier tests, 4,110 core tests, scoped formatter/
lint, AI noEmit typecheck and isolated AI build. Existing expectations retained.

Call sites: TicketClassifier.heuristicClassify is called by classify in
classifier.ts and the fallback in pipeline.ts; the public signature and
exports are unchanged. Pipeline.classifyTicket is consumed by the queue
AI-response handler. New guards are local to heuristicClassify.

Combines isolated R14-AI-A05 commits 903ac8c, d6e8404 and 0b6e993 after
root verification; source and tests exactly match the final isolated tree.
proseOutsideFences refuses a line whose leading backtick run is followed by
another backtick, unless the parser reports a code span opening at exactly that
line and column. collectCodeNodes recorded openings only, so every later line of
a span the grammar had already closed carried a run the guard could not account
for. `Use `` a\n```b `` here.` was refused as an ambiguous fence while the chat
renderer publishes it as one paragraph holding one <code> element and no fence
at all, and the reply was discarded and the ticket escalated.

Record the lines each span already covers where they begin — every line after
the one that opened it, through the line its closing run is on — and skip the
refusal on those. The docstring at :137 already states the rationale the guard
now follows: only a run left open is the ambiguity it exists for.

Scope is the refusal alone. A span crossing a line is still not masked, so its
contents stay conservatively read as prose: an ungrounded address inside one is
still refused, now by the link check that owns that question rather than by the
fence guard. A run genuinely left open still refuses, including on a line that
also carries a closed span. The unclosed-fence footer probe, the column-free
block detection from R13-AI-A02 and the same-line masking from R14-AI-A02 are
untouched.

Call sites of every symbol changed, all module-private to support-reply.ts:
  - CodeNodes (interface, new `spanContinuations` field): declared :96,
    constructed :160, read :174 and :178. No other file names the type.
  - collectCodeNodes: defined :108, recursive call :118, called :161.
  - proseOutsideFences: defined :157, sole caller validateProse :529.
  No export surface changes; nothing outside this module references any of them.

Red before the fix: the three accepted rows threw 'Invalid code fence in support
reply' at support-reply.ts:163; two refusal rows threw the fence error instead of
the link error that should own them. The renderer fixtures in
apps/web/src/__tests__/qa-components.test.tsx pin the published form — one <p>,
one <code>, no <pre>, no anchor — against the app's real ReactMarkdown +
remark-gfm.

R14-AI-A04. Owns L14-3's multiline-structure rows.
A citation of a scheme-less `www.` host was held to the evidence set only
when the host was written in lowercase. GFM linkifies that prefix in any
case, and the scan that finds these addresses matches any case, but the
scheme rebuilt before the comparison was rebuilt only for `www.` — so
`WWW.copilotkit.ai/reference/provider` reached `canonicalSourceUrl` with no
scheme, parsed as nothing, and was refused, while the grammar-derived
destination for that same sentence carried `http://` and grounded. The two
halves of one check disagreed about which addresses exist, and a reply that
cited its evidence exactly was discarded and escalated to a human over a
link whose reader would have landed on the validated source. A
sentence-initial capital is the ordinary trigger: the schema permits it and
the instructions do not forbid it.

The renderer is the oracle for why accepting these is correct, not a
case-folding assumption. Its anchor for each spelling carries the address in
the case it was written, over http:// — recorded against the app's real
ReactMarkdown + remark-gfm in apps/web/src/__tests__/qa-components.test.tsx
— and a URL host is case-insensitive, so every one of those hrefs resolves
to the cited evidence URL. Each accept row asserts that resolution.

Only the prefix test changes, and only its case sensitivity. The host is
still folded by the URL parser rather than here, so nothing compares a path
or a query case-insensitively: `Www.copilotkit.ai/reference/Provider` is
still refused. So are an ungrounded host in every case, a different path on
the evidence host, and the bare evidence host with its path dropped. The
recorded reason for choosing http:// over https:// is preserved and
extended rather than replaced.

Red before the change, against this same source otherwise unmodified: the
three capitalized accept rows threw 'Support reply link URL must belong to
validated source evidence' from checkUrl, reached through the
grammar-derived autolink-literal loop — the structural autolink fix in
b21d810 moved where the refusal surfaces but not its cause. The eight
refusal rows passed before and after.

Call sites of the changed closure, all seven inside `validateProse`, which
is the only reader of this scheme reconstruction: the inline-destination,
reference-definition and autolink-literal loops, the two raw-URL scans
(the second the only caller passing `allowProsePunctuation`), and the
grammar-derived `publishedDestinations` loop. `validateProse` has one
caller, `validateSupportReply`, which runs it over summary, details and
appliesTo; that function is called in production only at
ai/src/support-agent.ts:347 and is re-exported from ai/src/index.ts. No app
package imports it.
…tails

Compose applicability from parser-resolved spans and bound literal URLs only when escaping changes their destinations. Escape nested images and escaped-bang adjacency so metadata cannot introduce remote images. Keep final composed-output grounding validation intact.

Call sites checked: resolvedSpans and composeAppliesTo are private to publishedAppliesTo; boundedAutolink is called only from composeAppliesTo; supportReplyDetails remains consumed by formatter, serialization and barrel exports. Public signatures unchanged.

Verified isolated red-green regressions, 4036 core and 830 web tests, typechecks and build. Combined branch targeted support-reply and formatter suites: 445 passing. Full combined verification follows.
…nsumer

`negativeObservationPrefix`'s trailing object phrase repeated three arms whose
whitespace overlapped: `\s*[,/]\s*` carried a run on both sides and the word and
coordinator arms both opened with `\s+`. Every gap in a separator run could
therefore be split two ways, so a prefix the group cannot accept made the engine
enumerate 2^gaps partitions before failing.

`heuristicClassify` runs synchronously in the queue worker's AI_RESPONSE handler
on the raw, uncapped body — only the model call is truncated — so a merely
comma-heavy ticket (a pasted log, a quoted CSV) pinned a worker CPU. Measured on
the emitted method: 10 separators 2.4ms, 20 separators 90.8ms, 24 separators
1.45s, 26 separators unfinished in 3s.

Each arm now consumes only the whitespace preceding its own token. A token is
never whitespace, so every leading run is pinned to a whole gap. The `and`/`or`
arm must still assert a following gap, so it takes a single `\s` and leaves the
remainder to the next arm — `\s` plus the next arm's leading run accepts exactly
what `\s+…\s+` did. The accepted language is unchanged; the exponent is gone
rather than bounded, and the body is neither capped nor truncated.

Verified equivalent, not assumed: 1,050,256 exhaustive regex inputs (token
sequences to depth 3 over gap widths 0-3) plus 333,453 separator-only inputs to
depth 4, 0 divergences; the 608 distinct bodies in classifier.test.ts and
500,000 randomised negated-observation sentences replayed through both builds,
0 divergences. Pumping every other repeated group in the file — the `without`
and `no` filler lists, the negated-predicate opener, the remaining-incident
list, the takedown determiners, the outage adverbs, `\w+n't`, `\w+ly` — found no
second occurrence: all stay flat to 20,000 repetitions before and after.

The regression runs in a hard-bounded child process. In-process it could not
fail, only hang: Vitest's testTimeout is a timer and the timer needs the event
loop the blocked regex is holding. It asserts completion within a generous
liveness budget rather than a duration, so it does not depend on machine speed.
Against the unchanged guard it fails in 60s with "did not finish"; the eleven
semantic controls — four negated separator-joined object phrases, each paired
with the same sentence minus its negation — pass identically on both sides.

Call-site verification before the original commit is recorded in the raw
13-symbol-call-sites.txt report. negativeObservationPrefix has one unchanged
caller in hasNonIncidentPrefix. separatedObjectToken has one caller in that
regex; incidentMention gains that one reader. New bounded-process helpers and
types are referenced only by classifier-backtracking.test.ts and its harness.
No public runtime exports changed. Root adds this omitted record to the
commit message only; the tested source tree is unchanged.
A reference definition can write its destination on the line after `]:`, where
the block container it sits in re-states the markers the parser already
consumed. Those markers are not part of the node the parser reports, so they
are still inside the raw range it gives, and skipping whitespace alone stopped
on the '>' and reported the marker as the destination. The marker was masked
and the destination was not, so a URL the definition check had already approved
decoded reached the raw scans in the spelling they cannot accept: an
entity-encoded query and a destination ending in an escaped ')' were both
refused, and the reply was escalated to a human over the link its own evidence
backed. A literal destination passed by accident, which is why the suite could
not see it.

The padding is now stepped over by the grammar's own bound: indentation, one
line ending, then one '>' per open block quote with its optional space, and a
'>' only ever at the start of a line. An angle destination may hold no line
ending, so its scan stops at one rather than closing on a later line's marker.
The span still covers the destination alone, so a raw URL written into a label
or a title stays subject to the scans, and the canonical, escape and entity
comparisons are untouched.

Call sites checked: `destinationSpan` has one caller, `referenceDefinitions`,
whose signature and `ReferenceDefinition` shape are unchanged;
`continuationPadding` is new and private to `destinationSpan`. No exported
symbol changed.

Verified red before green: 12 of the 27 new matrix rows failed on the previous
locator, all 415 prior support-reply controls held. 4224 core and 852 web tests
pass, with cold AI and web `--noEmit`, an isolated AI build, and scoped
Prettier and ESLint clean.
`no` and `without` cancelled an incident mention unconditionally, so "No
data loss root cause has been identified yet.", "No production outage
postmortem has been written." and "Without a data loss postmortem we
cannot close the incident." all landed at HIGH, indistinguishable from
the genuine absence report "No data loss has been reported.".

The file already decides this reading twice: `negatedIncidentObjectHead`
requires the mention to head the negated object rather than sit as a
modifier inside it, and the `incidentToolingHeads` rationale names
"postmortem", "root cause" and "mitigation" as heads that presuppose an
instance and "must leave it standing". The two determiner arms were the
ones that never consulted the suffix.

`negatedIncidentObjectHead` is not reusable here: it is a verb's object,
where the phrase runs to the clause end, so any further bare noun
continues it. `no`/`without` take a determiner phrase in subject or
complement position, where a predicate legitimately follows — "No data
loss has been reported." must stay cancelled. So the test is inverted:
enumerate the heads that end the cancellation instead of the predicates
that keep it. `incidentPresupposingHeads` is that enumeration, the
complement of `incidentToolingHeads`, and an unlisted continuation keeps
the absence reading.

"report" and "incident" are named as presupposing in the same comment but
stay out of the set: `objectPhraseHandoff` already reads those two words
as heading an absence report about the incident ("no data loss reports"),
and splitting those readings is a separate decision, recorded as unsettled
rather than taken here.

Callsite enumeration (performed before this commit):
  incidentPresupposingHeadSuffix -> hasNonIncidentPrefix (classifier.ts:247), sole reference
  hasNonIncidentPrefix           -> classifier.ts:449, consumed at :468, sole callsite
  heuristicClassify              -> classify() floor (classifier.ts:54, :72-78)
                                 -> AIPipeline.classifyTicket (pipeline.ts:454)
                                 -> queue/src/handlers/ai-response.ts:779, persisting
                                    classification.priority onto Ticket.priority
  test-only: ai/src/test-utils/bounded-heuristic-classify.child.ts:37,41
  no other module references the new symbols

Tests: the three reported spellings join
observedAbsenceVersusStandingIncidentContrasts paired with absence
controls, plus a closed grid —
cancellingDeterminerVersusPresupposingHeadContrasts — crossing the two
determiners with the three heads the suite already pins and the three
incident terms. 12 rows / 24 sentences, each pair also asserted not to
collapse. Red on the base source: 48 failures (12 standing sentences x 3
boundaries + 12 relations), every absence control already green. Green
after the fix.
Brings the Slack ticket mirror (#150), the worker health/boot rework (#262)
and the worker poll-loop claim fencing (#224) under this branch's support-agent
and AI-response-lifecycle work.

One content conflict, resolved by union:

  packages/outpost/queue/src/__tests__/ai-response.test.ts

Both sides edited the file's import preamble and nothing else collided. Ours
imported `Message` (@prisma/client) and `EscalationPayload` for the migration /
support tests; upstream widened the vitest import to include beforeAll/afterAll
and added a suite-level SHADOW_MODE guard that clears an inherited
SHADOW_MODE=true and restores it afterwards, because seven of its claim-fencing
and mirror tests assert the non-shadow path.

Resolution keeps every element of both: the 8-name vitest import, both of our
type imports, and upstream's ambient-shadow guard. The only deviation from
either side is that `JobHandlerContext`'s type import is hoisted to sit with the
other imports rather than below the guard block — same semantics, no statement
between imports. Both `Message` and `EscalationPayload` were confirmed still
referenced in the merged body (PersistedResponse, escalationPayload()).

No test was lost. Block counts reconcile exactly: base 98, ours +10, upstream
+8, merged 116, and a title-set comparison shows no block from either side
dropped. The single base title upstream still carries and the merge does not
("succeeds without claiming a handoff when the response is already DELIVERED")
is a base test upstream never touched and this branch deliberately rewrote into
two stricter cases, so taking our deletion is correct.

packages/outpost/queue/src/handlers/ai-response.ts auto-merged; reviewed
against both contracts and both hold:

  - Upstream's mirror contract is intact. `delivery: SlackMirrorDelivery` is
    declared before the delivery branches and claimed by every branch upstream
    claimed it in — 'shadow' before the shadow try, 'withheld'/'delivered' on a
    successful post, 'post-failed' in the catch — and the SLACK_MIRROR enqueue
    still runs after reportProgress(85), gated on isMirrorableSource +
    isSlackMirrorEnabled.
  - This branch's contract is intact. nonDeliveryEscalationReason is still
    computed before the message.create that persists it, escalationReason still
    gives delivery failure precedence, the PENDING-guarded updateMany and
    escalationReasonPersistenceError threading survive, and the
    both-writes-failed early return is unchanged.
  - The two do not interfere: the mirror renders aiMessage.content
    (pipelineResult.response), which this branch never retargeted — its
    publishableResponse change only touched ticket.suggestedResponse and the
    shadow SYSTEM row.

Symbol audit before committing: updateJobProgress gained a required third
claimToken argument upstream; every callsite is upstream's own and already
passes it, and this branch only ever uses context.reportProgress.
JobHandlerContext is unchanged. JobType carries both SLACK_MIRROR (upstream)
and PENDING_RESPONSE_SWEEP (ours). The new Job.claimToken column and its two
migrations do not collide with this branch, which adds no migration.
The test file's platforms mock already stubs readSlackMirrorConfig,
isSlackMirrorEnabled and isMirrorableSource.

.env.example and docs/deployment.md auto-merged as disjoint additive sections;
both sides' edits verified present.

Emergent behavior worth noting, not altered here: when both the DELIVERED state
write and the deliveryConfirmed marker write fail, this branch's early return
now also skips the mirror enqueue, which upstream documents as running
"regardless of whether the draft was delivered". That path is a double Prisma
failure, the enqueue is itself a DB write that would fail and only log, and the
job returns failure for retry — so it is benign, but it is a real interaction
between the two changes rather than something either side chose.
R15-AI-A04. `parseSourceUrl` accepts `'` and '`' in an evidence URL and this
renderer publishes both in an href — `…/provider&#x27;s` and
`…/provider%60name`, each the `new URL(evidence)` canonicalization of the
cited address. The raw-URL scan in `validateProse` runs an address to the
next character outside a class that excludes both, and the CommonMark `<…>`
autolink was the one link syntax with no grammar-derived span masking it from
that scan: an inline destination, a reference definition and a bare GFM
literal all have one. So the scan read a truncated prefix of the cited
address, failed to find that prefix in the evidence set, and escalated to a
human a reply whose only citation was its own evidence and whose reader would
have clicked straight through to it.

`uriAutolinks` asks the parser the renderer runs where each autolink's
address begins and ends, and `validateProse` holds it to the evidence in the
spelling this syntax publishes before masking that span — the same
check-then-mask order the three syntaxes above it use, which is what keeps
the pattern scans no looser than they were. The class is not widened, the
delimiters are not stripped, nothing else is masked, and canonicalization is
untouched: an ungrounded address, a non-HTTP or credentialed URI, an address
the autolink production does not close around, and an address in a code span
crossing a line all stay refused.

Private-helper audit, done before committing:
- `uriAutolinks` is new and file-local; one callsite, support-reply.ts:698,
  inside `validateProse`.
- `collectInlineLinks` is reused unchanged; callsites are now
  support-reply.ts:377, :415, :457 and :853.
- No other helper in the file is touched. Every helper here is file-local —
  the module exports only `supportReplySchema`, `SupportReply`,
  `parseSourceUrl`, `validateSupportReply`, `supportReplyDetails` and
  `supportReplyText`, none of which changed signature or behaviour outside
  the refusal above.

Red on b9f889d: 8 failed / 454 passed in ai/src/support-reply.test.ts.
Green here: 462 passed in that file; core 77 files / 4475 tests; web 53
files / 856 tests; both cold typechecks; fresh non-incremental AI emit at 23
modules / 69 files with no test material in dist.
@jerelvelarde
jerelvelarde marked this pull request as ready for review October 7, 2026 18:00

@NathanTarbert NathanTarbert left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Jerel, this is a huge amount of careful work. The writeup is honest about what the live runs do and don't show, which makes it much easier to review. The paragraph-plus-details format reads really well, and the MCP headers answer is exactly the kind of reply we want to be sending.

The first thing I'd like to look at is what happens when read_source is pointed at something that isn't a file. If the model asks for a directory like packages/react-core/src/hooks, GitHub returns a list instead of a single file. The parse in support-agent.ts around line 253 then throws. That error isn't one of the two the pipeline turns into a handoff, so it comes back up as a failure. In the worker, the job retries and fails the same way each time. In web QA, the user sees the raw validation error. Symlinks and submodules hit the same path, and so does a bad path around line 230. If these came back to the model as tool errors, it could correct itself and keep going.

The second is that the GitHub calls in support-agent.ts go out without a token. GitHub caps unauthenticated requests at a low hourly limit per IP, and each read_source makes two calls. A busy worker will hit that cap, and when it does, the investigation fails outright instead of telling the model the source is unavailable. Passing the token the rest of the app already uses would fix the limit. Turning a 403 into a tool result would cover the rest.

The third is about escalation after an interrupted run. In ai-response.ts around line 839, the escalation reason is now written when the row is created, before anything is posted. If the worker stops between creating the row and posting, the retry escalates with only the low-confidence reason. The human who picks it up isn't told the reporter got no reply, and the delayed takeover that waits for an in-flight post is skipped. Main handles this case today by flagging the delivery as failed, so it would be good to keep that.

A smaller one: for web tickets, formatStructured puts the footer in text, and the details get appended after it. So the saved suggestion and the shadow record end with the footer, then the technical details. The dashboard chat isn't affected because it renders details separately.

I also had a question. The investigator runs with store: false and doesn't request reasoning.encrypted_content. My understanding is that multi-turn runs with reasoning sometimes need that to carry reasoning between turns. Your September 21 runs clearly made tool calls fine, so I may be wrong here. Did you see any "item not found" errors along the way?

Last, I'd like us to agree on the default before this merges. As it stands, merging moves every AI call to Luna unless AI_RESPONSE_PROVIDER is set to anthropic. Since the description says posting should stay off while evaluation continues, it might be safer to land this with Anthropic as the default and flip it once the matched old/new comparison is done. Happy to talk it through if you see it differently.

Thanks again for all of this!

read_source and read_release set errorFunction=null, so every non-404 GitHub
failure escaped the model loop and aborted the whole investigation instead of
letting the agent correct itself. Six predictable failures escaped:

- 403/429/5xx and any other non-ok status threw `GitHub evidence request failed`
- a transport failure (fetch reject) threw a raw TypeError
- a non-JSON body threw SyntaxError out of response.json()
- a directory path returned an entry array and failed `size` parsing (ZodError)
- symlink/submodule/other non-file blobs failed the base64 parse (ZodError)
- an unsafe path threw InvalidSupportReplyError, ending the run

githubJson now reports these as a discriminated result, and both tools turn them
into bounded tool output the model can act on: `invalid_path` (returned before
any URL is built, so an unsafe path is never fetched), `not_a_file`, `unreadable`
and `unavailable` with a sanitized reason — access_denied, rate_limited,
upstream_error, invalid_response or transport_error. No GitHub response body,
header or address reaches the model or the trace.

Preserved deliberately:
- repository/ref/path allowlists, and `not_found` + `too_large` as they were
- the size check still precedes base64 decoding
- only a validated single file resolves to a pinned sha and enters `remember`
- a failed call still spends one of the six; the budget never resets
- errorFunction stays null, so programmer errors still surface; AbortError and
  the 60s deadline still terminate the run rather than becoming tool output
- InvalidSupportReplyError on final output and InvestigationBudgetError unchanged

Callsite audit:
- pipeline.ts:171-194 is the only consumer of investigate(). It rethrows
  anything that is not InvalidSupportReplyError/InvestigationBudgetError, which
  is why a 403 previously failed the ticket and triggered a worker retry; those
  cases now route through the agent instead. Its two handled error types are
  untouched, so pipeline-openai.test.ts is unaffected.
- index.ts re-exports both error classes; neither signature changed.
- docs/support-agent.md documented "API outages still raise errors" and is
  updated to the new status vocabulary.

Tests: real SDK multi-turn aimock runs where a failed or wrong-resource read is
followed by a corrective grounded source read that succeeds, plus budget
exhaustion on six failed reads, sanitization assertions, and abort termination.
All 16 new cases fail on the parent commit. Two tests that asserted the old
escaping behavior (bad path throwing, malformed payload throwing) are converted
to assert the structured result; no assertion was weakened.

Verified: packages/outpost suite 4491 passed / 77 files (baseline 4475 / 77),
cold typecheck clean, eslint clean, prettier clean.
`escalationRequiredReason` is stamped on the primary response when the row
is created, before any publication. The prior-response gate read it first
and went straight to `recoverRequiredEscalation`, so an attempt interrupted
anywhere around the platform post — PENDING row, owed reason, no
`deliveryConfirmed`, no `responseError` — was settled as if the answer had
already gone out. Two consequences on exactly the rows that most need care:

  - the delayed takeover was bypassed. That delay exists so recovery never
    races a post still in flight; the owed marker predates the post, so
    acting on it summoned a human against an answer about to land.
  - the handoff reported `deliveryFailed: false` and handed the human the
    bare generation-time reason ("Low AI confidence (25%)"), which reads as
    "a weak answer was sent" to someone whose reporter may have nothing.

Delivery was also genuinely unrecorded for these rows: a successful post
that owes a handoff keeps the row PENDING, so it took neither the DELIVERED
transition nor the `deliveryConfirmed` flag, and nothing distinguished it
from an attempt that died before posting.

  - record `deliveryConfirmed` when publication succeeds on a row held
    PENDING by its owed handoff (non-fatal; the escalation below it is what
    the reporter is owed).
  - route an owed handoff with no recorded delivery outcome, still owned by
    this job, through the existing `schedulePendingResponseRecovery`. The
    takeover re-enters the same branch after the delay and escalates.
  - report `deliveryFailed` from proven delivery rather than from a recorded
    error, so unknown reads as failed instead of as delivered.
  - compose the reason in one shared helper, `recoveredHandoffReason`: the
    stored promise verbatim when the outcome is known, and the promise plus
    "the reporter may have received no response at all / a human must verify
    the thread" when it is not.
  - PENDING_RESPONSE_SWEEP: an owed handoff now outranks `deliveryConfirmed`.
    Repairing such a row to DELIVERED dropped the promised human AND left the
    marker on a settled row — the pair the gate's first branch fails on — and
    it uses the same helper so a sweep and a takeover read identically.

Preserved: proven delivery never reports a failure, a recorded delivery
failure keeps its diagnostic and escalates with no delay, escalated rows stay
idempotent through the compare-and-set, no path reposts, and the claim
fencing (`responseJobId` = this job), database-clock recovery window, and
single CAS/transaction are untouched.

Callsite audit — `deliveryConfirmed` / `escalationRequiredReason` are written
and read only by these two handlers (repo-wide grep; no app, API or dashboard
reader). `recoverRequiredEscalation` has one caller, updated for its new
ticketSource argument. `hasConfirmedDelivery` widened to a Pick, all four
callers still type-check. `recoveredHandoffReason` is consumed by the handler
and the sweep. `schedulePendingResponseRecovery` gains a caller whose row
satisfies the same PRIMARY/PENDING/own-claim preconditions its CAS requires.
The schema change is comment only — no column, no migration.

Tests: nine added (interrupted low/medium/suppressed rows deferred then
escalated undelivered; a same-job retry waiting out an in-flight post;
delivery recorded on a PENDING owed-handoff row; a recorded failure escalating
at once; a stale duplicate not deferring a claim it lacks; sweep escalating a
confirmed delivery that owes a handoff; sweep keeping a failure diagnostic).
Seven of them fail on c602d6c. Two existing fixtures gained
`deliveryConfirmed: true` because the handler now writes it on the rows they
model, with their assertions unchanged.
…p installation

read_source and read_release were the only GitHub callers in this package that
sent no credential — githubJson set an Accept header and nothing else — so every
ref resolution, file read and release lookup ran on the anonymous 60 req/hour
budget shared by the whole host, while the worker already holds App installation
credentials for the same two repositories.

Reuse those credentials rather than introducing a new one. A new
github-evidence-auth module reads GITHUB_APP_ID / GITHUB_PRIVATE_KEY /
GITHUB_INSTALLATION_ID exactly as shared/platforms/registry.ts and
apps/github-app/src/lib/github-client.ts already do, and mints an installation
token through @octokit/auth-app (already a dependency of this package — no
dependency change). No gh CLI token, no personal access token, no key file.

Behaviour:
  - Tokens are narrowed to `contents: read`; metadata:read is inherent. The
    strategy is built on first use, so a run that never reads GitHub signs no JWT.
  - auth() is invoked per request so @octokit/auth-app owns caching and pre-expiry
    refresh; nothing here holds a bearer string that would die an hour into uptime.
  - Anonymous public reads remain available, but only when none of the three
    variables is set. A partial or malformed configuration rejects at the point of
    use instead of degrading into a quietly rate-limited anonymous success.
  - A signing or token-exchange failure is re-thrown carrying only the underlying
    error's class name — not its message and not as `cause` — because that error
    reaches worker logs and run reports and the underlying one can quote the PEM
    or the freshly minted token.
  - SupportAgent defaults githubAuth from the environment, so pipeline consumers
    (pipeline.ts constructs `new SupportAgent({ pathfinder, model })`) authenticate
    without opting in. The option exists as a test seam only.

Integrated with the recoverable-evidence-failure change already on this branch
(c6f4922), which this commit was authored against a sibling of. Credential
resolution now happens inside githubJson's recoverability boundary, so a
configured-but-unusable credential becomes bounded tool output rather than an
escaping throw that fails the ticket:
  - githubEvidenceHeaders() is awaited in its own try, ahead of the fetch, and a
    GitHubEvidenceAuthError becomes `unavailable` with the new sanitized reason
    `auth_unavailable`. Only the bare reason crosses to the model; the message
    naming the missing variable, and the underlying error, stay worker-side.
  - The catch converts GitHubEvidenceAuthError and nothing else. Any other fault
    in the resolver is a programmer error and is re-thrown, preserving the
    branch's errorFunction=null contract; rethrowIfTerminal still runs first, so
    AbortError and the 60s deadline terminate the run as before.
  - The request is abandoned on a credential failure, never retried without the
    credential, so a misconfigured host cannot silently fall back to a
    rate-limited anonymous read.
  - Everything else the recovery commit established is retained unchanged: the
    sanitized GithubResult union, safeParse on every payload, not_found /
    invalid_path / not_a_file / too_large / unreadable shaping, the six-call
    budget that a failed call still spends, and terminal abort propagation.

The existing auth test that asserted a thrown GitHubEvidenceAuthError is
deliberately re-pointed at the integrated behaviour: it now asserts the model
receives a bounded `unavailable` / `auth_unavailable` tool result, that no
request reached api.github.com (proving no anonymous retry), and that the
missing-variable name does not appear in what the model sees. The other two auth
tests and all 18 github-evidence-auth unit tests are unchanged.

Callsite audit — every githubJson caller, all in support-agent.ts:
  - L322 read_source  ref -> commit resolution ...... passes githubAuth
  - L341 read_source  contents read ................. passes githubAuth
  - L414 read_release releases/tags ................. passes githubAuth
No other call exists; grep for api.github.com / raw.githubusercontent across
packages/outpost/ai/src finds no second GitHub fetch. pipeline.ts:106 is the only
production SupportAgent construction and takes the env default. The origin stays
hardcoded and repositories stay enum-allowlisted, so the model can neither supply
a URL nor pass an Authorization header to an arbitrary source; headers are
assembled in one place and a token can only ride on a request to GitHub's API.

docs/support-agent.md gains the three App variables under Configuration and
`auth_unavailable` in the documented failure-reason vocabulary.

Tests use synthetic placeholders only. Verified: support-agent.test.ts 46 passed
(43 recovery cases preserved unchanged + 3 auth cases), github-evidence-auth.test.ts
18 passed, packages/outpost suite green, cold ai typecheck clean, eslint and
prettier clean on the touched files.

(cherry picked from commit 42e55fb5929c24a4149df2f5784e9ecbc22b6f56)
…e shaping

The cherry-pick of the App-installation auth change (78d6f8c) onto the
recoverable-evidence-failure change (c6f4922) created a boundary neither parent
commit could test on its own: credential resolution now sits inside githubJson's
recoverability try. Two properties of that seam were only implied by the resolved
code, so they are asserted here rather than left to a future reader's reading.

  - A token-exchange failure must not leak the signing key into model-visible
    output. The auth module already strips the underlying error to its class name
    (github-evidence-auth.test.ts covers that in isolation), but nothing proved
    the sanitized value is what actually reaches the tool result after shaping.
    This drives a real githubEvidenceAuthFromEnv resolver — through the
    SupportAgent test seam — whose exchange throws an error quoting a synthetic
    PEM, then asserts the tool result the model receives carries
    `auth_unavailable` and neither the PEM header nor its body.

  - A non-credential fault in the resolver must still escape the model loop. The
    catch converts GitHubEvidenceAuthError and re-throws everything else, which is
    what keeps the branch's errorFunction=null contract intact; without a test, a
    later widening of that catch to `catch (error)` would silently launder
    programmer errors into a status the investigator routes past. A resolver that
    throws a TypeError is asserted to reject the investigation.

Both also assert no request reached api.github.com, so neither path can regress
into an anonymous retry of a request whose credential failed.

Synthetic placeholders only; no credential here works anywhere. Mutation-checked
against the resolved source: widening the catch to all errors fails the second
test; falling back to an anonymous read or hoisting the headers call out of the
recoverability try fails the first (and the re-pointed partial-config test from
78d6f8c). support-agent.test.ts 48 passed, packages/outpost 4523 passed / 78
files, cold ai typecheck clean, eslint and prettier clean.
`formatStructured(reply, 'web')` returns a two-pane result: `text` is the
summary pane and ALREADY ends in WEB_FOOTER, `details` is the disclosure pane.
That split is a UI contract, not a serialization — but the two sinks that hold
exactly one string reassembled it by appending `details` to `text`:

  packages/outpost/queue/src/handlers/ai-response.ts  completePublishableResponse
  packages/outpost/ai/src/pipeline.ts                 streaming fallback composer

so the durable `suggestedResponse`, the SHADOW_MODE SYSTEM record and the web
string stream all read summary → footer → details, with the footer and the
disclaimer stranded mid-response and the sources printed after the sign-off.

Fixed at the formatter, the only place that knows where the trailing matter
goes. `FormattedResponse` gains an optional `completeText`, composed by the web
branch from the same pieces in reading order (summary, details/sources,
disclaimer, footer last) and left absent wherever `text`/`parts` is already the
whole response. Both composers now read it through one shared `publishableText`
helper instead of re-deriving the join in two places.

Deliberately NOT done: splitting the footer back out of text the formatter did
not write, and smuggling details through `parts` — the latter would have
duplicated them in the streaming composer, which joins parts and then appends
details.

Callsite audit — every reader of a FormattedResponse:
  queue handlers/ai-response.ts:841  suggestedResponse  -> completeText (FIXED)
  queue handlers/ai-response.ts:906  shadow SYSTEM row  -> completeText (FIXED)
  ai pipeline.ts:489  generateStreamingResponse         -> completeText (FIXED)
  apps/web api/qa/route.ts:78,91  streams `text`, sends `details` as metadata
      -> UNCHANGED; the native QA disclosure still renders the two panes
  apps/github-app github-poster.ts:53  `text` + feedback  -> UNCHANGED
  shared platforms/{discord,slack,teams,github}.ts postResponse
      -> UNCHANGED; Discord keeps `parts` precedence and ordering, GitHub keeps
         its single collapsed <details>, neither carries a `details` field
  shared platforms/types.ts FormattedResponse (adapter-side copy)
      -> untouched; adapters never read `details`, so nothing to mirror

A FormattedResponse without `completeText` — a hand-built fixture, or a value
from an older run — still serializes exactly as before: `publishableText` keeps
the historical text/parts-then-details join as its fallback branch. Suppressed
runs go through `format()`, which sets neither `details` nor `completeText`, so
no withheld draft or internal handoff reason gains a new way out.

Tests: red first on all three suites (footer at index 50 ahead of the sources at
77 in the web stream; the shadow record asserting the stranded-footer string),
green after. The pre-existing queue fixtures that assert the old join are
correct as written — they carry no footer in `text`, so they pin the
compatibility branch rather than the bug, and are left alone. The queue module
mock now spreads the real `@copilotkit/outpost/ai` so these assertions exercise
the real serialization rather than a stub.

Verified: 77 files / 4493 tests pass, `pnpm typecheck` clean across db+ai+queue+
shared, apps/{web,github-app,worker} `tsc --noEmit` clean, eslint and prettier
clean on every changed file.

(cherry picked from commit fd0a4238a46f6e5715f048a851b2bb8e2ec9e609)
…dline

githubJson awaited githubEvidenceHeaders(auth) before the fetch that carries the
investigation's 60-second AbortSignal, and authorization() took no signal at all.
A stalled App token exchange therefore held the investigation open past its own
deadline: the bound only ever covered the request, never the credential the
request was waiting for.

Thread the investigation signal through the headers seam and race the
authorization against it. The pending exchange is detached rather than cancelled:
@octokit/auth-app serves every concurrent caller with the same installation and
permissions from one shared in-flight request (verified against 7.2.2 — two
concurrent auth() calls issue a single HTTP request), so cancelling on behalf of
one abandoned investigation would reject the others. Detaching lets it finish and
fill the SDK cache for whoever is still waiting, and leaves SDK-owned caching and
pre-expiry refresh untouched.

The investigation's signal never reaches the SDK. The exchange gets a fixed
10-second HTTP deadline of its own instead, applied per attempt through the
`request` option — the only seam the strategy exposes, since auth() forwards no
per-call request. Without it a hung exchange is permanent: the SDK hands the same
stalled promise to every later caller.

Cancellation is terminal. It rejects with the signal's own reason, so a cancelled
authorization is indistinguishable from the cancelled fetch it was about to
authorize, and can never be mistaken for a retryable GitHubEvidenceAuthError.
An already-aborted signal does no credential work at all — no strategy
construction, no JWT, no exchange — and abort listeners are dropped whether the
authorization succeeds or fails.

Also:
  - Replace `createAppAuth as unknown as InstallationTokenFactory` and the
    redundant Promise cast with a typed adapter. The overloaded AuthInterface is
    not structurally assignable to the one narrow signature this module needs;
    adapting lets both sides keep their real types. Test mocks drop `as never`
    along with it.
  - Lazy strategy construction now runs inside the same sanitized boundary as the
    exchange, so a synchronous createAppAuth validation failure cannot echo the
    PEM. Still no cause and no message forwarding — only the error's class name.

Mutation-checked: eight mutations, all killed, no survivors. Reverting the single
wiring line to its previous form fails `abandons a cancelled investigation
without issuing the evidence request` by timeout; reverting both production files
to 42e55fb fails 11 tests behaviourally, none by module resolution.

Default provider and reasoning settings are unchanged. Tool failure shaping is
untouched and still belongs to the concurrent tool-recovery finding, whose
terminal error guard continues to receive cancellation unwrapped.

(cherry picked from commit a67858c4790803f8f18b8a7d82f77b1409432085)
`formatDiscord` split at `DISCORD_MAX_LENGTH - 50`, then appended a
109-UTF-16-unit `STANDARD_FOOTER` to the last part and cut the result back
to fit with `slice(0, 1997) + '...'`. The reserve was less than half the
thing it reserved for, so the cut was the normal path for a band of
response sizes rather than an overflow guard — and the end of the last
message is the footer, so what it removed was always published copy.

Verified against the real formatter, not read off the source. A 1892-char
body lost the 👎 half of the feedback prompt; 1949 and 1950 lost the Docs
URL mid-link; 3898-3900 took the same damage one part further along. At
1896 the cut landed between the two code units of 👍 and published a bare
\uD83D. Sweeping body lengths 1-6000 across six shapes and eleven
adversarial inputs (36,011 bodies), the previous code broke the posting
contract on 3,229 of them: 1,068 trailing ellipses, 1,068 mangled footers,
25 unpaired surrogates. It is now 0.

The splitter is told what the footer costs and returns a last part with
room for it, so there is nothing left to truncate and the truncation
branch is gone. Closing a fence at a split point is charged against the
budget too, rather than fitting by accident in the 46 characters that
happened to be spare. The fallback hard cut steps back a unit rather than
halving a surrogate pair.

Same splitter, second defect: the continuation was built with
`slice(splitAt).trimStart()`, which stripped all leading whitespace and
not just the newline the split consumed — inside a fenced block that is
the snippet's own indentation, so code a reader was told to copy came out
rewritten. `splitPoint` now reports how much of the break it consumed and
exactly that much is dropped.

Call sites audited. `shared/platforms/discord.ts:177` posts `parts` and
attaches the buttons to the last message, which now also carries a
complete footer; its own `splitMessage` only runs when `parts` is absent,
which is the <=2000 single-message case. `publishableText` and its two
consumers — `queue/handlers/ai-response.ts:947` and `pipeline.ts:491` —
keep their shape and stop recording a truncated response durably.
`github-poster.ts` is on the GitHub path and untouched.

`truncated: true` is deliberately left on a split response: no runtime
code branches on it (only fixtures and the type do), and changing it would
alter a public field for every split Discord response, outside this fix.
It now means "split", which is worth renaming separately.

Not fixed, same class, different file: `splitMessage` in
`shared/platforms/discord.ts:244` hard-splits at 2000 with no surrogate
guard on the `postMessage` path.

Tests first: 7 failed / 41 passed before, 48 / 48 after. The one new test
that was green at baseline is labelled as a regression guard rather than
evidence — plain prose was never the part that got cut.
…aders

`githubJson` mapped every 403 to `access_denied`. GitHub does not give a rate
limit a status of its own: an exhausted primary limit answers 403 *or* 429 with
`x-ratelimit-remaining: 0`, and a secondary limit answers 403 or 429 and may add
`retry-after`.
https://docs.github.com/en/rest/using-the-rest-api/rate-limits-for-the-rest-api

So a throttled `read_source`/`read_release` reached the model as a permission
denial. That is the one failure the model cannot reason about correctly: a
denial invites switching to another repository or path, while a throttle means
every GitHub read in this investigation is spent and the remaining budget
belongs on search evidence. The reply built on top of that is wrong for a
reason the model cannot see.

A 403 is now `rate_limited` when the headers prove a throttle, and
`access_denied` otherwise. 429 is unchanged and still unconditionally
`rate_limited`.

Bounded deliberately:

- Headers only, read positively. An absent, empty or unparsable header leaves a
  403 a permission denial, so a malformed header can never manufacture a
  throttle. `x-ratelimit-remaining` must equal `0`; `retry-after` must parse as
  digits. A duplicated header (`Headers.get` joins to "0, 0") fails both and
  falls through to `access_denied`.
- The response body is never inspected. It is untrusted and already excluded
  from the model and the trace; classifying on its text would both reintroduce
  that coupling and misread an ordinary denial whose message happens to mention
  a limit.
- No sleeping, backoff, retry or transport framework. The tool still returns one
  recoverable result and the model still decides what to do next.
- No new vocabulary. `rate_limited` was already in `GithubFailureReason` and
  already documented in docs/support-agent.md; only 403 now reaches it.

Caller and pattern audit:

- `githubJson` has exactly three callers — `read_source` (commits, contents) and
  `read_release`. All three forward a non-`not_found` failure to
  `unavailableEvidence`, unchanged.
- The only branch on `reason` anywhere is `=== 'not_found'`. `access_denied` and
  `rate_limited` are pure pass-through, so no control flow changes: the fix
  alters only the word the model is told.
- No other rate-limit, 403 or 429 handling exists in production source. The
  `apps/web` 403 checks are Next.js session auth, and
  `apps/web/src/lib/auth.ts` hits a different GitHub endpoint for org
  membership — neither shares this path.
- Terminal behaviour untouched: `rethrowIfTerminal`, the investigation deadline,
  cancellation, the six-call budget and `errorFunction: null` propagation all
  unchanged. A failed call still spends one of the six.
- docs/support-agent.md describes the reason vocabulary without asserting a
  status mapping, so it stays accurate.

Tests: 7 added, table-driven over the same stub helpers. Primary-exhausted 403,
secondary 403 with `retry-after`, 429 with and without headers, 403 with budget
remaining, and 403 with unparsable and with empty headers. Each asserts the
model receives a recoverable failure, then corrects the ref and still
synthesizes an answer, and that no header value or body text reaches the model.

Red before green: at parent 1239b8b the two header-signalled 403 rows fail
(2 failed | 54 passed); the five guard rows pass, so the fix had to be narrow.
Mutation sweep: 7 mutations, 0 survivors.
An evidence-auth failure gave the operator nothing to act on: the model-facing
recoverable reason stayed auth_unavailable, so an operator could not tell a
missing credential from a rejected one without reading raw credential state.

Classify the failure into an allowlisted operator category and surface that
category only. The diagnostic is built from a snapshot read once per call
rather than once per branch, so a credential that changes mid-call cannot
produce a self-contradictory category, and no raw credential value, token or
key material reaches the message. The recoverable model-facing
auth_unavailable reason is unchanged.
The web response sanitized validated code on its way out, so a reply whose
code had already passed validation could reach the requester altered and no
longer runnable. Both the single-string response and the summary pane took
this path, and each serialized independently, so fixing one still left the
other able to corrupt the same snippet.

Preserve validated code verbatim in both validated fields and in both
serializations. Sanitization now applies only to the unvalidated disclaimer,
which is the untrusted prose it was meant for. The legacy formatter path and
the existing completeText output are unchanged.
…knows

Two ways a reason the pipeline derived itself failed to reach the human who
picks up the escalation.

An explicit investigator diagnosis short-circuited the deterministic reason
instead of joining it. On the support-agent path `explicitHandoffReason ||
legacySuppressionReason` meant that whenever the model supplied its own
handoffReason, the groundedness and lint findings the pipeline had already
computed about that same draft were dropped. The two explain different things
— why the investigator gave up, versus what the gate independently caught in
what it produced — so neither substitutes for the other. They are now joined,
deterministic first, which is also what survives the 2000-character bound when
a diagnosis runs long. The legacy path already combined them; this makes the
two paths agree.

A forced escalation carried no reason at all. `groundedness.forcesEscalation`
clamps the score below the escalation gate, so a human is committed — but the
reason was gated on `suppressed`, which that outcome deliberately does not set
(the draft still publishes). The result reached the reviewer with the
escalation and no statement of which assertion to check.

Reasons are attached to those two deterministic outcomes only. A reply that
merely scored under the gate stays reason-free: its score is not a finding, and
restating it would bury the real ones. Suppression, published body, disclaimer
selection and confidence arithmetic are untouched.

The reason now crosses into a durable consumer.

Producing a reason on the forced-escalation outcome is only half of it. That
outcome publishes, so its handoff comes down the queue handler's
low-confidence arm, and `nonDeliveryEscalationReason` consulted `handoffReason`
on the suppressed arm alone. The newly-produced reason would have been dropped
at that seam — and dropped for good, because that expression is what the
durable `escalationRequiredReason` column is written from, before any
publication. Neither the owning job's retry via `recoveredHandoffReason` nor
the PENDING_RESPONSE_SWEEP backstop could recover a reason never written;
both read the stored row. So the handler now carries it too, and the rules
that govern it there match the ones above:

  - Appended to the generic sentence, never substituted. Existing readers key
    on the percentage — including the sweep, whose recovered text asserts the
    stored reason leads it — so the numeric prefix stays first and literal.
  - An absent, empty, or whitespace-only reason contributes nothing, so a
    merely low-scoring reply reads exactly as it always has. Pinned by three
    negative controls.
  - Only the text changes, never the null-ness. The condition deciding who
    queues a human is bit-for-bit identical, as are suppression, publication,
    delivery, the compare-and-set, idempotency, phase faults and shadow mode.
    No new routing, no new escalation, no new public export.
  - The delivery-failure arm still replaces the owed reason with its more
    actionable diagnostic, as documented. That precedence is deliberate and
    unchanged.
  - The 2000-character bound is restated at the handler seam. It duplicates
    the producer's own slice on purpose: past this point the value lives in a
    durable column and a queued job payload, neither of which should be able
    to grow because a future producer skipped its bound.
  - The reason is internal. It reaches the durable row and the queue payload,
    and is asserted absent from the arguments the platform post is built from.

Handler tests are on the handler, not the pipeline: four were red against the
unchanged handler for exactly this reason, covering the seeded durable marker,
the actual ESCALATION job payload, a failed enqueue recovered on retry with
proven delivery, and the interrupted takeover that proves the reason survives
to a recovered escalation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants