Repository navigation
fix: replay OpenAI Responses reasoning from encrypted content - #63
ibetitsmike wants to merge 8 commits into
Conversation
Defer synthetic user media until all contiguous tool-result messages have been emitted. OpenAI rejects media inserted between replies to parallel tool calls. Preserve accompanying text and the order of media payloads. Cover both Chat Completions serializers with separate and grouped tool results, two media results, and following conversation messages. > Xum acted on Mike's behalf.
Capture encrypted_content, item ID and summary from the completed reasoning output item (stream output_item.done and Generate output) and mark that metadata Finalized. Replay finalized metadata with a blob as a full reasoning input item regardless of store, keeping its position before function, computer and web_search items. Unfinalized metadata (stream placeholders and rows persisted before this field) keeps the previous behavior: item_reference with store=true, skipped otherwise. Update the recorded follow-up request bodies of the summary-thinking cassettes to include the replayed reasoning items.
Empty summaries remain empty arrays rather than invented summary text entries.
…nses Record whether the source response was stored on finalized reasoning metadata and only replay a following web_search_call item reference when it was. Reasoning from an unstored response is still replayed inline from its encrypted content.
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Codex Review: Didn't find any major issues. What shall we delve into next? Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
mafredri
left a comment
There was a problem hiding this comment.
Capturing the reasoning item from output_item.done and replaying it inline regardless of store is the right shape, and gating inline replay on Finalized keeps the output_item.added placeholders already stored in Coder chats away from the API. One concern, one check, two suggestions.
Suggestions:
- The six summary-thinking cassettes (Azure gpt-5-mini, OpenAI gpt-5, OpenAI o4-mini) were edited offline, and the UAT through Coder ran gpt-5-mini and gpt-5.4. Could we re-record them, so each edited follow-up request has been accepted by the real API?
- charmbracelet#407 replays the same item fields inline, but only for
store=false, and has noFinalizedgate, so it would also replay the.addedplaceholder blobs in rows persisted before the fix. Could we proposeFinalizedthere, so the fork does not have to carry it?
Didn't review the tests in depth.
🤖 This review was automatically generated with Coder Agents.
OpenAI echoes store=false for zero data retention organizations even when the request asked to store, so a later web_search_call item reference to that response would 404. Record SourceStoreEnabled from the requested store unless the response (Generate) or response.created event (Stream) echoes store=false. An absent echo keeps the requested value, and an echo never upgrades a requested false.
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep it up! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Thanks for the review and the pointer, @mafredri. I agree the I added it to charmbracelet#407 with the same field name and JSON tag ( charmbracelet#407 still replays inline only when I re-recorded the four OpenAI summary-thinking cassettes (gpt-5, o4-mini) against the live API: the follow-up requests with inline encrypted reasoning return 200. The two Azure cassettes are still edited offline, because I don't have Azure access. |
coder/fantasy#63 now targets coder_2_33 and merges in #64 and #65, so the pin carries the merged base plus the reasoning replay changes.
OpenAI Responses reasoning was only replayed as an
item_reference, which needsstore=true. With the defaultstore=falseit was dropped from later requests, so reasoning models lost their chain of thought between tool steps and turns.This captures the completed reasoning item (ID,
encrypted_content, summary) fromresponse.output_item.doneand from the Generate output, marks that metadataFinalized, and replays it as a full reasoning input item regardless ofstore, before its function, computer, and web search items. Completed summaries are kept verbatim, so an empty summary stays[]. Unfinalized metadata (stream placeholders and rows persisted before this field) keeps the old behavior: an item reference withstore=true, skipped otherwise.Finalized metadata also records whether the source response was stored: the requested
store, unless the response echoesstore: false(as it does for zero data retention organizations). A followingweb_search_callitem reference is replayed only when both the source and the destination request store; otherwise it is omitted, because a reference to an unstored item returns 404, while the reasoning is still replayed inline.The follow-up request bodies of the six summary-thinking cassettes were updated offline to include the replayed reasoning items; they were not re-recorded live.
Targets
coder_2_33. The branch merges incoder_2_33(#64 and #65) at 88cffb5, which coder/coder#30298 pins, so coder getscoder_2_33plus this change.Testing
go test ./...,go build ./...,go vet ./..., golangci-lint, and targeted race tests forproviders/openaipass.responses_reasoning_replay_test.gocovers Generate and Stream capture, replay understoretrue and false, unfinalized fallbacks, and web search reference eligibility from stored and unstored sources. Red-green verified by disabling each gate.Review record
store: falseis fixed in d55a867 (Generate and Stream tests, red-green verified; not tested against a ZDR organization). Replaying web searches that had no preceding reasoning is a pre-existing gap, tracked in chatd: OpenAI web searches without a preceding reasoning item are never replayed coder#30528.govulncheckfails oncoder_2_33itself with the same 11 IDs (otel GO-2026-6505, grpc GO-2026-6348,x/net, and Go 1.26.6 standard library); this PR does not changego.mod.