feat(llm): model-generated image output via OpenAI Responses (#255 Phase 5) - #536
initializ-mk wants to merge 1 commit into
Conversation
…ase 5) Forge can now return model-GENERATED images as file parts in the A2A response, completing the multipart-output direction. Opt-in and Responses-only (Anthropic/ Gemini/Bedrock don't generate images through their chat/messages APIs). - llm: StreamDelta gains Parts (non-text output); ClientConfig gains EnableImageGeneration. - Responses provider: when EnableImageGeneration is set, buildRequest sends the `image_generation` built-in tool; readStream reads the image_generation_call base64 `result` from the authoritative response.completed frame (not fragile partial-image events) and emits it as an image ContentPart; Chat aggregates delta.Parts into resp.Message.Parts. responsesTool.Name is now omitempty so the built-in tool serializes without a name. - Executor: llmMessageToA2A projects response image/document Parts into A2A `file` parts, and the assistant message's generated media is persisted like inbound media (URI in history, no base64 bloat; Bytes stay inline so the image is surfaced this turn). The empty-assistant-turn placeholder now respects a media-only response. - Config: ModelRef.ImageGeneration (yaml image_generation) → ClientConfig, off by default (changes provider behavior + cost). Tests: Responses parses a generated image from a synthesized response.completed SSE frame → resp.Message.Parts; image_generation tool sent only when opted in (and carries no name); llmMessageToA2A emits the generated image as a file part. Docs: runtime-engine output section, forge-yaml-schema image_generation, skill. Note: the OpenAI Responses image wire-shape follows the documented format and is fixture-tested; warrants live-API verification before production reliance.
initializ-mk
left a comment
There was a problem hiding this comment.
Self-review — verified against source ✅
Phase-5 model-generated image output; posting as a COMMENT (own PR). No blocking code issues.
Verified good:
- Opt-in, off-by-default (
image_generation: true→ClientConfig); built-in tool sent only when enabled, no function name (omitempty). Tests pin both. - Defensive parsing — reads the authoritative
response.completedoutput array (not partial-image events); base64-decode failure → byteless part → skipped. Genuine fixture test (synthesizedresponse.completedframe → decoded bytes + mime round-trip). - Projection surfaces generated media as A2A file parts; media-only responses bypass the empty-assistant placeholder (
!HasMedia()); generated media reuses the Phase-4 secure content-addressed persist.
Watch-item (author-flagged, agreed): the live-API verification. The image_generation_call wire-shape is fixture-tested against OpenAI's documented format — a live check (auth, exact result field, output-item ordering, mime) is a real pre-production gate before relying on it. The defensive design (terminal frame + graceful degradation) bounds a mismatch to "no image" rather than a crash, so it's not a code-review blocker.
Two LOW/informational notes inline; extending the #255 item-3b GC note to cover generated media.
| if out.Type == "image_generation_call" && out.Result != "" { | ||
| if raw, derr := base64.StdEncoding.DecodeString(out.Result); derr == nil && len(raw) > 0 { | ||
| delta.Parts = append(delta.Parts, llm.NewMediaContentPart(llm.ContentPartImage, llm.MediaRef{ | ||
| MimeType: "image/png", // Responses image_generation defaults to PNG |
There was a problem hiding this comment.
LOW/informational: mime hardcoded image/png. Correct today — forge sends image_generation with no output_format, which defaults to PNG. If output_format (jpeg/webp) is ever made configurable, this mime needs to follow it (and persistInboundMedia's extForMIME already handles those). Fine as-is; noting the coupling.
| // to disk and record its URI, so history stores the reference not the | ||
| // base64 (#255 Phase 5). Bytes stay inline so finalizeResponse still | ||
| // surfaces the image as a file part in the A2A response this turn. | ||
| _ = persistInboundMedia(ctx, &assistantMsg) |
There was a problem hiding this comment.
LOW/informational: persistInboundMedia now persists outbound (model-generated) media too. Two small consequences: (1) the #255 item-3b retention/GC follow-up now covers generated media as well — inbound/ accumulates both received and generated media; (2) naming/semantics — generated output lands in a dir called inbound/ via a function named persistInbound*. Functionally fine (it's the shared content-addressed media store), but a rename to persistMedia and/or a generated/ subdir would read more clearly. Cosmetic.
Part of #255. Phase 5 — model-generated image output, the last core phase. Forge can now return images the model produces as
fileparts in the A2A response, completing the multipart-output direction. Opt-in and Responses-only (Anthropic/Gemini/Bedrock don't generate images through their chat/messages APIs).What this does
llmtypes —StreamDeltagainsParts(non-text output);ClientConfiggainsEnableImageGeneration.buildRequestsends theimage_generationbuilt-in tool;readStreamreads theimage_generation_callbase64resultfrom the authoritativeresponse.completedframe (not the fragile partial-image stream events) and emits it as an imageContentPart;Chataggregatesdelta.Partsintoresp.Message.Parts.responsesTool.Nameis nowomitemptyso the built-in tool serializes without a name.llmMessageToA2Aprojects response image/documentPartsinto A2Afileparts; the assistant message's generated media is persisted like inbound media (URI in history, no base64 bloat; Bytes stay inline so the image is surfaced this turn). The empty-assistant-turn placeholder now respects a media-only response.ModelRef.ImageGeneration(yamlimage_generation: true) →ClientConfig, off by default (changes provider behavior + cost).Tests
response.completedSSE frame →resp.Message.Parts(mime + decoded bytes).image_generationtool sent only when opted in, and carries no function name.llmMessageToA2Aemits the generated image as an A2Afilepart alongside the text.go test(core runtime + llm + types + cli runtime),golangci-lint0 issues, nogo.work.sumchurn.Docs
runtime-engine.md(model-generated output section),forge-yaml-schema.md(image_generationfield),forge.mdskill (synced).I can't hit the live OpenAI Responses API from here, so the
image_generation_callwire-shape follows OpenAI's documented format and is fixture-tested (synthesized SSE). It warrants a quick live-API check (auth, exactresultfield, output-item ordering) before production reliance. The design is defensive: reading from the terminalresponse.completedoutput array is the stable signal, and any decode failure degrades gracefully (byteless part → skipped).#255 status
All core phases (0–5) are now implemented. Remaining follow-ups: OpenAI Responses
input_filedocument input, Gemini/Bedrock media, text-extraction fallback, remote-session media replay, and streaming image output (partial-image events).