Skip to content

feat(llm): model-generated image output via OpenAI Responses (#255 Phase 5) - #536

Open
initializ-mk wants to merge 1 commit into
mainfrom
feat/multimodal-image-output
Open

initializ-mk wants to merge 1 commit into
mainfrom
feat/multimodal-image-output

Conversation

@initializ-mk

Copy link
Copy Markdown
Contributor

Part of #255. Phase 5 — model-generated image output, the last core phase. Forge can now return images the model produces as file parts in the A2A response, completing the multipart-output direction. Opt-in and Responses-only (Anthropic/Gemini/Bedrock don't generate images through their chat/messages APIs).

What this does

  • llm types — StreamDelta gains Parts (non-text output); ClientConfig gains EnableImageGeneration.
  • Responses provider — when enabled, buildRequest sends the image_generation built-in tool; readStream reads the image_generation_call base64 result from the authoritative response.completed frame (not the fragile partial-image stream events) and emits it as an image ContentPart; Chat aggregates delta.Parts into resp.Message.Parts. responsesTool.Name is now omitempty so the built-in tool serializes without a name.
  • Executor — llmMessageToA2A projects response image/document Parts into A2A file parts; the assistant message's generated media is persisted like inbound media (URI in history, no base64 bloat; Bytes stay inline so the image is surfaced this turn). The empty-assistant-turn placeholder now respects a media-only response.
  • Config — ModelRef.ImageGeneration (yaml image_generation: true) → ClientConfig, off by default (changes provider behavior + cost).

Tests

  • Responses parses a generated image from a synthesized response.completed SSE frame → resp.Message.Parts (mime + decoded bytes).
  • image_generation tool sent only when opted in, and carries no function name.
  • llmMessageToA2A emits the generated image as an A2A file part alongside the text.
  • gofmt, go test (core runtime + llm + types + cli runtime), golangci-lint 0 issues, no go.work.sum churn.

Docs

runtime-engine.md (model-generated output section), forge-yaml-schema.md (image_generation field), forge.md skill (synced).

⚠️ Verification caveat

I can't hit the live OpenAI Responses API from here, so the image_generation_call wire-shape follows OpenAI's documented format and is fixture-tested (synthesized SSE). It warrants a quick live-API check (auth, exact result field, output-item ordering) before production reliance. The design is defensive: reading from the terminal response.completed output array is the stable signal, and any decode failure degrades gracefully (byteless part → skipped).

#255 status

All core phases (0–5) are now implemented. Remaining follow-ups: OpenAI Responses input_file document input, Gemini/Bedrock media, text-extraction fallback, remote-session media replay, and streaming image output (partial-image events).

…ase 5)

Forge can now return model-GENERATED images as file parts in the A2A response,
completing the multipart-output direction. Opt-in and Responses-only (Anthropic/
Gemini/Bedrock don't generate images through their chat/messages APIs).

- llm: StreamDelta gains Parts (non-text output); ClientConfig gains
  EnableImageGeneration.
- Responses provider: when EnableImageGeneration is set, buildRequest sends the
  `image_generation` built-in tool; readStream reads the image_generation_call
  base64 `result` from the authoritative response.completed frame (not fragile
  partial-image events) and emits it as an image ContentPart; Chat aggregates
  delta.Parts into resp.Message.Parts. responsesTool.Name is now omitempty so
  the built-in tool serializes without a name.
- Executor: llmMessageToA2A projects response image/document Parts into A2A
  `file` parts, and the assistant message's generated media is persisted like
  inbound media (URI in history, no base64 bloat; Bytes stay inline so the image
  is surfaced this turn). The empty-assistant-turn placeholder now respects a
  media-only response.
- Config: ModelRef.ImageGeneration (yaml image_generation) → ClientConfig, off
  by default (changes provider behavior + cost).

Tests: Responses parses a generated image from a synthesized response.completed
SSE frame → resp.Message.Parts; image_generation tool sent only when opted in
(and carries no name); llmMessageToA2A emits the generated image as a file part.
Docs: runtime-engine output section, forge-yaml-schema image_generation, skill.

Note: the OpenAI Responses image wire-shape follows the documented format and is
fixture-tested; warrants live-API verification before production reliance.

@initializ-mk initializ-mk left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review — verified against source ✅

Phase-5 model-generated image output; posting as a COMMENT (own PR). No blocking code issues.

Verified good:

  • Opt-in, off-by-default (image_generation: true → ClientConfig); built-in tool sent only when enabled, no function name (omitempty). Tests pin both.
  • Defensive parsing — reads the authoritative response.completed output array (not partial-image events); base64-decode failure → byteless part → skipped. Genuine fixture test (synthesized response.completed frame → decoded bytes + mime round-trip).
  • Projection surfaces generated media as A2A file parts; media-only responses bypass the empty-assistant placeholder (!HasMedia()); generated media reuses the Phase-4 secure content-addressed persist.

Watch-item (author-flagged, agreed): the live-API verification. The image_generation_call wire-shape is fixture-tested against OpenAI's documented format — a live check (auth, exact result field, output-item ordering, mime) is a real pre-production gate before relying on it. The defensive design (terminal frame + graceful degradation) bounds a mismatch to "no image" rather than a crash, so it's not a code-review blocker.

Two LOW/informational notes inline; extending the #255 item-3b GC note to cover generated media.

if out.Type == "image_generation_call" && out.Result != "" {
if raw, derr := base64.StdEncoding.DecodeString(out.Result); derr == nil && len(raw) > 0 {
delta.Parts = append(delta.Parts, llm.NewMediaContentPart(llm.ContentPartImage, llm.MediaRef{
MimeType: "image/png", // Responses image_generation defaults to PNG

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LOW/informational: mime hardcoded image/png. Correct today — forge sends image_generation with no output_format, which defaults to PNG. If output_format (jpeg/webp) is ever made configurable, this mime needs to follow it (and persistInboundMedia's extForMIME already handles those). Fine as-is; noting the coupling.

// to disk and record its URI, so history stores the reference not the
// base64 (#255 Phase 5). Bytes stay inline so finalizeResponse still
// surfaces the image as a file part in the A2A response this turn.
_ = persistInboundMedia(ctx, &assistantMsg)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LOW/informational: persistInboundMedia now persists outbound (model-generated) media too. Two small consequences: (1) the #255 item-3b retention/GC follow-up now covers generated media as well — inbound/ accumulates both received and generated media; (2) naming/semantics — generated output lands in a dir called inbound/ via a function named persistInbound*. Functionally fine (it's the shared content-addressed media store), but a rename to persistMedia and/or a generated/ subdir would read more clearly. Cosmetic.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant