Skip to content

feat(llm): image input to vision models (#255 Phase 2) - #531

Merged
initializ-mk merged 1 commit into
mainfrom
feat/multimodal-image-input
Sep 25, 2026
Merged

initializ-mk merged 1 commit into
mainfrom
feat/multimodal-image-input

Conversation

@initializ-mk

Copy link
Copy Markdown
Contributor

Part of #255 (multimodal input/output). Phase 2 — image input reaches the model. Builds on Phase 0 (#527, the reject gate) and Phase 1 (#528, ChatMessage.Parts + history references), both merged.

What this does

An inbound image file part is now forwarded to the model as native vision input when the resolved model is vision-capable, instead of being rejected at the ingest gate.

  • Capability table (forge-core/runtime/model_capabilities.go) — ModelSupportsVision (prefix map mirroring ModelContextWindows: OpenAI gpt-4o/gpt-4.1/gpt-5/o1/o3/o4, Anthropic Claude 3+, Gemini 1.5/2) + IsImageMIME/NormalizeImageMIME. Single runtime source of truth (the catalog can't describe Claude).
  • Anthropic — anthropicContentBlock gains a Source field; convertMessage serializes Parts as a block array (text + {type:image, source:{base64, media_type, data}}).
  • OpenAI — openaiMessage.Content *string → any; multimodal messages serialize as an image_url content-parts array (inline data: URLs). Gemini inherits this via the OpenAI-compat client.
  • Projection (a2aMessageToLLM) — image file parts → ChatMessage.Parts; the flattened text stays in Content as the text-of-record for the guardrail/intent scanners (PromptText still omits file bytes — images aren't text-scannable, an accepted limitation, documented).
  • Gate flip (checkInboundMedia) — images pass on a vision model; images on a text-only model reject (model_not_vision_capable); documents/video still reject (unsupported_media_type). Fails closed when no model is resolved.

Text-only messages keep Parts empty and marshal byte-identically to before (pinned by tests) — no prompt-cache prefix drift.

Tests

  • Capability + MIME matrix; Anthropic image-source block + OpenAI image_url + text-only byte-identical wire (both providers) + assistant-tool-call still omits content.
  • Projection: image→Parts (text block + image, mime normalized, bytes carried), text-only→Parts nil, non-image file skipped.
  • Gate matrix: image/vision accept, image/non-vision reject (names mime + reason), document reject, text-only pass, nil-model fail-closed. Phase 0 e2e reject test still green.
  • gofmt, go test ./... (core) + ./runtime/ (cli), golangci-lint 0 issues, make sync-knowledge, no go.work.sum churn.

Scope boundary — request-body cap unchanged (deliberate)

The inbound body cap stays at 2 MiB (both transports), so images larger than ~1.5 MiB raw currently get a 413. Per the #527 review, the raise to 32 MiB must land with the media DoS controls (per-part count/size limits, image decode-dimension bounds, a concurrency semaphore) — a flat MaxBytesReader alone isn't media DoS protection. That's the immediate follow-up PR; keeping it separate keeps this one focused on "images reach the model" and keeps the cap-raise bundled with its safeguards. Small images work today.

Not in this PR

  • Larger images + DoS controls (next).
  • Document input (Phase 3), persist-inbound-for-tools (Phase 4), model-generated image output (Phase 5).
  • Cross-turn image replay: Phase 2 sends images inline in their turn; history persistence of images (URI references) lands with Phase 4's disk persistence.

Forge now forwards inbound image file parts to the model as native vision
input, instead of rejecting them at the ingest gate. Builds on Phase 0
(reject gate) + Phase 1 (ChatMessage.Parts).

- Capability table (forge-core/runtime/model_capabilities.go): ModelSupportsVision
  (prefix map, mirrors ModelContextWindows) + IsImageMIME / NormalizeImageMIME.
- Anthropic provider: anthropicContentBlock gains a Source field;
  convertMessage serializes ChatMessage.Parts as a block array (text +
  {type:image, source:{base64,media_type,data}}).
- OpenAI provider: openaiMessage.Content string -> any; multimodal messages
  serialize as an image_url content-parts array (data: URLs). Gemini inherits
  this via the OpenAI-compat client. Text-only messages stay byte-identical.
- Projection (a2aMessageToLLM): image file parts -> ChatMessage.Parts (text-of-
  record stays in Content for the guardrail/intent scanners; PromptText still
  omits file bytes).
- Ingest gate (checkInboundMedia): images pass on a vision-capable model;
  images on a text-only model reject with reason model_not_vision_capable;
  documents/video still reject with unsupported_media_type. Fails closed when
  no model is resolved.

Tests: capability + MIME matrix; Anthropic image-source + OpenAI image_url
(and text-only byte-identical wire); projection (image->Parts, text-only nil,
non-image skipped); gate matrix (image/vision accept, image/non-vision reject,
document reject, text-only pass, nil-model fail-closed). Docs: runtime-engine
multimodal section, audit-logging reason codes, forge.md skill (synced).

@initializ-mk initializ-mk left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review — verified against source ✅

Phase-2 image input; posting as a COMMENT (own PR). No blocking issues.

Verified good:

  • Gate flip: fail-closed on nil model, images accepted only on vision-capable, non-image (doc/video) + image-on-text-model rejected loudly.
  • Capability table conservative (unknown → text-only → reject).
  • Anthropic {type:image, source:{base64, media_type, data}} + OpenAI image_url data-URL both match spec; byte-less media skipped.
  • Byte-identical text-only wire preserved (OpenAI *string→any marshals identically; Anthropic simple-text path) — test-pinned, no prompt-cache prefix drift.
  • Projection normalizes the mime (image/jpg→image/jpeg) so Anthropic accepts it; keeps Content = PromptText() as the text-of-record.
  • Body cap deliberately stays 2 MiB — the raise ships with the media DoS controls (per the #527 review), not here.

One security consideration + one accuracy note inline (both non-blocking); image-content-scanning tracked as a future item on #255.

return llm.ChatMessage{
out := llm.ChatMessage{
Role: role,
Content: msg.PromptText(),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security consideration (documented & bounded, non-blocking): image-borne prompt injection bypasses the text scanners. Content = PromptText() omits image bytes, so the guardrail/intent scanners see only text — instructions embedded in an image reach the vision model unscanned. The compensating control holds: the action-layer guardrails (PathValidator, egress enforcement, PDP) don't depend on prompt scanning, so an image can influence the model's reasoning but can't bypass what it's allowed to do. Correctly documented as an accepted limitation; tracking image-content scanning (OCR / vision moderation) as future hardening on #255. Flagging so the scanner-coverage gap is a conscious, recorded decision.

var visionCapablePrefixes = []string{
// OpenAI (chat + reasoning models with vision).
"gpt-4o", "gpt-4.1", "gpt-4-turbo", "gpt-4-vision", "gpt-5",
"o1", "o3", "o4",

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LOW (accuracy): these prefixes also match -mini variants. o1 matches o1-mini, o3 matches o3-mini — verify those actually support vision. If a -mini under one of these prefixes is text-only, an image gets sent and yields a provider error rather than forge's clean model_not_vision_capable reject. Loud failure, not silent/insecure (so LOW) — but consider excluding known text-only minis, or accept the provider-error fallback as documented behavior.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant