feat(llm): image input to vision models (#255 Phase 2) - #531
Conversation
Forge now forwards inbound image file parts to the model as native vision
input, instead of rejecting them at the ingest gate. Builds on Phase 0
(reject gate) + Phase 1 (ChatMessage.Parts).
- Capability table (forge-core/runtime/model_capabilities.go): ModelSupportsVision
(prefix map, mirrors ModelContextWindows) + IsImageMIME / NormalizeImageMIME.
- Anthropic provider: anthropicContentBlock gains a Source field;
convertMessage serializes ChatMessage.Parts as a block array (text +
{type:image, source:{base64,media_type,data}}).
- OpenAI provider: openaiMessage.Content string -> any; multimodal messages
serialize as an image_url content-parts array (data: URLs). Gemini inherits
this via the OpenAI-compat client. Text-only messages stay byte-identical.
- Projection (a2aMessageToLLM): image file parts -> ChatMessage.Parts (text-of-
record stays in Content for the guardrail/intent scanners; PromptText still
omits file bytes).
- Ingest gate (checkInboundMedia): images pass on a vision-capable model;
images on a text-only model reject with reason model_not_vision_capable;
documents/video still reject with unsupported_media_type. Fails closed when
no model is resolved.
Tests: capability + MIME matrix; Anthropic image-source + OpenAI image_url
(and text-only byte-identical wire); projection (image->Parts, text-only nil,
non-image skipped); gate matrix (image/vision accept, image/non-vision reject,
document reject, text-only pass, nil-model fail-closed). Docs: runtime-engine
multimodal section, audit-logging reason codes, forge.md skill (synced).
initializ-mk
left a comment
There was a problem hiding this comment.
Self-review — verified against source ✅
Phase-2 image input; posting as a COMMENT (own PR). No blocking issues.
Verified good:
- Gate flip: fail-closed on nil model, images accepted only on vision-capable, non-image (doc/video) + image-on-text-model rejected loudly.
- Capability table conservative (unknown → text-only → reject).
- Anthropic
{type:image, source:{base64, media_type, data}}+ OpenAIimage_urldata-URL both match spec; byte-less media skipped. - Byte-identical text-only wire preserved (OpenAI
*string→anymarshals identically; Anthropic simple-text path) — test-pinned, no prompt-cache prefix drift. - Projection normalizes the mime (
image/jpg→image/jpeg) so Anthropic accepts it; keepsContent = PromptText()as the text-of-record. - Body cap deliberately stays 2 MiB — the raise ships with the media DoS controls (per the #527 review), not here.
One security consideration + one accuracy note inline (both non-blocking); image-content-scanning tracked as a future item on #255.
| return llm.ChatMessage{ | ||
| out := llm.ChatMessage{ | ||
| Role: role, | ||
| Content: msg.PromptText(), |
There was a problem hiding this comment.
Security consideration (documented & bounded, non-blocking): image-borne prompt injection bypasses the text scanners. Content = PromptText() omits image bytes, so the guardrail/intent scanners see only text — instructions embedded in an image reach the vision model unscanned. The compensating control holds: the action-layer guardrails (PathValidator, egress enforcement, PDP) don't depend on prompt scanning, so an image can influence the model's reasoning but can't bypass what it's allowed to do. Correctly documented as an accepted limitation; tracking image-content scanning (OCR / vision moderation) as future hardening on #255. Flagging so the scanner-coverage gap is a conscious, recorded decision.
| var visionCapablePrefixes = []string{ | ||
| // OpenAI (chat + reasoning models with vision). | ||
| "gpt-4o", "gpt-4.1", "gpt-4-turbo", "gpt-4-vision", "gpt-5", | ||
| "o1", "o3", "o4", |
There was a problem hiding this comment.
LOW (accuracy): these prefixes also match -mini variants. o1 matches o1-mini, o3 matches o3-mini — verify those actually support vision. If a -mini under one of these prefixes is text-only, an image gets sent and yields a provider error rather than forge's clean model_not_vision_capable reject. Loud failure, not silent/insecure (so LOW) — but consider excluding known text-only minis, or accept the provider-error fallback as documented behavior.
Part of #255 (multimodal input/output). Phase 2 — image input reaches the model. Builds on Phase 0 (#527, the reject gate) and Phase 1 (#528,
ChatMessage.Parts+ history references), both merged.What this does
An inbound image
filepart is now forwarded to the model as native vision input when the resolved model is vision-capable, instead of being rejected at the ingest gate.forge-core/runtime/model_capabilities.go) —ModelSupportsVision(prefix map mirroringModelContextWindows: OpenAIgpt-4o/gpt-4.1/gpt-5/o1/o3/o4, Anthropic Claude 3+, Gemini 1.5/2) +IsImageMIME/NormalizeImageMIME. Single runtime source of truth (the catalog can't describe Claude).anthropicContentBlockgains aSourcefield;convertMessageserializesPartsas a block array (text+{type:image, source:{base64, media_type, data}}).openaiMessage.Content*string→any; multimodal messages serialize as animage_urlcontent-parts array (inlinedata:URLs). Gemini inherits this via the OpenAI-compat client.a2aMessageToLLM) — image file parts →ChatMessage.Parts; the flattened text stays inContentas the text-of-record for the guardrail/intent scanners (PromptTextstill omits file bytes — images aren't text-scannable, an accepted limitation, documented).checkInboundMedia) — images pass on a vision model; images on a text-only model reject (model_not_vision_capable); documents/video still reject (unsupported_media_type). Fails closed when no model is resolved.Text-only messages keep
Partsempty and marshal byte-identically to before (pinned by tests) — no prompt-cache prefix drift.Tests
image_url+ text-only byte-identical wire (both providers) + assistant-tool-call still omits content.Parts(text block + image, mime normalized, bytes carried), text-only→Partsnil, non-image file skipped.go test ./...(core) +./runtime/(cli),golangci-lint0 issues,make sync-knowledge, nogo.work.sumchurn.Scope boundary — request-body cap unchanged (deliberate)
The inbound body cap stays at 2 MiB (both transports), so images larger than ~1.5 MiB raw currently get a 413. Per the #527 review, the raise to 32 MiB must land with the media DoS controls (per-part count/size limits, image decode-dimension bounds, a concurrency semaphore) — a flat
MaxBytesReaderalone isn't media DoS protection. That's the immediate follow-up PR; keeping it separate keeps this one focused on "images reach the model" and keeps the cap-raise bundled with its safeguards. Small images work today.Not in this PR