Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/skills/forge.md
Original file line number Diff line number Diff line change
Expand Up @@ -1200,7 +1200,7 @@ when OTel tracing is enabled (OTel v1 / Phase 4 / #105). Both use
| `AuditScheduleModify` | `schedule_modify` | Schedule mutated at runtime |
| `EventAuthVerify` | `auth_verify` | Inbound request authenticated (`provider`, `user_id`, `org_id`, `token_kind`; `email` when the identity carries one). **Channel invoker:** for a channel-originated request the transport credential is the loopback token (`provider:internal`/`user_id:forge-internal`, recorded truthfully) and the human sender is stamped as `channel`/`channel_user`/`channel_email` from the `X-Forge-Channel*` headers — honored only for the runtime-internal identity (same trust gate as `applyChannelOnBehalfOf`). Slack/Teams resolve `channel_email`; Telegram (numeric id) & WhatsApp (msisdn) carry `channel_user` only |
| `EventAuthFail` | `auth_fail` | Inbound request rejected (`reason`, `token_kind`) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable). **Persistence/replay** (`persistInboundMedia` + `RehydrateMedia`): inbound media written to `.forge/files/inbound/<sha256>.<ext>`, path stored as `MediaRef.URI`; history keeps URI-only (`Bytes` json:"-"), rehydrated per turn so multi-turn convos replay media without base64 bloat |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable). **Persistence/replay** (`persistMedia` + `RehydrateMedia`): media written to `.forge/files/{inbound,generated}/<sha256>.<ext>`, path stored as `MediaRef.URI`; history keeps URI-only (`Bytes` json:"-"), rehydrated per turn so multi-turn convos replay media without base64 bloat. **Model-generated image OUTPUT** (#255 Phase 5): opt-in `image_generation: true` on an `openai-responses` model → sends the `image_generation` built-in tool; the `image_generation_call` base64 `result` (from `response.completed`) → `resp.Message.Parts` image → `llmMessageToA2A` emits a response `file` part (persisted like inbound). Responses-only (Anthropic/Gemini/Bedrock don't generate images via chat) |
| `EventMCPServerStarted` | `mcp_server_started` | MCP server handshake succeeded |
| `EventMCPServerFailed` | `mcp_server_failed` | MCP server dial / handshake failed |
| `EventMCPServerDegraded` | `mcp_server_degraded` | MCP server in soft-fail |
Expand Down
6 changes: 4 additions & 2 deletions docs/core-concepts/runtime-engine.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,9 @@ A media `file` part is forwarded to the model as native input when the resolved

`a2aMessageToLLM` projects supported parts into `llm.ChatMessage.Parts` (the flattened text stays in `Content` as the text-of-record for the scanners). A text-only message keeps `Parts` empty and marshals byte-identically to before.

**Persistence & cross-turn replay.** Inbound media is written to `.forge/files/inbound/<sha256>.<ext>` (`persistInboundMedia`) and the on-disk path is recorded as the part's `MediaRef.URI` — content-addressed, so identical uploads dedup and re-writes are idempotent, and the files are available on disk for tools. The bytes stay inline for the current turn's request; because `MediaRef.Bytes` is `json:"-"`, session history persists only the URI (never base64). On a later turn, `RehydrateMedia` reloads the bytes from the URI before the request is built, so a multi-turn conversation keeps seeing earlier images/PDFs without bloating the session file. A file that can't be reloaded (deleted, or an absolute path from another host — remote/distributed session replay is a follow-up) is left byteless and skipped by the provider serializers, degrading to text rather than failing the turn. Persistence and rehydration are best-effort: a failure never fails the turn (media is still fed inline the turn it arrives).
**Model-generated image output.** A model can also *produce* an image, which forge returns as a `file` part in the A2A response. For OpenAI **Responses** (`provider: openai-responses`), set `image_generation: true` on the model (`ModelRef.ImageGeneration` → `ClientConfig.EnableImageGeneration`); the client then sends the `image_generation` built-in tool, and the base64 `result` from the `image_generation_call` output (read from the authoritative `response.completed` frame) is decoded into an image `ContentPart` on `resp.Message.Parts`. `llmMessageToA2A` projects those into response `file` parts, and the generated image is persisted to `.forge/files/generated/` (URI in history, no base64 bloat). Off by default (it changes provider behavior and cost). Anthropic/Gemini/Bedrock don't generate images through their chat/messages APIs, so this is Responses-only today.

**Persistence & cross-turn replay.** Media is written to `.forge/files/<inbound|generated>/<sha256>.<ext>` (`persistMedia`) — received uploads under `inbound/`, model-generated output under `generated/`. The on-disk path is recorded as the part's `MediaRef.URI` — content-addressed, so identical uploads dedup and re-writes are idempotent, and the files are available on disk for tools. The bytes stay inline for the current turn's request; because `MediaRef.Bytes` is `json:"-"`, session history persists only the URI (never base64). On a later turn, `RehydrateMedia` reloads the bytes from the URI before the request is built, so a multi-turn conversation keeps seeing earlier images/PDFs without bloating the session file. A file that can't be reloaded (deleted, or an absolute path from another host — remote/distributed session replay is a follow-up) is left byteless and skipped by the provider serializers, degrading to text rather than failing the turn. Persistence and rehydration are best-effort: a failure never fails the turn (media is still fed inline the turn it arrives).

Media the model can't consume is **rejected loudly, never silently dropped** (the `checkInboundMedia` ingest gate): an image on a text-only model, a PDF on a non-document model, or an unsupported type (other documents/video) returns a 4xx and emits the `input_media_rejected` audit event. Note media **bytes** are not text-scannable, so guardrail/intent scanning still applies only to the text/data projection; this is an accepted limitation.

Expand Down Expand Up @@ -346,7 +348,7 @@ docker run -e KUBECONFIG="$(cat ~/.kube/config)" my-agent

## File Output Directory

The runtime configures a `FilesDir` for tool-generated files (e.g., from `file_create`). This directory defaults to `<WorkDir>/.forge/files/` and is injected into the execution context so tools can write files that other tools can reference by path. Inbound media (uploaded images/PDFs) is persisted under `<FilesDir>/inbound/` — see [Image and document input](#image-and-document-input-multimodal).
The runtime configures a `FilesDir` for tool-generated files (e.g., from `file_create`). This directory defaults to `<WorkDir>/.forge/files/` and is injected into the execution context so tools can write files that other tools can reference by path. Inbound uploads are persisted under `<FilesDir>/inbound/` and model-generated media under `<FilesDir>/generated/` — see [Image and document input](#image-and-document-input-multimodal).

```
<WorkDir>/
Expand Down
1 change: 1 addition & 0 deletions docs/reference/forge-yaml-schema.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ model:
aws_region: "" # Required when auth_scheme: aws_sigv4 — issue #202
auth_header_name: "" # apikey_header[_only] custom header name; default "apikey" — issue #302
disable_store: false # openai-responses only: send store=false so OpenAI doesn't retain responses — issue #383
image_generation: false # openai-responses only: enable the image_generation tool so the model can return images as file parts — issue #255
fallbacks: # Fallback providers (optional)
- provider: "anthropic"
name: "claude-sonnet-4-20250514"
Expand Down
2 changes: 1 addition & 1 deletion forge-cli/internal/surface/knowledge/forge.md
Original file line number Diff line number Diff line change
Expand Up @@ -1200,7 +1200,7 @@ when OTel tracing is enabled (OTel v1 / Phase 4 / #105). Both use
| `AuditScheduleModify` | `schedule_modify` | Schedule mutated at runtime |
| `EventAuthVerify` | `auth_verify` | Inbound request authenticated (`provider`, `user_id`, `org_id`, `token_kind`; `email` when the identity carries one). **Channel invoker:** for a channel-originated request the transport credential is the loopback token (`provider:internal`/`user_id:forge-internal`, recorded truthfully) and the human sender is stamped as `channel`/`channel_user`/`channel_email` from the `X-Forge-Channel*` headers — honored only for the runtime-internal identity (same trust gate as `applyChannelOnBehalfOf`). Slack/Teams resolve `channel_email`; Telegram (numeric id) & WhatsApp (msisdn) carry `channel_user` only |
| `EventAuthFail` | `auth_fail` | Inbound request rejected (`reason`, `token_kind`) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable). **Persistence/replay** (`persistInboundMedia` + `RehydrateMedia`): inbound media written to `.forge/files/inbound/<sha256>.<ext>`, path stored as `MediaRef.URI`; history keeps URI-only (`Bytes` json:"-"), rehydrated per turn so multi-turn convos replay media without base64 bloat |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable). **Persistence/replay** (`persistMedia` + `RehydrateMedia`): media written to `.forge/files/{inbound,generated}/<sha256>.<ext>`, path stored as `MediaRef.URI`; history keeps URI-only (`Bytes` json:"-"), rehydrated per turn so multi-turn convos replay media without base64 bloat. **Model-generated image OUTPUT** (#255 Phase 5): opt-in `image_generation: true` on an `openai-responses` model → sends the `image_generation` built-in tool; the `image_generation_call` base64 `result` (from `response.completed`) → `resp.Message.Parts` image → `llmMessageToA2A` emits a response `file` part (persisted like inbound). Responses-only (Anthropic/Gemini/Bedrock don't generate images via chat) |
| `EventMCPServerStarted` | `mcp_server_started` | MCP server handshake succeeded |
| `EventMCPServerFailed` | `mcp_server_failed` | MCP server dial / handshake failed |
| `EventMCPServerDegraded` | `mcp_server_degraded` | MCP server in soft-fail |
Expand Down
7 changes: 7 additions & 0 deletions forge-core/llm/client.go
Original file line number Diff line number Diff line change
Expand Up @@ -136,4 +136,11 @@ type ClientConfig struct {
// unset so the API applies its own default (#383). The ChatGPT OAuth
// path forces this true regardless (Codex backend requires it).
DisableStore bool

// EnableImageGeneration opts an OpenAI Responses request into the
// `image_generation` built-in tool, letting the model emit images that
// forge surfaces as `file` parts in the A2A response (#255). Off by
// default; only the openai-responses client honors it (other clients
// ignore the field). Opt-in because it changes provider behavior and cost.
EnableImageGeneration bool
}
30 changes: 29 additions & 1 deletion forge-core/llm/providers/responses.go
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ import (
"bufio"
"bytes"
"context"
"encoding/base64"
"encoding/json"
"fmt"
"io"
Expand All @@ -26,6 +27,7 @@ type ResponsesClient struct {
authHeaderName string
client *http.Client
disableStore bool // set store=false in requests (required for ChatGPT Codex backend)
enableImageGen bool // send the image_generation built-in tool (#255)
}

// NewResponsesClient creates a new Responses API client.
Expand Down Expand Up @@ -57,6 +59,7 @@ func NewResponsesClient(cfg llm.ClientConfig) *ResponsesClient {
authScheme: cfg.AuthScheme,
authHeaderName: cfg.AuthHeaderName,
disableStore: cfg.DisableStore,
enableImageGen: cfg.EnableImageGeneration,
client: httpClient,
}
}
Expand Down Expand Up @@ -84,6 +87,9 @@ func (c *ResponsesClient) Chat(ctx context.Context, req *llm.ChatRequest) (*llm.
if delta.Content != "" {
result.Message.Content += delta.Content
}
if len(delta.Parts) > 0 {
result.Message.Parts = append(result.Message.Parts, delta.Parts...)
}
for _, tc := range delta.ToolCalls {
existing, ok := toolCallMap[tc.ID]
if !ok {
Expand Down Expand Up @@ -207,7 +213,7 @@ type responsesInput struct {
// responsesTool is the Responses API tool format (flat, not nested under "function").
type responsesTool struct {
Type string `json:"type"`
Name string `json:"name"`
Name string `json:"name,omitempty"` // omitted for built-in tools (e.g. image_generation)
Description string `json:"description,omitempty"`
Parameters json.RawMessage `json:"parameters,omitempty"`
}
Expand Down Expand Up @@ -274,6 +280,11 @@ func (c *ResponsesClient) buildRequest(req *llm.ChatRequest, stream bool) respon
Parameters: t.Function.Parameters,
})
}
// Opt-in: let the model emit images via the image_generation built-in tool
// (#255). Surfaced back as file parts in the A2A response.
if c.enableImageGen {
tools = append(tools, responsesTool{Type: "image_generation"})
}

// The Responses API requires the instructions field. If no system
// message was provided (e.g. summarization calls), use a minimal default
Expand Down Expand Up @@ -319,6 +330,9 @@ type responsesOutput struct {
CallID string `json:"call_id,omitempty"`
Name string `json:"name,omitempty"`
Arguments string `json:"arguments,omitempty"`

// For image_generation_call outputs: base64-encoded image bytes (#255).
Result string `json:"result,omitempty"`
}

type responsesContentPart struct {
Expand Down Expand Up @@ -473,6 +487,20 @@ func (c *ResponsesClient) readStream(r io.Reader, ch chan<- llm.StreamDelta) {
TotalTokens: ev.Response.Usage.TotalTokens,
}
}
// Surface any model-generated images from the final output array as
// content parts (#255). The completed event carries the authoritative
// output[], so we read the image_generation_call `result` (base64)
// here rather than reassembling partial-image stream events.
for _, out := range ev.Response.Output {
if out.Type == "image_generation_call" && out.Result != "" {
if raw, derr := base64.StdEncoding.DecodeString(out.Result); derr == nil && len(raw) > 0 {
delta.Parts = append(delta.Parts, llm.NewMediaContentPart(llm.ContentPartImage, llm.MediaRef{
MimeType: "image/png", // Responses image_generation defaults to PNG

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LOW/informational: mime hardcoded image/png. Correct today — forge sends image_generation with no output_format, which defaults to PNG. If output_format (jpeg/webp) is ever made configurable, this mime needs to follow it (and persistInboundMedia's extForMIME already handles those). Fine as-is; noting the coupling.

Bytes: raw,
}))
}
}
}
// Determine finish reason from output
for _, out := range ev.Response.Output {
if out.Type == "function_call" {
Expand Down
Loading
Loading