Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/skills/forge.md
Original file line number Diff line number Diff line change
Expand Up @@ -1200,7 +1200,7 @@ when OTel tracing is enabled (OTel v1 / Phase 4 / #105). Both use
| `AuditScheduleModify` | `schedule_modify` | Schedule mutated at runtime |
| `EventAuthVerify` | `auth_verify` | Inbound request authenticated (`provider`, `user_id`, `org_id`, `token_kind`; `email` when the identity carries one). **Channel invoker:** for a channel-originated request the transport credential is the loopback token (`provider:internal`/`user_id:forge-internal`, recorded truthfully) and the human sender is stamped as `channel`/`channel_user`/`channel_email` from the `X-Forge-Channel*` headers — honored only for the runtime-internal identity (same trust gate as `applyChannelOnBehalfOf`). Slack/Teams resolve `channel_email`; Telegram (numeric id) & WhatsApp (msisdn) carry `channel_user` only |
| `EventAuthFail` | `auth_fail` | Inbound request rejected (`reason`, `token_kind`) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `image_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image` source blocks / OpenAI `image_url` data URLs. Docs/video still rejected (Phase 3+). **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP (`CheckImageLimits`, header-only decode defuses bombs); ≤20 images/msg; ≤4 concurrent media requests (excess shed with 429/unavailable) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable) |
| `EventMCPServerStarted` | `mcp_server_started` | MCP server handshake succeeded |
| `EventMCPServerFailed` | `mcp_server_failed` | MCP server dial / handshake failed |
| `EventMCPServerDegraded` | `mcp_server_degraded` | MCP server in soft-fail |
Expand Down
15 changes: 11 additions & 4 deletions docs/core-concepts/runtime-engine.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,19 +33,26 @@ An inbound A2A message is a list of typed parts. `a2a.Message.PromptText()` proj

The same projection feeds the inbound guardrail and intent-alignment scanners, so the security checks see exactly what the model sees — a payload carried in a data part can't reach the LLM while bypassing them. Each data part's projected block is capped (~16KiB, rune-safe) — the cap applies identically to the scanners and the prompt, so truncation can't open a divergence.

#### Image input (multimodal)
#### Image and document input (multimodal)

An image `file` part (`image/png`, `image/jpeg`, `image/gif`, `image/webp`) is forwarded to the model as native vision input when the resolved model is **vision-capable** (`runtime.ModelSupportsVision` — OpenAI `gpt-4o`/`gpt-4.1`/`gpt-5`/`o1`/`o3`/`o4`, Anthropic Claude 3+, Gemini 1.5/2). `a2aMessageToLLM` projects such parts into `llm.ChatMessage.Parts` (the flattened text stays in `Content` as the text-of-record for the scanners), and each provider serializes them natively — Anthropic `image` source blocks, OpenAI/Gemini `image_url` data URLs. A text-only message keeps `Parts` empty and marshals byte-identically to before.
A media `file` part is forwarded to the model as native input when the resolved model supports that modality:

Media the model can't consume is **rejected loudly, never silently dropped** (the `checkInboundMedia` ingest gate): an image on a text-only model, or a document/video part (not yet supported), returns a 4xx and emits the `input_media_rejected` audit event. Note the image **bytes** themselves are not text-scannable, so guardrail/intent scanning still applies only to the text/data projection; this is an accepted limitation.
- **Images** (`image/png`, `image/jpeg`, `image/gif`, `image/webp`) → **vision-capable** models (`runtime.ModelSupportsVision` — OpenAI `gpt-4o`/`gpt-4.1`/`gpt-5`/`o1`/`o3`/`o4`, Anthropic Claude 3+, Gemini 1.5/2). Serialized as Anthropic `image` source blocks / OpenAI/Gemini `image_url` data URLs.
- **PDFs** (`application/pdf`) → **document-capable** models (`runtime.ModelSupportsPDF` — Anthropic Sonnet 3.5+, Opus 4+, Haiku 4.5+, and Fable 5; the older Claude 3.0 and 3.5 Haiku are excluded, so a PDF sent to those gets a clean reject rather than a provider error). Serialized as Anthropic `document` source blocks. OpenAI Responses `input_file`, Gemini documents, and text-extraction fallback for non-native models are deferred follow-ups.

**DoS bounds.** Because inline images raise the inbound-body cap to 32 MiB (both transports), the gate also enforces per-image and per-message limits, and a concurrency semaphore bounds how many media-bearing requests run at once — a flat body cap alone is not media DoS protection:
`a2aMessageToLLM` projects supported parts into `llm.ChatMessage.Parts` (the flattened text stays in `Content` as the text-of-record for the scanners). A text-only message keeps `Parts` empty and marshals byte-identically to before.

Media the model can't consume is **rejected loudly, never silently dropped** (the `checkInboundMedia` ingest gate): an image on a text-only model, a PDF on a non-document model, or an unsupported type (other documents/video) returns a 4xx and emits the `input_media_rejected` audit event. Note media **bytes** are not text-scannable, so guardrail/intent scanning still applies only to the text/data projection; this is an accepted limitation.

**DoS bounds.** Because inline media raises the inbound-body cap to 32 MiB (both transports), the gate also enforces per-part and per-message limits, and a concurrency semaphore bounds how many media-bearing requests run at once — a flat body cap alone is not media DoS protection:

| Bound | Limit | On breach |
|-------|-------|-----------|
| Per-image bytes | 5 MiB (`MaxImagePartBytes`) | 4xx `image_limit_exceeded` |
| Decoded dimensions | 100 000 px per side **and** 50 MP total (`MaxImagePixels`, png/jpeg/gif via header-only `DecodeConfig`; webp bounded by bytes) | 4xx `image_limit_exceeded` (defuses decompression bombs; the per-side bound also keeps the pixel product from overflowing int64) |
| Images per message | 20 (`MaxImagePartsPerMessage`) | 4xx `too_many_image_parts` |
| Per-document bytes | 32 MiB (`MaxDocumentPartBytes`), plus a `%PDF-` format sniff | 4xx `document_limit_exceeded` |
| Documents per message | 5 (`MaxDocumentPartsPerMessage`) | 4xx `too_many_document_parts` |
| Concurrent media requests | 4 (`maxConcurrentMediaRequests`) | `429`/unavailable — request is shed, not queued |

**Note for guardrail pattern authors:** parts join with **newlines** (matching what the model sees). A pattern intended to match content that may span a part boundary should use `\s+` rather than a literal space — a payload split across two text parts joins as `…end\nstart…`.
Expand Down
2 changes: 1 addition & 1 deletion docs/security/audit-logging.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ All runtime security events are emitted as structured NDJSON to stderr with corr
| `mcp_auth_resolved` | The parked call's consent arrived (or the wait was canceled) and it resumed (#330). Carries `server`, `subject`, `wait_ms`. Emitted **once**, attributed to the parked invocation (#366). |
| `mcp_auth_timeout` | No consent within the window; the parked MCP call fails `no_token` (#330). Carries `server`, `subject`, `wait_ms`, `decision`. |
| `auth_fail` | Inbound request rejected (with `reason`, `token_kind`). No `task_id` (none is ever created), but carries `workflow_execution_id` when the request had the execution header — so a rejected request is still attributable to its workflow run (#278). |
| `input_media_rejected` | An inbound `tasks/send`/`sendSubscribe` message carried `file` parts the runtime cannot forward to the model, so the request was rejected with a 4xx (JSON-RPC invalid-params / HTTP 400) instead of silently dropping the attachment and returning a plausible answer that ignored it (#255). Carries `fields.dropped` (`["file:<mimeType>", …]`), `fields.count`, and `fields.reason` — `model_not_vision_capable` (image sent to a text-only model), `unsupported_media_type` (documents/video — not yet accepted), `too_many_image_parts` (over the per-message image limit), or `image_limit_exceeded` (an image over the per-image byte or pixel bound). **Images** (`png`/`jpeg`/`gif`/`webp`) within those bounds, sent to a **vision-capable model**, are NOT rejected: they are projected into the model request as inline vision input and do not emit this event. A separate load-shedding path returns an unavailable/`429` (no audit event) when the runner is already at its concurrent-media-request cap. |
| `input_media_rejected` | An inbound `tasks/send`/`sendSubscribe` message carried `file` parts the runtime cannot forward to the model, so the request was rejected with a 4xx (JSON-RPC invalid-params / HTTP 400) instead of silently dropping the attachment and returning a plausible answer that ignored it (#255). Carries `fields.dropped` (`["file:<mimeType>", …]`), `fields.count`, and `fields.reason` — `model_not_vision_capable` (image on a text-only model), `model_not_document_capable` (PDF on a model without native document support), `unsupported_media_type` (a non-PDF document, video, … — not yet accepted), `too_many_image_parts` / `too_many_document_parts` (over the per-message count), or `image_limit_exceeded` / `document_limit_exceeded` (a part over its byte/pixel bound or failing its format sniff). **Images** (`png`/`jpeg`/`gif`/`webp`) on a **vision-capable model**, and **PDFs** on a **document-capable model** (Anthropic Claude 3.5+), within those bounds, are NOT rejected: they are projected into the model request as native vision / document blocks and do not emit this event. A separate load-shedding path returns an unavailable/`429` (no audit event) when the runner is already at its concurrent-media-request cap. |
| `agent_card_published` | Agent Card finalized at startup or hot-reload (with `name`, `version`, `protocol_version`, `url`, `skill_count`, `capabilities`, `security_schemes`, `card_size_bytes`, `card_sha256`). See [Agent Card reference](../reference/a2a-agent-card.md). |
| `policy_loaded` | One per non-empty policy layer at startup (system / user / workspace). Carries `fields.layer`, `source` (file path), deny-list size counts, and max bounds. See [Platform Policy](platform-policy.md). |
| `policy_violation_at_build_time` | One per violation when `forge.yaml` conflicts with any policy layer. Agent refuses to start. Carries `fields.violation_kind` / `offending_value` / `forge_yaml_field` plus `layer` + `source` identifying the enforcing file. See [Platform Policy](platform-policy.md). |
Expand Down
2 changes: 1 addition & 1 deletion forge-cli/internal/surface/knowledge/forge.md
Original file line number Diff line number Diff line change
Expand Up @@ -1200,7 +1200,7 @@ when OTel tracing is enabled (OTel v1 / Phase 4 / #105). Both use
| `AuditScheduleModify` | `schedule_modify` | Schedule mutated at runtime |
| `EventAuthVerify` | `auth_verify` | Inbound request authenticated (`provider`, `user_id`, `org_id`, `token_kind`; `email` when the identity carries one). **Channel invoker:** for a channel-originated request the transport credential is the loopback token (`provider:internal`/`user_id:forge-internal`, recorded truthfully) and the human sender is stamped as `channel`/`channel_user`/`channel_email` from the `X-Forge-Channel*` headers — honored only for the runtime-internal identity (same trust gate as `applyChannelOnBehalfOf`). Slack/Teams resolve `channel_email`; Telegram (numeric id) & WhatsApp (msisdn) carry `channel_user` only |
| `EventAuthFail` | `auth_fail` | Inbound request rejected (`reason`, `token_kind`) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `image_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image` source blocks / OpenAI `image_url` data URLs. Docs/video still rejected (Phase 3+). **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP (`CheckImageLimits`, header-only decode defuses bombs); ≤20 images/msg; ≤4 concurrent media requests (excess shed with 429/unavailable) |
| `AuditInputMediaRejected` | `input_media_rejected` | Inbound `file` parts the runtime can't forward to the model → rejected 4xx, not silently dropped (#255). Fields: `dropped` (`["file:<mime>"]`), `count`, `reason` (`model_not_vision_capable` \| `model_not_document_capable` \| `unsupported_media_type` \| `too_many_image_parts` \| `too_many_document_parts` \| `image_limit_exceeded` \| `document_limit_exceeded`). Gate: `Runner.checkInboundMedia`. **Images** (png/jpeg/gif/webp) on a **vision model** (`coreruntime.ModelSupportsVision`) and **PDFs** on a **doc model** (`ModelSupportsPDF` = Anthropic Sonnet 3.5+/Opus 4+/Haiku 4.5+/Fable 5; Claude 3.0 & 3.5-Haiku excluded, fail-closed), within the DoS bounds, are NOT rejected — `a2aMessageToLLM` projects them into `llm.ChatMessage.Parts` → Anthropic `image`/`document` source blocks, OpenAI `image_url` data URLs. Non-PDF docs/video still rejected; OpenAI-Responses PDF + extraction fallback are follow-ups. **DoS controls** (`media_limits.go` + `mediaSem`): body cap 32 MiB both transports; per-image ≤5 MiB & ≤50 MP/≤100k-per-side (`CheckImageLimits`); per-PDF ≤32 MiB + `%PDF-` sniff (`CheckDocumentLimits`); ≤20 images & ≤5 docs/msg; ≤4 concurrent media requests (excess shed with 429/unavailable) |
| `EventMCPServerStarted` | `mcp_server_started` | MCP server handshake succeeded |
| `EventMCPServerFailed` | `mcp_server_failed` | MCP server dial / handshake failed |
| `EventMCPServerDegraded` | `mcp_server_degraded` | MCP server in soft-fail |
Expand Down
33 changes: 28 additions & 5 deletions forge-cli/runtime/media_gate_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -66,11 +66,34 @@ func TestCheckInboundMedia_Phase2(t *testing.T) {
}
})

t.Run("document rejected even on vision model", func(t *testing.T) {
r := runnerWithModel("gpt-4o")
got := r.checkInboundMedia(ctx, fileMsg("application/pdf", []byte("%PDF-1.7")), nil)
if got == "" || !strings.Contains(got, "application/pdf") {
t.Errorf("a document part must still be rejected in Phase 2; got %q", got)
t.Run("pdf accepted on a document-capable model", func(t *testing.T) {
r := runnerWithModel("claude-sonnet-5")
if got := r.checkInboundMedia(ctx, fileMsg("application/pdf", []byte("%PDF-1.7\nbody")), nil); got != "" {
t.Errorf("pdf on a document-capable model must be accepted, got reject: %q", got)
}
})

t.Run("pdf rejected on a non-document model", func(t *testing.T) {
r := runnerWithModel("gpt-4o") // vision-capable but not PDF-capable
got := r.checkInboundMedia(ctx, fileMsg("application/pdf", []byte("%PDF-1.7\nbody")), nil)
if got == "" || !strings.Contains(got, "application/pdf") || !strings.Contains(got, "does not support PDF") {
t.Errorf("pdf on a non-document model must be rejected with a clear reason; got %q", got)
}
})

t.Run("mislabeled pdf rejected", func(t *testing.T) {
r := runnerWithModel("claude-sonnet-5")
got := r.checkInboundMedia(ctx, fileMsg("application/pdf", []byte("PK\x03\x04 zip")), nil)
if got == "" {
t.Error("a non-PDF payload labeled application/pdf must be rejected")
}
})

t.Run("unsupported document type rejected", func(t *testing.T) {
r := runnerWithModel("claude-sonnet-5")
got := r.checkInboundMedia(ctx, fileMsg("application/vnd.ms-excel", []byte("x")), nil)
if got == "" {
t.Error("a non-PDF document type must be rejected (only PDF supported)")
}
})

Expand Down
Loading
Loading