You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 97342d6
Browse filesBrowse the repository at this point in the historyBrowse files
feat(client): grounded judge context and per-judge diagnostics
Callers can now hand judges evidence about what actually happened during a
request. config() takes a lazy judge_context callback because the value does
not exist yet when config() is called: the caller's tools fill it while the
primary handler runs. The SDK resolves it exactly once, right after the
primary handler succeeds and before output-format parsing, so every request
with a callback freezes the same snapshot whether or not a judge is sampled,
and whether or not skip_judges is set, on invoke() and stream() alike.
A resolved context must be acyclic JSON of at most 64 KiB encoded. It is
returned unchanged on ProviderResponse.judge_context and reaches a judge only
through that judge's message_history variable, between the
UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN and UNTRUSTED_ACTUATOR_EVIDENCE_END lines.
It never reaches the primary model, the track data, or a span. The block is
one more part of judge_scoring.build_message_history, after the answer and
before the formatting instructions, so the inline path and run_judge on a
JudgeTask still show a judge the same conversation, trajectory included. With
no callback the judge prompt is byte-identical to before.
Judges are now isolated from each other and from the primary result. Config
lookup, provider call, parse and tracking each sit behind their own boundary,
bounded by judge_timeout_ms, and a failure produces one JudgeDiagnostic
instead of discarding work that already succeeded. A judge no registered
handler can serve, and a malformed judgeConfiguration block or entry, are
reported the same way rather than raised or dropped. A judge that beats the
clock and then fails to track keeps its result. A judge that misses the clock
has its late completion consumed silently, so it can neither mutate results
nor emit the score metric. Diagnostics carry codes only, never exception text,
and the same codes are shared with the TypeScript SDK.
graph().invoke() and the final graph().stream() done event forward the
graph-level judge's diagnostics. Graph nodes do not receive a caller judge
context in v1.
run_judges and build_judge_tasks now return result objects carrying both the
results and the diagnostics, which is a breaking change for direct callers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BREAKING CHANGE: run_judges and build_judge_tasks now return a result
object (RunJudgesResult with judge_results/judge_diagnostics, and
BuildJudgeTasksResult with judge_tasks/judge_diagnostics/judge_context)
instead of the bare value. The per-entry shape inside judge_results is
unchanged. Both are exported from the package.
|`track_data`|`TrackData`| Tracking payload from this invocation (run ID, config key, etc.). Carried inside each `JudgeTask` so background judge results are attributed to the originating request. |
212
212
|`judge_results`|`dict[str, JudgeResult]?`| Results from inline judge evaluations. Present when `skip_judges=False` (default) and judges ran. |
213
-
|`judge_tasks`|`list[JudgeTask]?`| Pre-packaged judge tasks. Present (as a list) when `skip_judges=True`. Each task is a serialisable dataclass ready to pass to a background thread running `run_judge(task, handlers)`. `None` when `skip_judges=False`. |
213
+
|`judge_tasks`|`list[JudgeTask]?`| Pre-packaged judge tasks. Present (as a list) when `skip_judges=True`. Each task is a serialisable dataclass ready to pass to a background thread running `run_judge(task, handlers)`, and carries the resolved `judge_context` so the worker injects the identical evidence block. `None` when `skip_judges=False`. |
214
+
|`judge_context`|`JsonValue?`| The value the `judge_context` callback returned, unchanged. `None` when no callback was given, it returned `None` (no evidence), or it failed validation. |
215
+
|`judge_diagnostics`|`list[JudgeDiagnostic]?`| Why judges were skipped or failed. `None` when nothing went wrong (never `[]`). |
216
+
217
+
#### `JsonValue`
218
+
219
+
`None | bool | int | float | str | list[JsonValue] | dict[str, JsonValue]` — a value that survives a `json.dumps` / `json.loads` round trip unchanged.
220
+
221
+
#### `JudgeDiagnostic`
222
+
223
+
One reason a judge produced no result, or a partial one. The strings are shared with the TypeScript SDK on the wire; never rename them. A diagnostic carries no raw exception text.
224
+
225
+
| Field | Type | Description |
226
+
|---|---|---|
227
+
|`status`|`"skipped" \| "failed"`| Skipped means the judge never ran. |
228
+
|`stage`|`"context" \| "config" \| "provider" \| "parse" \| "track" \| "timeout"`| Where it went wrong. |
|`judge_key`|`str?`| The judge this concerns. Absent for context-stage diagnostics and for a malformed `judgeConfiguration` entry with no string key. |
231
+
232
+
#### Stream `done` event
233
+
234
+
The final event of `config().stream()`: `{"type": "done", "response": str, "usage": UsageDict, "judge_context": JsonValue \| None, "judge_results": dict[str, JudgeResult] \| None, "judge_diagnostics": list[JudgeDiagnostic] \| None}`. Exactly one is yielded, after every `chunk`. Empty results and diagnostics are `None`.
214
235
215
236
#### `ProviderGraphResponse`
216
237
@@ -221,6 +242,7 @@ The value returned by `graph().invoke()`. A dataclass with attribute access.
221
242
|`response`|`str`| The final text output (from the last node executed). |
222
243
|`usage`|`UsageDict`| Aggregate token counts across all nodes. |
223
244
|`judge_results`|`dict[str, JudgeResult]?`| Results from a graph-level judge, if configured. |
245
+
|`judge_diagnostics`|`list[JudgeDiagnostic]?`| Diagnostics from the graph-level judge. Graph nodes do not receive a caller judge context in v1. |
224
246
225
247
#### `ConfigArgs`
226
248
@@ -232,6 +254,8 @@ Arguments accepted by `config()`.
232
254
|`handler`|`ProviderHandler \| list[ProviderHandler]`? | One handler or an ordered list of handlers. Routing selects the match by provider + mode. |
233
255
|`tool_handlers`|`dict[str, Callable \| NativeTool]?`| Map of tool name → implementation function (or `NativeTool` sentinel). |
234
256
|`registry`|`Registry?`| Registry to source handlers and tools from. Local `handler`/`tool_handlers` take precedence. |
257
+
|`judge_context`|`Callable[[], JsonValue \| Awaitable[JsonValue]]?`| Lazily resolves JSON-safe evidence for the judges. Lazy on purpose: the value does not exist when `config()` is called, and the caller's tools fill it while the primary handler runs. |
|`skip_judges`|`bool`? | When `True`, `invoke()` does not run judges inline. Instead it returns `judge_tasks: list[JudgeTask]` — pre-packaged tasks ready for background thread execution via `run_judge(task, handlers)`. Default: `False`. |
236
260
237
261
#### `TrackData`
@@ -362,10 +386,11 @@ Returns a `ConfigInstance` with:
362
386
2. Selects the handler by matching on `[config.provider.name, normalized mode]`. Selection priority: (a) exact provider match, (b) wildcard `['*', mode]` fallback for multi-provider adapters (e.g. LangChain). Raises if no matching handler is found.
363
387
3. Invokes the selected handler with the config, user input, tool handlers, variables, and history. The `context` passed to `.invoke()` is automatically merged into `variables` under the key `ldContext`, so templates can reference `{{ldContext.key}}`, `{{ldContext.email}}`, etc. If `history` is provided, it is passed to the handler as the 5th positional argument — messages-mode handlers splice it into the messages array; agent-mode handlers append it to the system prompt.
-**Default (`skip_judges=False`):** runs each configured judge inline at its `samplingRate`. Results are returned in `ProviderResponse.judge_results`.
367
-
-**`skip_judges=True`:** builds serialisable `JudgeTask` objects for each judge (no AI calls). Returns them in `ProviderResponse.judge_tasks`. Pass each task to a background thread running `run_judge(task, handlers)`.
368
-
6. Returns a `ProviderResponse` (always includes `response`, `usage`, and `track_data`).
389
+
5. If a `judge_context` callback was given, resolves it exactly once — immediately after the handler succeeds, before output-format parsing, whether or not any judge is sampled. The value must be acyclic JSON of at most 64 KiB encoded; it is returned unchanged on `ProviderResponse.judge_context` and is never truncated or transformed. An invalid value produces a `context` diagnostic and skips every judge.
390
+
6. If `judgeConfiguration` is present:
391
+
-**Default (`skip_judges=False`):** runs sampled judges sequentially, in configured order, only the first occurrence of each key. Results are returned in `ProviderResponse.judge_results`. Each judge is isolated: a failure adds one `JudgeDiagnostic` and never erases the primary result or another judge's result. Each judge is bounded by `judge_timeout_ms`, and its reasoning is capped at 4 KiB. The resolved context reaches a judge only through its `message_history` variable, between the lines `UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN` and `UNTRUSTED_ACTUATOR_EVIDENCE_END`: never the primary model, never track data, never a span.
392
+
-**`skip_judges=True`:** builds serialisable `JudgeTask` objects for each judge (no AI calls). Returns them in `ProviderResponse.judge_tasks`, with the resolved context on each task and any build-step diagnostics on `ProviderResponse.judge_diagnostics`. Pass each task to a background thread running `run_judge(task, handlers)`.
393
+
7. Returns a `ProviderResponse` (always includes `response`, `usage`, and `track_data`).
Copy file name to clipboardExpand all lines: packages/client/agents.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -123,7 +123,7 @@ Handlers may return any of these — the client normalizes them before emitting
123
123
- On success: emits `$ld:ai:generation:success` + token tracks
124
124
- On error: emits `$ld:ai:generation:error` then re-raises
125
125
3. If `judge_configuration.judges` is present, runs each judge handler (sampled by `sampling_rate`) against the primary response, tracks `evaluation_metric_key`, and emits a `gen_ai.evaluation.result` span event on the judge's `invoke_agent` span (`gen_ai.evaluation.name` / `.score.value` / `.explanation`).
126
-
4. Returns `ProviderResponse`: `{ response: str, usage: UsageDict, track_data: TrackData, judge_results?: dict[str, JudgeResult], judge_tasks?: list[JudgeTask] }`. `judge_results` is populated when `skip_judges=False` (default) and judges ran; `judge_tasks` is populated when `skip_judges=True`.
126
+
4. Returns `ProviderResponse`: `{ response: str, usage: UsageDict, track_data: TrackData, judge_context?: JsonValue, judge_diagnostics?: list[JudgeDiagnostic], judge_results?: dict[str, JudgeResult], judge_tasks?: list[JudgeTask] }`. `judge_context` is the caller callback's value, resolved once after the primary handler and injected only into each judge's `message_history`; `judge_diagnostics` says why a judge was skipped or failed. `judge_results` is populated when `skip_judges=False` (default) and judges ran; `judge_tasks` is populated when `skip_judges=True`.
0 commit comments