Skip to content

Commit 97342d6

Browse files
committed
feat(client): grounded judge context and per-judge diagnostics
Callers can now hand judges evidence about what actually happened during a request. config() takes a lazy judge_context callback because the value does not exist yet when config() is called: the caller's tools fill it while the primary handler runs. The SDK resolves it exactly once, right after the primary handler succeeds and before output-format parsing, so every request with a callback freezes the same snapshot whether or not a judge is sampled, and whether or not skip_judges is set, on invoke() and stream() alike. A resolved context must be acyclic JSON of at most 64 KiB encoded. It is returned unchanged on ProviderResponse.judge_context and reaches a judge only through that judge's message_history variable, between the UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN and UNTRUSTED_ACTUATOR_EVIDENCE_END lines. It never reaches the primary model, the track data, or a span. The block is one more part of judge_scoring.build_message_history, after the answer and before the formatting instructions, so the inline path and run_judge on a JudgeTask still show a judge the same conversation, trajectory included. With no callback the judge prompt is byte-identical to before. Judges are now isolated from each other and from the primary result. Config lookup, provider call, parse and tracking each sit behind their own boundary, bounded by judge_timeout_ms, and a failure produces one JudgeDiagnostic instead of discarding work that already succeeded. A judge no registered handler can serve, and a malformed judgeConfiguration block or entry, are reported the same way rather than raised or dropped. A judge that beats the clock and then fails to track keeps its result. A judge that misses the clock has its late completion consumed silently, so it can neither mutate results nor emit the score metric. Diagnostics carry codes only, never exception text, and the same codes are shared with the TypeScript SDK. graph().invoke() and the final graph().stream() done event forward the graph-level judge's diagnostics. Graph nodes do not receive a caller judge context in v1. run_judges and build_judge_tasks now return result objects carrying both the results and the diagnostics, which is a breaking change for direct callers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> BREAKING CHANGE: run_judges and build_judge_tasks now return a result object (RunJudgesResult with judge_results/judge_diagnostics, and BuildJudgeTasksResult with judge_tasks/judge_diagnostics/judge_context) instead of the bare value. The per-entry shape inside judge_results is unchanged. Both are exported from the package.
1 parent 627f1fa commit 97342d6

14 files changed

Lines changed: 2031 additions & 255 deletions

File tree

‎AGENTS.md‎

Lines changed: 30 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -210,7 +210,28 @@ The value returned to callers of `config().invoke()`. A `dataclass` with the fol
210210
| `usage` | `UsageDict` | Normalized token counts (`input`, `output`, `total`). |
211211
| `track_data` | `TrackData` | Tracking payload from this invocation (run ID, config key, etc.). Carried inside each `JudgeTask` so background judge results are attributed to the originating request. |
212212
| `judge_results` | `dict[str, JudgeResult]?` | Results from inline judge evaluations. Present when `skip_judges=False` (default) and judges ran. |
213-
| `judge_tasks` | `list[JudgeTask]?` | Pre-packaged judge tasks. Present (as a list) when `skip_judges=True`. Each task is a serialisable dataclass ready to pass to a background thread running `run_judge(task, handlers)`. `None` when `skip_judges=False`. |
213+
| `judge_tasks` | `list[JudgeTask]?` | Pre-packaged judge tasks. Present (as a list) when `skip_judges=True`. Each task is a serialisable dataclass ready to pass to a background thread running `run_judge(task, handlers)`, and carries the resolved `judge_context` so the worker injects the identical evidence block. `None` when `skip_judges=False`. |
214+
| `judge_context` | `JsonValue?` | The value the `judge_context` callback returned, unchanged. `None` when no callback was given, it returned `None` (no evidence), or it failed validation. |
215+
| `judge_diagnostics` | `list[JudgeDiagnostic]?` | Why judges were skipped or failed. `None` when nothing went wrong (never `[]`). |
216+
217+
#### `JsonValue`
218+
219+
`None | bool | int | float | str | list[JsonValue] | dict[str, JsonValue]` — a value that survives a `json.dumps` / `json.loads` round trip unchanged.
220+
221+
#### `JudgeDiagnostic`
222+
223+
One reason a judge produced no result, or a partial one. The strings are shared with the TypeScript SDK on the wire; never rename them. A diagnostic carries no raw exception text.
224+
225+
| Field | Type | Description |
226+
|---|---|---|
227+
| `status` | `"skipped" \| "failed"` | Skipped means the judge never ran. |
228+
| `stage` | `"context" \| "config" \| "provider" \| "parse" \| "track" \| "timeout"` | Where it went wrong. |
229+
| `code` | `"context_callback_failed" \| "context_invalid_json" \| "context_too_large" \| "judge_duplicate_key" \| "judge_config_failed" \| "judge_provider_failed" \| "judge_response_invalid" \| "judge_tracking_failed" \| "judge_timed_out"` | The specific reason. |
230+
| `judge_key` | `str?` | The judge this concerns. Absent for context-stage diagnostics and for a malformed `judgeConfiguration` entry with no string key. |
231+
232+
#### Stream `done` event
233+
234+
The final event of `config().stream()`: `{"type": "done", "response": str, "usage": UsageDict, "judge_context": JsonValue \| None, "judge_results": dict[str, JudgeResult] \| None, "judge_diagnostics": list[JudgeDiagnostic] \| None}`. Exactly one is yielded, after every `chunk`. Empty results and diagnostics are `None`.
214235

215236
#### `ProviderGraphResponse`
216237

@@ -221,6 +242,7 @@ The value returned by `graph().invoke()`. A dataclass with attribute access.
221242
| `response` | `str` | The final text output (from the last node executed). |
222243
| `usage` | `UsageDict` | Aggregate token counts across all nodes. |
223244
| `judge_results` | `dict[str, JudgeResult]?` | Results from a graph-level judge, if configured. |
245+
| `judge_diagnostics` | `list[JudgeDiagnostic]?` | Diagnostics from the graph-level judge. Graph nodes do not receive a caller judge context in v1. |
224246

225247
#### `ConfigArgs`
226248

@@ -232,6 +254,8 @@ Arguments accepted by `config()`.
232254
| `handler` | `ProviderHandler \| list[ProviderHandler]`? | One handler or an ordered list of handlers. Routing selects the match by provider + mode. |
233255
| `tool_handlers` | `dict[str, Callable \| NativeTool]?` | Map of tool name → implementation function (or `NativeTool` sentinel). |
234256
| `registry` | `Registry?` | Registry to source handlers and tools from. Local `handler`/`tool_handlers` take precedence. |
257+
| `judge_context` | `Callable[[], JsonValue \| Awaitable[JsonValue]]?` | Lazily resolves JSON-safe evidence for the judges. Lazy on purpose: the value does not exist when `config()` is called, and the caller's tools fill it while the primary handler runs. |
258+
| `judge_timeout_ms` | `int` | Per-judge budget covering config lookup, provider call and parse. Default: `30_000`. |
235259
| `skip_judges` | `bool`? | When `True`, `invoke()` does not run judges inline. Instead it returns `judge_tasks: list[JudgeTask]` — pre-packaged tasks ready for background thread execution via `run_judge(task, handlers)`. Default: `False`. |
236260

237261
#### `TrackData`
@@ -362,10 +386,11 @@ Returns a `ConfigInstance` with:
362386
2. Selects the handler by matching on `[config.provider.name, normalized mode]`. Selection priority: (a) exact provider match, (b) wildcard `['*', mode]` fallback for multi-provider adapters (e.g. LangChain). Raises if no matching handler is found.
363387
3. Invokes the selected handler with the config, user input, tool handlers, variables, and history. The `context` passed to `.invoke()` is automatically merged into `variables` under the key `ldContext`, so templates can reference `{{ldContext.key}}`, `{{ldContext.email}}`, etc. If `history` is provided, it is passed to the handler as the 5th positional argument — messages-mode handlers splice it into the messages array; agent-mode handlers append it to the system prompt.
364388
4. Emits LaunchDarkly telemetry events: duration (`$ld:ai:duration:total`), outcome (`$ld:ai:generation:success` / `$ld:ai:generation:error`), and token counts (`$ld:ai:tokens:*`).
365-
5. If `judgeConfiguration` is present:
366-
- **Default (`skip_judges=False`):** runs each configured judge inline at its `samplingRate`. Results are returned in `ProviderResponse.judge_results`.
367-
- **`skip_judges=True`:** builds serialisable `JudgeTask` objects for each judge (no AI calls). Returns them in `ProviderResponse.judge_tasks`. Pass each task to a background thread running `run_judge(task, handlers)`.
368-
6. Returns a `ProviderResponse` (always includes `response`, `usage`, and `track_data`).
389+
5. If a `judge_context` callback was given, resolves it exactly once — immediately after the handler succeeds, before output-format parsing, whether or not any judge is sampled. The value must be acyclic JSON of at most 64 KiB encoded; it is returned unchanged on `ProviderResponse.judge_context` and is never truncated or transformed. An invalid value produces a `context` diagnostic and skips every judge.
390+
6. If `judgeConfiguration` is present:
391+
- **Default (`skip_judges=False`):** runs sampled judges sequentially, in configured order, only the first occurrence of each key. Results are returned in `ProviderResponse.judge_results`. Each judge is isolated: a failure adds one `JudgeDiagnostic` and never erases the primary result or another judge's result. Each judge is bounded by `judge_timeout_ms`, and its reasoning is capped at 4 KiB. The resolved context reaches a judge only through its `message_history` variable, between the lines `UNTRUSTED_ACTUATOR_EVIDENCE_BEGIN` and `UNTRUSTED_ACTUATOR_EVIDENCE_END`: never the primary model, never track data, never a span.
392+
- **`skip_judges=True`:** builds serialisable `JudgeTask` objects for each judge (no AI calls). Returns them in `ProviderResponse.judge_tasks`, with the resolved context on each task and any build-step diagnostics on `ProviderResponse.judge_diagnostics`. Pass each task to a background thread running `run_judge(task, handlers)`.
393+
7. Returns a `ProviderResponse` (always includes `response`, `usage`, and `track_data`).
369394

370395
### `graph(key, **options)`
371396

‎TELEMETRY-CONTRACT.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -238,6 +238,10 @@ attributes require `captureContent` / `capture_content`, a handler-factory optio
238238
not receive. The reasoning is still returned to the caller in `judgeResults` / `judge_results`;
239239
only the telemetry copy is withheld. Exporting it needs its own opt-in.
240240

241+
The caller-supplied judge context (`judge_context` / `judgeContext`) is never recorded in
242+
telemetry either: it reaches the judge only through its prompt, and never becomes a span
243+
attribute, a span event, or track data.
244+
241245
---
242246

243247
## 5. Finish reasons

‎packages/client/agents.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -123,7 +123,7 @@ Handlers may return any of these — the client normalizes them before emitting
123123
- On success: emits `$ld:ai:generation:success` + token tracks
124124
- On error: emits `$ld:ai:generation:error` then re-raises
125125
3. If `judge_configuration.judges` is present, runs each judge handler (sampled by `sampling_rate`) against the primary response, tracks `evaluation_metric_key`, and emits a `gen_ai.evaluation.result` span event on the judge's `invoke_agent` span (`gen_ai.evaluation.name` / `.score.value` / `.explanation`).
126-
4. Returns `ProviderResponse`: `{ response: str, usage: UsageDict, track_data: TrackData, judge_results?: dict[str, JudgeResult], judge_tasks?: list[JudgeTask] }`. `judge_results` is populated when `skip_judges=False` (default) and judges ran; `judge_tasks` is populated when `skip_judges=True`.
126+
4. Returns `ProviderResponse`: `{ response: str, usage: UsageDict, track_data: TrackData, judge_context?: JsonValue, judge_diagnostics?: list[JudgeDiagnostic], judge_results?: dict[str, JudgeResult], judge_tasks?: list[JudgeTask] }`. `judge_context` is the caller callback's value, resolved once after the primary handler and injected only into each judge's `message_history`; `judge_diagnostics` says why a judge was skipped or failed. `judge_results` is populated when `skip_judges=False` (default) and judges ran; `judge_tasks` is populated when `skip_judges=True`.
127127

128128
---
129129

‎packages/client/src/launchdarkly_ai_server/__init__.py‎

Lines changed: 17 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -47,7 +47,15 @@
4747
image_block_to_url,
4848
is_content_blocks,
4949
)
50-
from .judges import build_judge_tasks, run_judge, run_judges
50+
from .judges import (
51+
BuildJudgeTasksResult,
52+
JudgeContextResolution,
53+
RunJudgesResult,
54+
build_judge_tasks,
55+
resolve_judge_context,
56+
run_judge,
57+
run_judges,
58+
)
5159
from .lifecycle import (
5260
extract_variation,
5361
get_client,
@@ -80,6 +88,8 @@
8088
HandlerStreamEvent,
8189
InitClientOptions,
8290
InputTokenDetails,
91+
JsonValue,
92+
JudgeDiagnostic,
8393
JudgeResult,
8494
JudgeRunResult,
8595
JudgeTask,
@@ -143,6 +153,10 @@
143153
"HandlerResult",
144154
"HandlerStreamEvent",
145155
"InitClientOptions",
156+
"BuildJudgeTasksResult",
157+
"JsonValue",
158+
"JudgeContextResolution",
159+
"JudgeDiagnostic",
146160
"JudgeResult",
147161
"JudgeRunResult",
148162
"JudgeTask",
@@ -242,7 +256,9 @@
242256
# judges
243257
"build_judge_tasks",
244258
"run_judge",
259+
"resolve_judge_context",
245260
"run_judges",
261+
"RunJudgesResult",
246262
# client
247263
"config",
248264
"ConfigInstance",

‎packages/client/src/launchdarkly_ai_server/client.py‎

Lines changed: 79 additions & 21 deletions
Original file line numberDiff line numberDiff line change
@@ -1,16 +1,23 @@
11
from __future__ import annotations
22

33
import json
4-
from collections.abc import AsyncGenerator, Callable
4+
from collections.abc import AsyncGenerator, Awaitable, Callable
55
from typing import Any
66

77
from .conversation import bind_conversation_id
8-
from .judges import build_judge_tasks, run_judges
8+
from .judges import (
9+
DEFAULT_JUDGE_TIMEOUT_MS,
10+
build_judge_tasks,
11+
resolve_judge_context,
12+
run_judges,
13+
)
914
from .lifecycle import extract_variation
1015
from .registry import resolve_handlers, resolve_tools
1116
from .tracking import execute_and_stream, execute_and_track
1217
from .types import (
1318
AiConfigRep,
19+
JsonValue,
20+
JudgeDiagnostic,
1421
LDContext,
1522
NativeTool,
1623
ProviderHandler,
@@ -52,12 +59,16 @@ def __init__(
5259
tool_handlers: dict[str, Callable[..., Any] | NativeTool] | None,
5360
registry: Any, # Registry | None
5461
skip_judges: bool = False,
62+
judge_context: Callable[[], JsonValue | Awaitable[JsonValue]] | None = None,
63+
judge_timeout_ms: int = DEFAULT_JUDGE_TIMEOUT_MS,
5564
) -> None:
5665
self._key = key
5766
self._handler = handler
5867
self._tool_handlers = tool_handlers
5968
self._registry = registry
6069
self._skip_judges = skip_judges
70+
self._judge_context = judge_context
71+
self._judge_timeout_ms = judge_timeout_ms
6172

6273
def _normalize_handlers(self) -> list[ProviderHandler] | None:
6374
if self._handler is None:
@@ -102,6 +113,10 @@ async def invoke(
102113
usage: dict[str, int] = result["usage"]
103114
track_data = result["track_data"]
104115

116+
# Freeze the caller's context the moment the primary handler succeeds, before output
117+
# parsing. Sampling controls judge execution, never this boundary.
118+
context_resolution = await resolve_judge_context(self._judge_context)
119+
105120
parsed_response = _resolve_output_format_response(
106121
raw_response,
107122
config.get("outputFormat") if isinstance(config, dict) else None,
@@ -116,7 +131,7 @@ async def invoke(
116131
usage_obj = to_usage_dict(usage)
117132

118133
if self._skip_judges:
119-
judge_tasks = await build_judge_tasks(
134+
build_result = await build_judge_tasks(
120135
config=config,
121136
user_context=context,
122137
handler=handler,
@@ -125,29 +140,48 @@ async def invoke(
125140
base_track_data=track_data,
126141
user_input=user_input,
127142
trajectory=result.get("trajectory", ""),
143+
context_resolution=context_resolution,
128144
)
129145
return ProviderResponse(
130146
response=parsed_response,
131147
usage=usage_obj,
132-
judge_tasks=judge_tasks,
148+
judge_context=build_result.judge_context,
149+
judge_diagnostics=build_result.judge_diagnostics or None,
150+
judge_tasks=build_result.judge_tasks,
133151
track_data=track_data,
134152
)
135153

136-
judge_results = await run_judges(
137-
config=config,
138-
user_context=context,
139-
handler=handler,
140-
handlers=resolved_handler_list,
141-
user_input=user_input,
142-
trajectory=result.get("trajectory", ""),
143-
llm_response=llm_str,
144-
base_track_data=track_data,
145-
tool_handlers=resolved_tools,
154+
diagnostics: list[JudgeDiagnostic] = (
155+
[context_resolution.diagnostic]
156+
if context_resolution.diagnostic is not None
157+
else []
146158
)
159+
if not context_resolution.failed:
160+
judge_run = await run_judges(
161+
config=config,
162+
user_context=context,
163+
handler=handler,
164+
handlers=resolved_handler_list,
165+
user_input=user_input,
166+
trajectory=result.get("trajectory", ""),
167+
llm_response=llm_str,
168+
base_track_data=track_data,
169+
tool_handlers=resolved_tools,
170+
judge_context=context_resolution.judge_context,
171+
judge_context_json=context_resolution.serialized,
172+
judge_timeout_ms=self._judge_timeout_ms,
173+
)
174+
diagnostics.extend(judge_run.judge_diagnostics)
175+
judge_results = judge_run.judge_results
176+
else:
177+
judge_results = {}
178+
147179
return ProviderResponse(
148180
response=parsed_response,
149181
usage=usage_obj,
150-
judge_results=judge_results if judge_results else None,
182+
judge_context=context_resolution.judge_context,
183+
judge_diagnostics=diagnostics or None,
184+
judge_results=judge_results or None,
151185
track_data=track_data,
152186
)
153187

@@ -209,10 +243,17 @@ async def _stream_events(
209243

210244
if done_event:
211245
track_data = done_event.get("track_data", {})
212-
judge_results = (
213-
{}
214-
if self._skip_judges
215-
else await run_judges(
246+
judge_results: dict[str, Any] = {}
247+
diagnostics: list[JudgeDiagnostic] = []
248+
249+
# Same freeze boundary as invoke(): resolved once the primary finishes,
250+
# whether or not judges run.
251+
context_resolution = await resolve_judge_context(self._judge_context)
252+
judge_context = context_resolution.judge_context
253+
if context_resolution.diagnostic is not None:
254+
diagnostics.append(context_resolution.diagnostic)
255+
if not self._skip_judges and not context_resolution.failed:
256+
judge_run = await run_judges(
216257
config=config,
217258
user_context=context,
218259
handler=handler,
@@ -222,13 +263,20 @@ async def _stream_events(
222263
llm_response=done_event.get("response", ""),
223264
base_track_data=track_data,
224265
tool_handlers=resolved_tools,
266+
judge_context=context_resolution.judge_context,
267+
judge_context_json=context_resolution.serialized,
268+
judge_timeout_ms=self._judge_timeout_ms,
225269
)
226-
)
270+
judge_results = judge_run.judge_results
271+
diagnostics.extend(judge_run.judge_diagnostics)
272+
227273
yield {
228274
"type": "done",
229275
"response": done_event.get("response", ""),
230276
"usage": done_event.get("usage"),
231-
"judge_results": judge_results if judge_results else None,
277+
"judge_context": judge_context,
278+
"judge_results": judge_results or None,
279+
"judge_diagnostics": diagnostics or None,
232280
}
233281

234282

@@ -239,6 +287,8 @@ def config(
239287
tool_handlers: dict[str, Callable[..., Any] | NativeTool] | None = None,
240288
registry: Any = None,
241289
skip_judges: bool = False,
290+
judge_context: Callable[[], JsonValue | Awaitable[JsonValue]] | None = None,
291+
judge_timeout_ms: int = DEFAULT_JUDGE_TIMEOUT_MS,
242292
) -> ConfigInstance:
243293
"""
244294
Creates a ``ConfigInstance`` bound to *key*. Accepts a single handler or a
@@ -251,11 +301,19 @@ def config(
251301
``.invoke()`` / ``.stream()``. When set, ``invoke()`` returns
252302
``judge_tasks: list[JudgeTask]`` — pre-packaged tasks ready for a background
253303
thread calling ``run_judge(task, handlers)``.
304+
305+
``judge_context`` is a callback returning JSON-safe evidence for the judges. It is lazy on
306+
purpose: the value does not exist yet when ``config()`` is called, and the caller's tools
307+
fill it while the primary handler runs. It resolves exactly once per request, right after
308+
the primary handler succeeds, and reaches the judges only through their ``message_history``
309+
variable. ``judge_timeout_ms`` bounds one judge's config lookup, provider call and parse.
254310
"""
255311
return ConfigInstance(
256312
key=key,
257313
handler=handler,
258314
tool_handlers=tool_handlers,
259315
registry=registry,
260316
skip_judges=skip_judges,
317+
judge_context=judge_context,
318+
judge_timeout_ms=judge_timeout_ms,
261319
)

0 commit comments

Comments
 (0)