From 657b6d16e98aaf0420dc09881d5a91b236c27c77 Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 18:23:11 +0900 Subject: [PATCH 01/13] docs: simulation tests section for agent tests Co-Authored-By: Claude Fable 5.1 --- agents/test/agent-tests.mdx | 88 ++++++++++++++++++++++++++++++++----- 1 file changed, 77 insertions(+), 11 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index b58c097..4493229 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -15,8 +15,8 @@ A **Single Turn** test hands your agent a scripted conversation and asks it to p 3. An LLM judge scores the reply against your **Expectation**, optionally calibrated by success and failure examples, and returns a verdict with its reasoning. - Tests run as text (no audio is synthesized), but the agent uses its full - draft configuration: the [knowledge base](/agents/build/knowledge-base) is + Tests run as text (no audio is synthesized), but the agent uses its full draft + configuration: the [knowledge base](/agents/build/knowledge-base) is consulted, and the agent can invoke its attached tools while generating the reply. [Webhook tools](/agents/build/webhook-tools) send real HTTP requests during a test, so point them at a staging endpoint. For end-to-end @@ -48,9 +48,9 @@ Tests are workspace-level resources, managed under **Library → Tests** in the - The type picker also offers **Tool** tests, available today, which check - whether the agent called a specific tool. **Multi Turn** tests are coming - soon. + The type picker also offers **Tool** tests, which check whether the agent + called a specific tool, and **Multi Turn** tests, which run a whole simulated + conversation. See [Simulation tests](#simulation-tests) below. ## Attach tests to agents @@ -79,6 +79,67 @@ If the run can't complete, the verdict shows **Error** with the error message in On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete), and the page header summarizes the latest batch, for example `4 passed, 1 failed on last run`. Use the row menu to re-run a single test, edit it, or remove it from the agent. +## Simulation tests + +A **Multi Turn** test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions. Use it for flows that only show up over several exchanges: collecting details step by step, handling objections, or deciding when to hand off. + +### Write the scenario + +The scenario is the simulated user's brief. Click **Insert template** to start from the four sections the simulator expects: + +| Section | What to write | +| ----------- | -------------------------------------------------------------------------------------------- | +| **PERSONA** | Who the user is and how they talk, for example an impatient customer on a lunch break. | +| **GOAL** | What they want out of the conversation. | +| **FACTS** | Details the agent may ask for, such as an order number or a date. Reveal one fact at a time. | +| **ENDING** | When the user hangs up: satisfied, out of patience, or after a fixed number of tries. | + +Two rules make scenarios reliable. Reveal one fact at a time, so the agent has to ask for what it needs instead of receiving everything in the first message. And give an explicit ending, otherwise the conversation runs until the turn limit and the judge has to guess whether the user was done. + +A **Conversation** script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message. + +### Success conditions + +Add one to ten **Success conditions**, each a plain sentence describing something that must happen in the conversation, for example `The agent confirms the refund amount before closing`. Give each a short name or leave it blank and the first words of the description become the name. + +The judge reads the whole transcript and marks every condition **Success**, **Failure**, or **Unknown**. A run passes only when every condition succeeds. An **Unknown** verdict means the transcript did not contain enough evidence either way, and the run is flagged **Needs review** so you read the transcript before trusting the result. Success and failure examples, when you add them, calibrate the judge across all conditions. + +### Tool mocks + +Simulated users usually run many times, so by default tool calls never reach your real endpoints: + +| Strategy | Behaviour | +| ----------------- | ----------------------------------------------------------------------------------------------------------------------- | +| **Mock all** | Every tool answers from a mock. Tools without their own entry use their default mock response. This is the default. | +| **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | +| **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | + +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). + +### Assertions + +Under **Advanced**, add deterministic checks that run alongside the judge: + +- **Required tool calls**: the agent must call the tool, optionally with parameters that match. +- **Forbidden tools**: the agent must not call the tool. +- **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. +- **Final node**: the workflow node the conversation must end on. + +A failed assertion fails the run regardless of the judge's verdict. Each assertion shows **Passed** or **Failed** with a short detail in the result. + +### Repeat count and pass rate + +Set **Repeat** to run the same scenario up to ten times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The row shows how many repeats passed, for example `2 / 3 passed`, and the page header shows the batch **pass rate**: green at 100%, amber from 80%, red below. + +### Read the result + +A simulation runs in the background and can take a few minutes. The sheet shows **Queued**, then **Running**, then the result: + +- The **Transcript** with every user and agent turn, and each tool call marked **Mock** when it was answered from a mock. +- Every **Success condition** with its verdict and the judge's rationale, plus the judge's summary. +- Every **Assertion** with its outcome. +- Who ended the conversation, how many turns were used, and token usage for the agent, the simulated user, and the judge. + ## Tests run against the draft Tests always exercise the agent's latest **draft** configuration, including unpublished changes to the system prompt. That makes the loop fast: @@ -98,12 +159,17 @@ Tests always exercise the agent's latest **draft** configuration, including unpu ## Limits -| Field | Limit | -| ------------------------- | ------------------------------------------- | -| Test name | 200 characters | -| Conversation message | 2,000 characters each, at least one message | -| Expectation | 400 characters | -| Success / failure example | 400 characters each | +| Field | Limit | +| ------------------------------- | --------------------------------------------------------------------- | +| Test name | 200 characters | +| Conversation message | 2,000 characters each, at least one message (optional for Multi Turn) | +| Expectation | 400 characters | +| Success / failure example | 400 characters each | +| Scenario (Multi Turn) | 6,000 characters | +| Success conditions (Multi Turn) | 1 to 10, description 500 characters each | +| Max turns (Multi Turn) | 30 | +| Repeat count (Multi Turn) | 10 | +| Dynamic variables | 50 per test | ## Going further From 05816f4db8f0cbf2bdd397cce5231ffaad7c5487 Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 19:04:44 +0900 Subject: [PATCH 02/13] docs: introduce the three agent test types up front instead of in a note Co-Authored-By: Claude Fable 5.1 --- agents/test/agent-tests.mdx | 120 ++++++++++++++++++------------------ 1 file changed, 59 insertions(+), 61 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 4493229..cbd6f98 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -1,26 +1,27 @@ --- title: "Agent Tests" -description: "Script conversations and let an LLM judge score your agent's replies to catch regressions before you publish" +description: "Script a conversation or simulate a user, then let an LLM judge score your agent to catch regressions before you publish" icon: "vial" --- -Write a conversation once, run it against any agent, and get a **Pass** or **Fail** verdict with the judge's reasoning. Tests live in a shared workspace library, run against your agent's current draft, and never touch what's published, so you can iterate on a prompt and re-run in seconds. +Write a test once, run it against any agent, and get a **Pass** or **Fail** verdict with the judge's reasoning. Tests live in a shared workspace library, run against your agent's current draft, and never touch what's published, so you can iterate on a prompt and re-run in seconds. -## How a test works +## Test types -A **Single Turn** test hands your agent a scripted conversation and asks it to produce the next reply: - -1. You script a conversation history of agent and user messages. -2. The agent generates the next reply using its current **draft** configuration and the same language model that answers in live conversations. -3. An LLM judge scores the reply against your **Expectation**, optionally calibrated by success and failure examples, and returns a verdict with its reasoning. +| Type | What it checks | Use it for | +| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | +| **Single Turn** | The next reply to a scripted conversation, scored by an LLM judge against your expectation. | Wording, tone, and one specific answer. | +| **Tool** | Whether the next reply calls a given tool (or no tool at all), optionally with matching parameters. | Routing decisions and argument extraction. | +| **Multi Turn** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. | Tests run as text (no audio is synthesized), but the agent uses its full draft configuration: the [knowledge base](/agents/build/knowledge-base) is - consulted, and the agent can invoke its attached tools while generating the - reply. [Webhook tools](/agents/build/webhook-tools) send real HTTP requests - during a test, so point them at a staging endpoint. For end-to-end - verification with voice, use [preview calls](/agents/test/preview-calls). + consulted, and the agent can invoke its attached tools. In Single Turn and + Tool tests, [webhook tools](/agents/build/webhook-tools) send real HTTP + requests, so point them at a staging endpoint. Multi Turn tests mock tools by + default. For end-to-end verification with voice, use [preview + calls](/agents/test/preview-calls). ## Create a test @@ -29,59 +30,38 @@ Tests are workspace-level resources, managed under **Library → Tests** in the - Open **Library → Tests** and click **New test**. Give it a name and keep the - type as **Single Turn**. - - - Under **Conversation**, click **Add message** to build the history the agent - sees. Each message is either an **Agent** or **User** turn. When the test - runs, the agent generates the reply that comes next. + Open **Library → Tests** and click **New test**. Give it a name and pick a + **Type**. - - Under **Judging**, write the **Expectation**: what a correct reply must do. - Optionally click **Add example** to provide success and failure examples; - they calibrate the judge but aren't required. + + Fill in the fields for that type. They are described in the sections below: + [Single Turn](#single-turn-tests), [Tool](#tool-tests), and [Multi + Turn](#multi-turn-tests). Click **Save**. The test is now in your library, ready to attach to agents. - - The type picker also offers **Tool** tests, which check whether the agent - called a specific tool, and **Multi Turn** tests, which run a whole simulated - conversation. See [Simulation tests](#simulation-tests) below. - +## Single Turn tests -## Attach tests to agents +A Single Turn test hands your agent a scripted conversation and asks it to produce the next reply: -A test only runs against agents it's attached to. Attach from either side: - -- **From the library**: open the test's **Access** tab and toggle it on for each agent. -- **From the Builder**: on the agent's **Tests** page, click **Add tests** and pick from the library. - -One test can be attached to many agents, and each agent keeps its own last result. **Remove from agent** detaches the test from that agent only; **Delete** in the library removes the test from all agents. - -## Run a single test - -Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. The verdict card shows: +1. You script a conversation history of agent and user messages. +2. The agent generates the next reply using its current **draft** configuration and the same language model that answers in live conversations. +3. An LLM judge scores the reply against your **Expectation**, optionally calibrated by success and failure examples, and returns a verdict with its reasoning. -| Result | What you see | -| --------------- | -------------------------------------------------- | -| Verdict | **Pass** (green) or **Fail** (red) | -| Latency | Time to generate and judge the reply, e.g. `1.8 s` | -| Agent reply | The full reply the agent generated | -| Judge reasoning | Why the judge passed or failed the reply | +Under **Conversation**, click **Add message** to build the history the agent sees. Each message is either an **Agent** or **User** turn, and at least one must be a user turn. Under **Judging**, write the **Expectation**: what a correct reply must do. Optionally click **Add example** to provide success and failure examples. They calibrate the judge but aren't required. -If the run can't complete, the verdict shows **Error** with the error message in place of the reply and reasoning. +## Tool tests -## Run every test for an agent +A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it. -On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete), and the page header summarizes the latest batch, for example `4 passed, 1 failed on last run`. Use the row menu to re-run a single test, edit it, or remove it from the agent. +Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called. -## Simulation tests +## Multi Turn tests -A **Multi Turn** test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions. Use it for flows that only show up over several exchanges: collecting details step by step, handling objections, or deciding when to hand off. +A Multi Turn test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions. ### Write the scenario @@ -96,17 +76,17 @@ The scenario is the simulated user's brief. Click **Insert template** to start f Two rules make scenarios reliable. Reveal one fact at a time, so the agent has to ask for what it needs instead of receiving everything in the first message. And give an explicit ending, otherwise the conversation runs until the turn limit and the judge has to guess whether the user was done. -A **Conversation** script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message. +**Max turns** caps how many user turns the simulation runs. A **Conversation** script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message. ### Success conditions -Add one to ten **Success conditions**, each a plain sentence describing something that must happen in the conversation, for example `The agent confirms the refund amount before closing`. Give each a short name or leave it blank and the first words of the description become the name. +Add one to ten **Success conditions**, each a plain sentence describing something that must happen in the conversation, for example `The agent confirms the refund amount before closing`. Give each a short name, or leave it blank and the first words of the description become the name. The judge reads the whole transcript and marks every condition **Success**, **Failure**, or **Unknown**. A run passes only when every condition succeeds. An **Unknown** verdict means the transcript did not contain enough evidence either way, and the run is flagged **Needs review** so you read the transcript before trusting the result. Success and failure examples, when you add them, calibrate the judge across all conditions. ### Tool mocks -Simulated users usually run many times, so by default tool calls never reach your real endpoints: +Simulated conversations usually run many times, so by default tool calls never reach your real endpoints: | Strategy | Behaviour | | ----------------- | ----------------------------------------------------------------------------------------------------------------------- | @@ -127,18 +107,36 @@ Under **Advanced**, add deterministic checks that run alongside the judge: A failed assertion fails the run regardless of the judge's verdict. Each assertion shows **Passed** or **Failed** with a short detail in the result. -### Repeat count and pass rate +### Repeat count -Set **Repeat** to run the same scenario up to ten times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The row shows how many repeats passed, for example `2 / 3 passed`, and the page header shows the batch **pass rate**: green at 100%, amber from 80%, red below. +Set **Repeat** to run the same scenario up to ten times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`. + +## Attach tests to agents -### Read the result +A test only runs against agents it's attached to. Attach from either side: + +- **From the library**: open the test's **Access** tab and toggle it on for each agent. +- **From the Builder**: on the agent's **Tests** page, click **Add tests** and pick from the library. -A simulation runs in the background and can take a few minutes. The sheet shows **Queued**, then **Running**, then the result: +One test can be attached to many agents, and each agent keeps its own last result. **Remove from agent** detaches the test from that agent only. **Delete** in the library removes the test from all agents. + +## Run a single test + +Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Multi Turn test runs in the background, shows **Queued** and then **Running**, and can take a few minutes. + +The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type: + +| Type | What you see | +| --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | +| **Tool** | Which tool the agent called, if any, against what the test expected. | +| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, token usage, and the full transcript with each tool call marked **Mock** when it answered from a mock. | + +A Multi Turn run with an **Unknown** condition shows **Needs review** instead of a verdict. If a run can't complete, the card shows **Error** with the error message. + +## Run every test for an agent -- The **Transcript** with every user and agent turn, and each tool call marked **Mock** when it was answered from a mock. -- Every **Success condition** with its verdict and the judge's rationale, plus the judge's summary. -- Every **Assertion** with its outcome. -- Who ended the conversation, how many turns were used, and token usage for the agent, the simulated user, and the judge. +On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Multi Turn tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. Use the row menu to re-run a single test, edit it, or remove it from the agent. ## Tests run against the draft From fd615085d04f3d278e7cec6ca1953addeea7146f Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 20:11:36 +0900 Subject: [PATCH 03/13] docs: simulation results show LLM cost per participant Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index cbd6f98..a6d02e2 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -126,11 +126,11 @@ Open the test and switch to its **Run** tab. Pick an agent that has access, then The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type: -| Type | What you see | -| --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | -| **Tool** | Which tool the agent called, if any, against what the test expected. | -| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, token usage, and the full transcript with each tool call marked **Mock** when it answered from a mock. | +| Type | What you see | +| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | +| **Tool** | Which tool the agent called, if any, against what the test expected. | +| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked **Mock** when it answered from a mock. | A Multi Turn run with an **Unknown** condition shows **Needs review** instead of a verdict. If a run can't complete, the card shows **Error** with the error message. From eaf9bf9515feb65029b7fb7e5021e8e2e2bf1ba0 Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 21:02:00 +0900 Subject: [PATCH 04/13] docs: an errored simulation run still shows its transcript Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index a6d02e2..2e1ca29 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -132,7 +132,7 @@ The verdict card shows **Pass** or **Fail** with the time the run took, followed | **Tool** | Which tool the agent called, if any, against what the test expected. | | **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked **Mock** when it answered from a mock. | -A Multi Turn run with an **Unknown** condition shows **Needs review** instead of a verdict. If a run can't complete, the card shows **Error** with the error message. +A Multi Turn run with an **Unknown** condition shows **Needs review** instead of a verdict. If a run can't complete, the card shows **Error** with the error message. When the conversation finished but could not be judged, the transcript is still shown below the error. ## Run every test for an agent From dcc59a4f00b6dd139a85d1c19c378e9c816e76fb Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 22:27:23 +0900 Subject: [PATCH 05/13] docs: simulation tests have no final node assertion Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 1 - 1 file changed, 1 deletion(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 2e1ca29..be2ce77 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -103,7 +103,6 @@ Under **Advanced**, add deterministic checks that run alongside the judge: - **Required tool calls**: the agent must call the tool, optionally with parameters that match. - **Forbidden tools**: the agent must not call the tool. - **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. -- **Final node**: the workflow node the conversation must end on. A failed assertion fails the run regardless of the judge's verdict. Each assertion shows **Passed** or **Failed** with a short detail in the result. From 10b2d5b8c9a41e074eebc67ca578a19adc3d940c Mon Sep 17 00:00:00 2001 From: Him188 Date: Sun, 27 Sep 2026 22:32:02 +0900 Subject: [PATCH 06/13] docs: simulation tests take a 10,000 character scenario, 50 turns and 20 repeats Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index be2ce77..e1da94f 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -108,7 +108,7 @@ A failed assertion fails the run regardless of the judge's verdict. Each asserti ### Repeat count -Set **Repeat** to run the same scenario up to ten times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`. +Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`. ## Attach tests to agents @@ -162,10 +162,10 @@ Tests always exercise the agent's latest **draft** configuration, including unpu | Conversation message | 2,000 characters each, at least one message (optional for Multi Turn) | | Expectation | 400 characters | | Success / failure example | 400 characters each | -| Scenario (Multi Turn) | 6,000 characters | +| Scenario (Multi Turn) | 10,000 characters | | Success conditions (Multi Turn) | 1 to 10, description 500 characters each | -| Max turns (Multi Turn) | 30 | -| Repeat count (Multi Turn) | 10 | +| Max turns (Multi Turn) | 50 | +| Repeat count (Multi Turn) | 20 | | Dynamic variables | 50 per test | ## Going further From 07ec123dfaec794aebe43dbd57ed3a000098a51f Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 03:11:10 +0900 Subject: [PATCH 07/13] docs(agent-tests): channels, call counts, repeats and broken conversations Co-Authored-By: Claude Opus 5.5 --- agents/build/custom-llm.mdx | 2 +- agents/build/dynamic-variables.mdx | 2 +- agents/build/webhook-tools.mdx | 2 +- agents/test/agent-tests.mdx | 26 +++++++++++++++++++------- 4 files changed, 22 insertions(+), 10 deletions(-) diff --git a/agents/build/custom-llm.mdx b/agents/build/custom-llm.mdx index 4cd4253..fa1691f 100644 --- a/agents/build/custom-llm.mdx +++ b/agents/build/custom-llm.mdx @@ -134,6 +134,6 @@ Voice conversations are latency sensitive, so aim for a time-to-first-token unde ## Limitations -- [Agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. +- Single Turn and Tool [agent tests](/agents/test/agent-tests) are not supported. A scripted test run refuses to execute rather than substitute a platform model for yours. Multi Turn tests run on your endpoint. - Configuration is API-only for now. A console UI comes later. - The prompt-level safety guardrails still ride the assembled context, but your model decides whether to honor them. diff --git a/agents/build/dynamic-variables.mdx b/agents/build/dynamic-variables.mdx index 4ed9f4a..a231ea2 100644 --- a/agents/build/dynamic-variables.mdx +++ b/agents/build/dynamic-variables.mdx @@ -103,7 +103,7 @@ The platform fills a handful of `{{system.*}}` placeholders itself, on every ses | Variable | Value | Notes | |---|---|---| -| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and agent test sessions are all `web_voice`. | +| `system.channel` | `phone_inbound`, `phone_outbound`, or `web_voice` | Web SDK, API, preview, and Single Turn and Tool test sessions are `web_voice`. A Multi Turn test uses its [channel](/agents/test/agent-tests#channel). | | `system.timezone` | The session's resolved IANA timezone, for example `Asia/Tokyo` | `UTC` when nothing resolves. Always equals the session's `timezone` field; see [which timezone a session uses](/agents/build/time-timezone#which-timezone-a-session-uses). | | `system.today` | Today's date in that timezone, ISO 8601 `YYYY-MM-DD` | The calendar date at session creation. There is no `system.now`; see [World context](#world-context). | | `system.language` | The session language code, one of the [52 supported languages](/agents/build/voice-language#speaking-language) | diff --git a/agents/build/webhook-tools.mdx b/agents/build/webhook-tools.mdx index a93aa72..b536e8c 100644 --- a/agents/build/webhook-tools.mdx +++ b/agents/build/webhook-tools.mdx @@ -110,7 +110,7 @@ Use `hide` when error responses might leak internal details you don't want spoke ## Mock responses -A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios, but they don't intercept anything yet: preview and live calls always hit the real endpoint. +A tool can store mock responses: canned payloads, each with a `name`, a `status_code` (100–599), a `content_type`, and a `body`. Mocks are saved with the tool's configuration for test scenarios. [Multi Turn agent tests](/agents/test/agent-tests#tool-mocks) that mock all tools answer the tool with its first mock response. Live calls always hit the real endpoint. ## Test your tool diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index e1da94f..b89bffe 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -30,7 +30,7 @@ Tests are workspace-level resources, managed under **Library → Tests** in the - Open **Library → Tests** and click **New test**. Give it a name and pick a + Open **Library → Tests** and click **Add test**. Give it a name and pick a **Type**. @@ -39,7 +39,8 @@ Tests are workspace-level resources, managed under **Library → Tests** in the Turn](#multi-turn-tests). - Click **Save**. The test is now in your library, ready to attach to agents. + Click **Create Test**. The test is now in your library, ready to attach to + agents. @@ -78,6 +79,10 @@ Two rules make scenarios reliable. Reveal one fact at a time, so the agent has t **Max turns** caps how many user turns the simulation runs. A **Conversation** script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message. +### Channel + +**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. A phone agent, for example, keeps replies short and reads numbers back digit by digit. **Auto**, the default, uses Phone inbound when the agent answers a phone number and Web voice otherwise. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. + ### Success conditions Add one to ten **Success conditions**, each a plain sentence describing something that must happen in the conversation, for example `The agent confirms the refund amount before closing`. Give each a short name, or leave it blank and the first words of the description become the name. @@ -90,17 +95,21 @@ Simulated conversations usually run many times, so by default tool calls never r | Strategy | Behaviour | | ----------------- | ----------------------------------------------------------------------------------------------------------------------- | -| **Mock all** | Every tool answers from a mock. Tools without their own entry use their default mock response. This is the default. | +| **Mock all** | Every tool answers from a mock. This is the default. See below for tools without an entry. | | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions in mock conditions use JavaScript syntax, so Python forms such as `(?i)` or `(?P...)` are refused when you save. + +Under **Mock all**, a tool without an entry in the test answers with its own first mock response, and a tool with neither, integrations included, returns an empty success result. The agent then carries on as if the call worked, so add an entry wherever the conversation depends on what the tool returns. + +Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked, and an assertion on it fails with _This tool is not on the agent_. ### Assertions Under **Advanced**, add deterministic checks that run alongside the judge: -- **Required tool calls**: the agent must call the tool, optionally with parameters that match. +- **Required tool calls**: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or to 0 to forbid calls with those arguments. - **Forbidden tools**: the agent must not call the tool. - **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. @@ -108,7 +117,9 @@ A failed assertion fails the run regardless of the judge's verdict. Each asserti ### Repeat count -Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`. +Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`, and the test's result lets you switch between repeats to read each conversation. + +Repeats only apply to **Run all**. Running the test from its **Run** tab, or re-running one test from the row menu, runs the conversation once. ## Attach tests to agents @@ -131,7 +142,7 @@ The verdict card shows **Pass** or **Fail** with the time the run took, followed | **Tool** | Which tool the agent called, if any, against what the test expected. | | **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked **Mock** when it answered from a mock. | -A Multi Turn run with an **Unknown** condition shows **Needs review** instead of a verdict. If a run can't complete, the card shows **Error** with the error message. When the conversation finished but could not be judged, the transcript is still shown below the error. +A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with the error message. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. ## Run every test for an agent @@ -160,6 +171,7 @@ Tests always exercise the agent's latest **draft** configuration, including unpu | ------------------------------- | --------------------------------------------------------------------- | | Test name | 200 characters | | Conversation message | 2,000 characters each, at least one message (optional for Multi Turn) | +| Conversation (Multi Turn) | 50 messages | | Expectation | 400 characters | | Success / failure example | 400 characters each | | Scenario (Multi Turn) | 10,000 characters | From 9603f069b0d68a6156eb64ce5ddf1f089b8ff1dd Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 04:25:30 +0900 Subject: [PATCH 08/13] docs(agent-tests): unmocked tools fail, integration mocks, channel rules, error causes Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 30 +++++++++++++++--------------- 1 file changed, 15 insertions(+), 15 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index b89bffe..9588f2d 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -58,7 +58,7 @@ Under **Conversation**, click **Add message** to build the history the agent see A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it. -Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called. +Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Multi Turn test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called. ## Multi Turn tests @@ -81,7 +81,7 @@ Two rules make scenarios reliable. Reveal one fact at a time, so the agent has t ### Channel -**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. A phone agent, for example, keeps replies short and reads numbers back digit by digit. **Auto**, the default, uses Phone inbound when the agent answers a phone number and Web voice otherwise. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. +**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. A phone agent, for example, keeps replies short and repeats important numbers back in small groups. **Auto**, the default, uses Phone inbound when a phone number is assigned to the agent and Web voice otherwise, so pick Phone outbound yourself for an agent that only places calls. As in production, client tools are not available on phone channels, and transfers only happen on phone channels. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. ### Success conditions @@ -95,21 +95,21 @@ Simulated conversations usually run many times, so by default tool calls never r | Strategy | Behaviour | | ----------------- | ----------------------------------------------------------------------------------------------------------------------- | -| **Mock all** | Every tool answers from a mock. This is the default. See below for tools without an entry. | +| **Mock all** | Every tool answers from a mock. This is the default. A tool without any mock fails, see below. | | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions in mock conditions use JavaScript syntax, so Python forms such as `(?i)` or `(?P...)` are refused when you save. +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions in mock conditions use JavaScript syntax, so Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. -Under **Mock all**, a tool without an entry in the test answers with its own first mock response, and a tool with neither, integrations included, returns an empty success result. The agent then carries on as if the call worked, so add an entry wherever the conversation depends on what the tool returns. +Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither returns the error `No mock is configured for in this test.`, so the conversation shows the gap instead of passing on a result the tool never produced. The test form lists the agent's tools that would fail this way. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. To match production, mock both, or have the write tool's mock return the final result. -Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked, and an assertion on it fails with _This tool is not on the agent_. +Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with _This tool is not on the agent_, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as "never call issue_refund" passing on agents that cannot call the tool. ### Assertions Under **Advanced**, add deterministic checks that run alongside the judge: -- **Required tool calls**: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or to 0 to forbid calls with those arguments. +- **Required tool calls**: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or set both the minimum and the maximum to 0 to forbid calls with those arguments. - **Forbidden tools**: the agent must not call the tool. - **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. @@ -117,7 +117,7 @@ A failed assertion fails the run regardless of the judge's verdict. Each asserti ### Repeat count -Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`, and the test's result lets you switch between repeats to read each conversation. +Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`, and opening the test from there lets you switch between the finished repeats of the latest batch to read each conversation. Repeats only apply to **Run all**. Running the test from its **Run** tab, or re-running one test from the row menu, runs the conversation once. @@ -136,17 +136,17 @@ Open the test and switch to its **Run** tab. Pick an agent that has access, then The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type: -| Type | What you see | -| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | -| **Tool** | Which tool the agent called, if any, against what the test expected. | -| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation and after how many turns, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked **Mock** when it answered from a mock. | +| Type | What you see | +| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | +| **Tool** | Which tool the agent called, if any, against what the test expected. | +| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **No mock** or **No matching mock** when it failed for lack of one. | -A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with the error message. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. +A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its language model stopped responding, or the platform did. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. ## Run every test for an agent -On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Multi Turn tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. Use the row menu to re-run a single test, edit it, or remove it from the agent. +On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Multi Turn tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent. ## Tests run against the draft From 03c8e9c092ba1e3865606785f3cc2c5a641a60d5 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 05:07:23 +0900 Subject: [PATCH 09/13] docs(agent-tests): one regex dialect, channel details, repeat counts, error causes Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 9588f2d..474be59 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -81,7 +81,7 @@ Two rules make scenarios reliable. Reveal one fact at a time, so the agent has t ### Channel -**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. A phone agent, for example, keeps replies short and repeats important numbers back in small groups. **Auto**, the default, uses Phone inbound when a phone number is assigned to the agent and Web voice otherwise, so pick Phone outbound yourself for an agent that only places calls. As in production, client tools are not available on phone channels, and transfers only happen on phone channels. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. +**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. On a phone channel, for example, the agent repeats important numbers back in small groups. Phone outbound speaks in the same phone register as Phone inbound and only `{{system.channel}}` differs, so if the agent should open an outbound call differently, branch on `{{system.channel}}` in its prompt. **Auto**, the default, uses Phone inbound when a phone number is assigned to the agent and Web voice otherwise, so pick Phone outbound yourself for an agent that only places calls. As in production, client tools are not available on phone channels, and transfers only happen on phone channels. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. No real call is placed, so `{{system.caller_number}}` and `{{system.dialed_number}}` are empty on every channel. ### Success conditions @@ -99,7 +99,7 @@ Simulated conversations usually run many times, so by default tool calls never r | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions in mock conditions use JavaScript syntax, so Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Unicode forms such as `\p{L}`, `\u{1F600}` or `\z` are refused too, because without Unicode mode they would match their literal text. List the characters in a class instead, for example `[A-Za-zÀ-ÿ]`, and write `\z` as `$`. Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither returns the error `No mock is configured for in this test.`, so the conversation shows the gap instead of passing on a result the tool never produced. The test form lists the agent's tools that would fail this way. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. To match production, mock both, or have the write tool's mock return the final result. @@ -111,13 +111,13 @@ Under **Advanced**, add deterministic checks that run alongside the judge: - **Required tool calls**: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or set both the minimum and the maximum to 0 to forbid calls with those arguments. - **Forbidden tools**: the agent must not call the tool. -- **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. +- **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. A transfer only happens on a phone channel, so on Web voice a transfer check always fails. A failed assertion fails the run regardless of the judge's verdict. Each assertion shows **Passed** or **Failed** with a short detail in the result. ### Repeat count -Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed, for example `2 / 3 passed`, and opening the test from there lets you switch between the finished repeats of the latest batch to read each conversation. +Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed out of those that reached a verdict, for example `2 / 3 passed`, or `2 / 2 passed · 1 error` when one repeat ended in **Error**. While repeats are still running, the row counts them instead, for example `1 passed · 2 running`. Opening the test from that page lets you switch between the finished repeats of the latest batch to read each conversation. The test opens on the first failed repeat. Repeats only apply to **Run all**. Running the test from its **Run** tab, or re-running one test from the row menu, runs the conversation once. @@ -142,7 +142,7 @@ The verdict card shows **Pass** or **Fail** with the time the run took, followed | **Tool** | Which tool the agent called, if any, against what the test expected. | | **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **No mock** or **No matching mock** when it failed for lack of one. | -A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its language model stopped responding, or the platform did. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. +A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. ## Run every test for an agent From e2a03a72261c628f7b0255dc8c496a821ff62337 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 05:39:04 +0900 Subject: [PATCH 10/13] docs(agent-tests): missing mocks need review, real calls, error entries, migration note Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 18 +++++++++++------- 1 file changed, 11 insertions(+), 7 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 474be59..104bd77 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -99,9 +99,13 @@ Simulated conversations usually run many times, so by default tool calls never r | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Unicode forms such as `\p{L}`, `\u{1F600}` or `\z` are refused too, because without Unicode mode they would match their literal text. List the characters in a class instead, for example `[A-Za-zÀ-ÿ]`, and write `\z` as `$`. +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Unicode forms such as `\p{L}`, `\u{1F600}` or `\z` are refused too, because without Unicode mode they would match their literal text. List the characters in a class instead, for example `[A-Za-zÀ-ÿ]`, and write `\z` as `$`. -Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither returns the error `No mock is configured for in this test.`, so the conversation shows the gap instead of passing on a result the tool never produced. The test form lists the agent's tools that would fail this way. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. To match production, mock both, or have the write tool's mock return the final result. +Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. A client tool that doesn't wait for an answer gets its usual acknowledgement instead. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there, so a tool you attach later still fails safely until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. To match production, mock both, or have the write tool's mock return the final result. + +A run in which a tool call failed for lack of a mock, or matched none of its tool's entries, never passes. It fails with **Needs review** and **No mock** badges, and the result names the tool and says the failure comes from the test setup, not the agent. This holds even when the judge and every assertion passed, because the conversation left the path it would take in production. The judge is told that such a failure is a gap in the test's setup and scores only how the agent handled it. Add the missing mock and run the test again. + +If you are used to tools without a mock calling their real endpoint, your first runs will show **No mock** on those calls. Add a mock for each tool the test form lists, or add read-only tools to **Call the real endpoint**. Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with _This tool is not on the agent_, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as "never call issue_refund" passing on agents that cannot call the tool. @@ -136,11 +140,11 @@ Open the test and switch to its **Run** tab. Pick an agent that has access, then The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type: -| Type | What you see | -| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | -| **Tool** | Which tool the agent called, if any, against what the test expected. | -| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **No mock** or **No matching mock** when it failed for lack of one. | +| Type | What you see | +| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. | +| **Tool** | Which tool the agent called, if any, against what the test expected. | +| **Multi Turn** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. | A Multi Turn run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged. From c62c16a612bea2157dbd5185d51ce06754d7a6f4 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 06:22:14 +0900 Subject: [PATCH 11/13] docs(agent-tests): unanswered calls, acknowledged tools, regex limits, Confirm write Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 104bd77..63b2b42 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -99,11 +99,11 @@ Simulated conversations usually run many times, so by default tool calls never r | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Unicode forms such as `\p{L}`, `\u{1F600}` or `\z` are refused too, because without Unicode mode they would match their literal text. List the characters in a class instead, for example `[A-Za-zÀ-ÿ]`, and write `\z` as `$`. +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as `\p{L}` or `\u{1F600}` (list the characters in a class instead, for example `[A-Za-zÀ-ÿ]`) and Python escapes such as `\A`, `\z` or `\N{...}` (write `\A` as `^` and `\z` as `$`). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII letters under `(?i:...)`, `\S` inside a class and `\c` are refused as well. -Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. A client tool that doesn't wait for an answer gets its usual acknowledgement instead. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there, so a tool you attach later still fails safely until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. To match production, mock both, or have the write tool's mock return the final result. +Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don't wait for an answer and webhook tools in fire and forget mode, get their usual acknowledgement instead and need no mock. Such a client tool can't be mocked at all, since production always acknowledges it. A simulation has no client, so a client tool that waits for an answer needs a mock under every strategy. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review. -A run in which a tool call failed for lack of a mock, or matched none of its tool's entries, never passes. It fails with **Needs review** and **No mock** badges, and the result names the tool and says the failure comes from the test setup, not the agent. This holds even when the judge and every assertion passed, because the conversation left the path it would take in production. The judge is told that such a failure is a gap in the test's setup and scores only how the agent handled it. Add the missing mock and run the test again. +A run in which a tool call went unanswered never passes: the tool had no mock, or none of its entries matched and neither the tool's own mock response nor **Call the real endpoint** could answer. The run fails with **Needs review** and **No mock** badges and lists the unanswered calls, even when the judge and every assertion passed, because the conversation left the path it would take in production. A call with no matching entry means either the agent passed arguments you didn't expect or the entry's conditions are too narrow. Calls to a tool the test forbids are not listed, they are the agent's failure and the assertion reports them. The judge is told the call failed because the test prepared no answer for it, and judges the agent as usual without treating the call as having run. Add or widen the mock and run the test again. If you are used to tools without a mock calling their real endpoint, your first runs will show **No mock** on those calls. Add a mock for each tool the test form lists, or add read-only tools to **Call the real endpoint**. From 56e2f7be9f360da13007563dadd19f2d7c178382 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 07:00:02 +0900 Subject: [PATCH 12/13] docs(agent-tests): client tools answer from mocks, fire-and-forget tools always acknowledge Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index 63b2b42..ed0d40b 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -93,15 +93,15 @@ The judge reads the whole transcript and marks every condition **Success**, **Fa Simulated conversations usually run many times, so by default tool calls never reach your real endpoints: -| Strategy | Behaviour | -| ----------------- | ----------------------------------------------------------------------------------------------------------------------- | -| **Mock all** | Every tool answers from a mock. This is the default. A tool without any mock fails, see below. | -| **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | -| **Mock none** | Every tool calls its real endpoint. Webhooks may create real side effects, so reserve this for a safe test environment. | +| Strategy | Behaviour | +| ----------------- | ----------------------------------------------------------------------------------------------------------------------------- | +| **Mock all** | Every tool answers from a mock. This is the default. A tool without any mock fails, see below. | +| **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | +| **Mock none** | Every tool with a real endpoint calls it. Webhooks may create real side effects, so reserve this for a safe test environment. | A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as `\p{L}` or `\u{1F600}` (list the characters in a class instead, for example `[A-Za-zÀ-ÿ]`) and Python escapes such as `\A`, `\z` or `\N{...}` (write `\A` as `^` and `\z` as `$`). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII letters under `(?i:...)`, `\S` inside a class and `\c` are refused as well. -Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don't wait for an answer and webhook tools in fire and forget mode, get their usual acknowledgement instead and need no mock. Such a client tool can't be mocked at all, since production always acknowledges it. A simulation has no client, so a client tool that waits for an answer needs a mock under every strategy. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review. +Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don't wait for an answer and webhook tools in fire and forget mode, need no mock: as in production, the agent only ever hears their acknowledgement, whether the call ran for real, answered from a mock, or was skipped. Such a client tool can't be mocked at all. A simulation has no client, so a client tool that waits for an answer always answers from its mock entries, under every strategy including **Mock none**, and needs one. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review. A run in which a tool call went unanswered never passes: the tool had no mock, or none of its entries matched and neither the tool's own mock response nor **Call the real endpoint** could answer. The run fails with **Needs review** and **No mock** badges and lists the unanswered calls, even when the judge and every assertion passed, because the conversation left the path it would take in production. A call with no matching entry means either the agent passed arguments you didn't expect or the entry's conditions are too narrow. Calls to a tool the test forbids are not listed, they are the agent's failure and the assertion reports them. The judge is told the call failed because the test prepared no answer for it, and judges the agent as usual without treating the call as having run. Add or widen the mock and run the test again. From a27c8b4d15188aa022bb7619f4f446a5df77d569 Mon Sep 17 00:00:00 2001 From: Him188 Date: Mon, 28 Sep 2026 07:17:07 +0900 Subject: [PATCH 13/13] docs(agent-tests): assertion regex limits wording Co-Authored-By: Claude Opus 5.5 --- agents/test/agent-tests.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/agents/test/agent-tests.mdx b/agents/test/agent-tests.mdx index ed0d40b..26fed0d 100644 --- a/agents/test/agent-tests.mdx +++ b/agents/test/agent-tests.mdx @@ -99,7 +99,7 @@ Simulated conversations usually run many times, so by default tool calls never r | **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. | | **Mock none** | Every tool with a real endpoint calls it. Webhooks may create real side effects, so reserve this for a safe test environment. | -A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as `\p{L}` or `\u{1F600}` (list the characters in a class instead, for example `[A-Za-zÀ-ÿ]`) and Python escapes such as `\A`, `\z` or `\N{...}` (write `\A` as `^` and `\z` as `$`). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII letters under `(?i:...)`, `\S` inside a class and `\c` are refused as well. +A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P...)` are refused when you save. Write a named group as `(?...)`, and ignore case with `(?i:...)` around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as `\p{L}` or `\u{1F600}` (list the characters in a class instead, for example `[A-Za-zÀ-ÿ]`) and Python escapes such as `\A`, `\z` or `\N{...}` (write `\A` as `^` and `\z` as `$`). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII characters under `(?i:...)`, `\S` inside a class without `\s` (`[\s\S]` is fine) and `\c` are refused as well. Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don't wait for an answer and webhook tools in fire and forget mode, need no mock: as in production, the agent only ever hears their acknowledgement, whether the call ran for real, answered from a mock, or was skipped. Such a client tool can't be mocked at all. A simulation has no client, so a client tool that waits for an answer always answers from its mock entries, under every strategy including **Mock none**, and needs one. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review.