Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions agents/test/agent-tests.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -107,7 +107,7 @@

If you are used to tools without a mock calling their real endpoint, your first runs will show **No mock** on those calls. Add a mock for each tool the test form lists, or add read-only tools to **Call the real endpoint**.

Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with _This tool is not on the agent_, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as "never call issue_refund" passing on agents that cannot call the tool.

Check warning on line 110 in agents/test/agent-tests.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/test/agent-tests.mdx#L110

Did you really mean 'issue_refund'?

### Assertions

Expand Down Expand Up @@ -169,6 +169,52 @@
</Step>
</Steps>

## Run tests from the API and CI

Everything on this page is also available over the REST API with an [API key](/developer-guide/getting-started/api-key), so you can keep tests next to your agent configuration and gate releases on them. See the [Tests API reference](/api-reference/endpoint/agent/list-tests) for every endpoint.

- **Export and import.** [Get Test](/api-reference/endpoint/agent/get-test) returns a test's full definition. Post that body to [Create Test](/api-reference/endpoint/agent/create-test) as is to recreate it, for example from JSON files kept in your repository. Read-only fields such as `test_id` and `created_at` are ignored. Tool references use your workspace's tool ids, so an export imports as is within the same team. [List Test Tools](/api-reference/endpoint/agent/list-test-tools) returns the ids a test can reference for an agent, including integration tools.

Check warning on line 176 in agents/test/agent-tests.mdx

View check run for this annotation

Mintlify / Mintlify Validation (hanabiaiinc) - vale-spellcheck

agents/test/agent-tests.mdx#L176

Did you really mean 'workspace's'?
- **Attach.** Pass `agent_ids` when you create a test, or attach it later with [Attach Test](/api-reference/endpoint/agent/attach-test).
- **Run and wait.** [Run Tests](/api-reference/endpoint/agent/run-tests) starts a batch with every attached test, or only the `test_ids` you pass, and `repeat_count` overrides each test's repeat count for that batch. Poll [Get Test Batch](/api-reference/endpoint/agent/get-test-batch) until `completed` is `true`, then check `passed`, `failed`, `errors`, and `pass_rate`. `pass_rate` leaves out runs that ended in **Error**, so decide in your pipeline whether an error should fail the build or be retried.
- **History.** [List Test Batches](/api-reference/endpoint/agent/list-test-batches) and [List Test Runs](/api-reference/endpoint/agent/list-test-runs) return past results, filterable by test, batch, and status.

Runs always use the agent's current draft. A pipeline that changes the agent should update the draft first ([Configure through the API](/agents/build/configuration#configure-through-the-api)), run the tests, and publish only when they pass.

This GitHub Actions workflow runs every test attached to an agent and fails the build if any run fails or errors:

```yaml .github/workflows/agent-tests.yml
name: Agent tests
on: [pull_request]

jobs:
agent-tests:
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- name: Run the agent's tests
env:
FISH_API_KEY: ${{ secrets.FISH_API_KEY }}
AGENT_ID: ${{ vars.AGENT_ID }}
run: |
set -euo pipefail
api=https://api.fish.audio/v1/agent
auth="Authorization: Bearer $FISH_API_KEY"

batch=$(curl -sf -X POST "$api/agents/$AGENT_ID/tests/run" \
-H "$auth" -H "Content-Type: application/json" -d '{}' | jq -r .batch_id)

while true; do
result=$(curl -sf "$api/agents/$AGENT_ID/test-batches/$batch" -H "$auth")
[ "$(jq -r .completed <<<"$result")" = "true" ] && break
sleep 15
done

jq '{passed, failed, errors, pass_rate}' <<<"$result"
jq -e '.failed == 0 and .errors == 0' <<<"$result" > /dev/null
```

Simulation tests take minutes, so give the job a timeout that covers your longest scenarios. To tolerate occasional failures of repeated Simulation tests, gate on `pass_rate` instead, for example `jq -e '.pass_rate >= 0.8'`.

## Limits

| Field | Limit |
Expand All @@ -182,6 +228,7 @@
| Success conditions (Simulation) | 1 to 10, description 500 characters each |
| Max turns (Simulation) | 50 |
| Repeat count (Simulation) | 20 |
| Runs per API batch | 1,000 (tests times repeats) |
| Dynamic variables | 50 per test |

## Going further
Expand Down
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/attach-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: put /v1/agent/agents/{agent_id}/tests/{test_id}
title: "Attach Test"
description: "Attach a test to an agent so the agent's test runs include it."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/create-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: post /v1/agent/tests
title: "Create Test"
description: "Create a Single Turn, Tool or Simulation test, optionally attached to agents."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/delete-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: delete /v1/agent/tests/{test_id}
title: "Delete Test"
description: "Delete a test and detach it from every agent."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/detach-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: delete /v1/agent/agents/{agent_id}/tests/{test_id}
title: "Detach Test"
description: "Detach a test from an agent. The test stays in your library."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/get-test-batch.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/agents/{agent_id}/test-batches/{batch_id}
title: "Get Test Batch"
description: "The runs of one batch with their pass, fail and error counts."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/get-test-run.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/test-runs/{run_id}
title: "Get Test Run"
description: "Fetch one test run with its transcript, verdicts and tool calls."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/get-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/tests/{test_id}
title: "Get Test"
description: "Fetch one test's full definition, which is also the export format."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/list-test-batches.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/agents/{agent_id}/test-batches
title: "List Test Batches"
description: "The agent's test batches, newest first, with their counts."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/list-test-runs.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/test-runs
title: "List Test Runs"
description: "List your team's test runs, newest first."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/list-test-tools.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/agents/{agent_id}/test-tools
title: "List Test Tools"
description: "The tool ids a test can reference for this agent."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/list-tests.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: get /v1/agent/tests
title: "List Tests"
description: "List your team's tests, newest first."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/run-tests.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: post /v1/agent/agents/{agent_id}/tests/run
title: "Run Tests"
description: "Start a batch that runs the agent's attached tests against its draft."
icon: "flask"
iconType: "solid"
---
7 changes: 7 additions & 0 deletions api-reference/endpoint/agent/update-test.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
openapi: patch /v1/agent/tests/{test_id}
title: "Update Test"
description: "Patch test fields. Omitted fields keep their value."
icon: "flask"
iconType: "solid"
---
Loading
Loading