diff --git a/.mintlify/skills/fish-audio-api/SKILL.md b/.mintlify/skills/fish-audio-api/SKILL.md index c6ba0c5..d6d945b 100644 --- a/.mintlify/skills/fish-audio-api/SKILL.md +++ b/.mintlify/skills/fish-audio-api/SKILL.md @@ -20,7 +20,7 @@ This file condenses those into rules an agent can apply directly. - Optional distributed tracing for inference APIs: see `https://docs.fish.audio/api-reference/observability`. - Get API keys: `https://fish.audio/app/api-keys` - Never hardcode keys. Read from an env var like `FISH_API_KEY`. -- Errors are JSON `{status, message}` for 401 / 402 / 404, and an array of `{loc, type, msg, ctx, in}` for 422 (validation). +- Errors are JSON with at least `{status, message}` (for example 401 / 402 / 404 / 429). Speech-to-text errors from `transcribe-1-pro` also carry `code` and `request_id`. Errors from the network edge in front of the API (some 413 and 5xx responses) may not be JSON. `422` validation errors, on the endpoints that return them, are an array of `{loc, type, msg, ctx, in}`. ## Endpoint map @@ -199,31 +199,92 @@ await pipeline(Readable.fromWeb(res.body), createWriteStream("out.mp3")); ## Speech-to-Text: `POST /v1/asr` -Required headers: `Authorization`. Content type: `multipart/form-data` or `application/msgpack`. +Synchronous: one audio file per request, and the JSON response arrives when the whole file has been transcribed. -Form fields: +Use `transcribe-1-pro`, the recommended model, and select it with the `model` header on every request. This section describes `transcribe-1-pro`; for the differences on `transcribe-1`, see "If you use `transcribe-1`" at the end of this section. -- `audio` (binary, required) -- `language` (string, optional; omit to auto-detect) -- `ignore_timestamps` (bool, default `true`; set `false` to get per-segment timestamps, which adds latency on clips < 30 s) +Guide: `https://docs.fish.audio/features/speech-to-text`. The `/v1/asr` entry in `openapi.json` currently does not list the `transcribe-1-pro` fields, `request_id`, `speaker_turns`, or the error `code`; where it is less complete or disagrees, follow this section. -Response (200): +Headers: + +- `Authorization: Bearer ` (required). +- `model: transcribe-1-pro` (recommended): recordings up to 60 minutes, including multi-speaker conversations, with speaker markers, `speaker_turns`, and emotion and vocal-event cues. Write the value exactly, in lowercase. The header is optional, but a missing or unrecognized value (for example `Transcribe-1-Pro` or `transcribe-1pro`) is served and billed as `transcribe-1`, and no error is returned; if you expected Pro but `text` has no speaker markers, check the header. `model` is a header only; a `model` form field is ignored. +- Body encoding: `multipart/form-data` (let the HTTP client set `Content-Type` and the boundary) or `Content-Type: application/msgpack`. Base64-encoded audio in a JSON body is not supported. + +Request fields (the same names in multipart and MessagePack): + +| Field | Default | Notes | +| ------------------------------ | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `audio` | — (required) | Exactly one audio file: a multipart file part, or MessagePack `bin`. The format is read from the bytes, not the file name. | +| `language` | unset | Optional hint, a lowercase ISO 639-1 code such as `en`, `zh`, or `ja`. Detection still runs, and the hint does not force the transcript language. Other forms, such as `en-US` or `English`, may be rejected with 400. | +| `ignore_timestamps` | `true` | `false` returns word-level `segments` and, unless `diarize=false`, `speaker_turns`; it adds processing time. Multipart: any value other than `true` (case-insensitive), including `1` or an empty value, means `false`. MessagePack: a boolean. | +| `tag_audio_events` | `true` | `false` removes bracketed cues such as `[laughter]` or `[高兴]` from `text` and `speaker_turns`. Timestamps, `duration`, and billing do not change. | +| `diarize` | `auto` | `auto` or `true`: return `speaker_turns` when timestamps are requested. `false`: omit `speaker_turns`; the transcript and its speaker markers do not change. | +| `num_speakers` | unset | Expected number of speakers, an integer ≥ 1. A best-effort hint that may have no effect on short recordings. Cannot be combined with `min_speakers` or `max_speakers`. | +| `min_speakers`, `max_speakers` | unset | Bounds on the number of speakers, integers ≥ 1, with `min_speakers` ≤ `max_speakers`. Same best-effort rule. | + +Multipart values: `tag_audio_events` is `true` or `false` and `diarize` is `auto`, `true`, or `false` (both case-insensitive); speaker counts are digits. MessagePack values: `tag_audio_events` is a boolean, `diarize` is a boolean or `"auto"`/`"true"`/`"false"`, and speaker counts are integers. An invalid value, `tag_audio_events`, `diarize`, or a speaker count sent twice in a multipart form, or a speaker count with `diarize=false` returns 400 `invalid_parameter`. + +Limits and formats: + +- `transcribe-1-pro` accepts recordings up to 60 minutes; longer audio returns 400 `audio_too_long`. Send long recordings as compressed audio (MP3, Opus, or AAC); an hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413. Send a whole conversation as one file: speaker labels are consistent within one response, not across requests. +- Processing time grows with the length of the recording, and long `transcribe-1-pro` requests can take several minutes. Set a generous client timeout (the examples use 15 minutes); many HTTP libraries default to much less (httpx: 5 s). If a long request fails with a 5xx error or the connection drops, retry it. +- `transcribe-1-pro` accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file; send the original file bytes. AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM return 400. + +Response (200), an illustrative `transcribe-1-pro` response with `ignore_timestamps=false`: ```json { - "text": "full transcript", - "duration": 12.34, - "segments": [{"text": "...", "start": 0.0, "end": 1.23}] + "text": "<|speaker:0|> Thanks for joining. <|speaker:1|> [laughter] Happy to be here.", + "duration": 4.2, + "segments": [ + { "text": "Thanks", "start": 0.16, "end": 0.48 }, + { "text": "for", "start": 0.48, "end": 0.64 }, + { "text": "joining", "start": 0.64, "end": 1.12 }, + { "text": "Happy", "start": 2.24, "end": 2.56 }, + { "text": "to", "start": 2.56, "end": 2.68 }, + { "text": "be", "start": 2.68, "end": 2.8 }, + { "text": "here", "start": 2.8, "end": 3.2 } + ], + "language": "English", + "language_code": "en", + "request_id": "3f6c2a1e-8b4d-4c7a-9e21-5d0b7f9a6c13", + "speaker_turns": [ + { + "speaker": "speaker:0", + "text": "Thanks for joining.", + "start": 0.16, + "end": 1.12 + }, + { + "speaker": "speaker:1", + "text": "[laughter] Happy to be here.", + "start": 2.24, + "end": 3.2 + } + ] } ``` +Response fields: + +- `text`: The full transcript, with inline `<|speaker:N|>` markers, usually with a space on each side, and, unless `tag_audio_events=false`, bracketed cues. Text before the first marker belongs to the first turn; a transcript with no marker is a single speaker (speaker 0). +- `duration`: Audio length in seconds, including silence. +- `segments`: Word-level timestamps `{ text, start, end }` in seconds. Usually one word (one or a few characters in Chinese and Japanese), with no punctuation, markers, or cues; it can be normalized (`35` for `3.5`), so it does not always match `text`. `start` can equal `end`. `[]` (never omitted) when `ignore_timestamps=true`, when no speech was found, or when timing is temporarily unavailable. Segments are not speaker turns. +- `language`: The detected language's English name, such as `English` or `Chinese`. Omitted when it cannot be determined. +- `language_code`: ISO 639-1 code for `language`, such as `en`. Omitted when unknown. If the language cannot be determined and you sent a `language` hint, it reports your hint, unchecked. Responses report one language, even for recordings that switch languages. +- `request_id`: Unique ID for the request, also in error bodies and in the `x-request-id` response header. Include it when you contact support. +- `speaker_turns`: Present only when `ignore_timestamps=false` and `diarize` is not `false`. Turns in the order they occur (`[]` when no speech was found), each `{ speaker: "speaker:N", text, start, end }`. `N` matches the `<|speaker:N|>` marker in `text`; labels identify speakers within one response only. Turn `text` has no markers and keeps cues unless `tag_audio_events=false`. If `segments` is empty, turn times are approximate and can cover the whole recording. Consecutive turns can have the same speaker; do not assume turns are contiguous or non-overlapping. Prefer `speaker_turns` over parsing `text`. + +`language` and `language_code` are omitted, never `null`. Do not depend on the order of keys in the JSON. + ### curl ```bash curl --request POST https://api.fish.audio/v1/asr \ --header "Authorization: Bearer $FISH_API_KEY" \ + --header "model: transcribe-1-pro" \ --form "audio=@input.wav" \ - --form "language=en" \ --form "ignore_timestamps=false" ``` @@ -235,15 +296,72 @@ import os, httpx with open("input.wav", "rb") as f: r = httpx.post( "https://api.fish.audio/v1/asr", - headers={"Authorization": f"Bearer {os.environ['FISH_API_KEY']}"}, + headers={ + "Authorization": f"Bearer {os.environ['FISH_API_KEY']}", + "model": "transcribe-1-pro", + }, files={"audio": f}, - data={"language": "en", "ignore_timestamps": "false"}, - timeout=120, + data={"ignore_timestamps": "false"}, + # Long Pro recordings can take several minutes; httpx defaults to 5 s. + timeout=httpx.Timeout(900.0, connect=10.0), ) -r.raise_for_status() -print(r.json()["text"]) +if r.is_error: + try: + err = r.json() # {status, message}; transcribe-1-pro adds code, request_id + except ValueError: + err = {"status": r.status_code, "message": r.text} # edge errors may not be JSON + raise RuntimeError(f"ASR failed: {err}") +result = r.json() +print(result["text"]) +for turn in result.get("speaker_turns", []): + print(f"{turn['speaker']} [{turn['start']:.2f}-{turn['end']:.2f}] {turn['text']}") +``` + +### Node.js (fetch) + +```js +import { readFile } from "node:fs/promises"; +// npm install undici@7 (Node.js 20.18.1+; undici 8 needs Node.js 22.19+) +import { Agent, setGlobalDispatcher } from "undici"; + +// Node's fetch waits only 300 s for response headers; long Pro requests can take longer. +setGlobalDispatcher( + new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 }) +); + +const form = new FormData(); +form.append("audio", new Blob([await readFile("input.wav")]), "input.wav"); +form.append("ignore_timestamps", "false"); + +const res = await fetch("https://api.fish.audio/v1/asr", { + method: "POST", + headers: { + Authorization: `Bearer ${process.env.FISH_API_KEY}`, + model: "transcribe-1-pro", + }, + body: form, + signal: AbortSignal.timeout(900_000), +}); +if (!res.ok) throw new Error(`${res.status} ${await res.text()}`); +const result = await res.json(); +console.log(result.text); +for (const turn of result.speaker_turns ?? []) { + console.log(`${turn.speaker} [${turn.start}-${turn.end}] ${turn.text}`); +} ``` +MessagePack instead of multipart: send `msgpack.packb({"audio": audio_bytes, "ignore_timestamps": False}, use_bin_type=True)` with `Content-Type: application/msgpack` and the same `model: transcribe-1-pro` header; values are typed (booleans, integers), and `audio` must be raw bytes (`bin`). + +### If you use `transcribe-1` + +`transcribe-1` is for general transcription of short recordings, and it serves every request whose `model` header is missing or not an exact match. It differs from `transcribe-1-pro` as follows: + +- Fields: send only `audio`, `language`, and `ignore_timestamps`. The other fields apply to `transcribe-1-pro`; send them only with `model: transcribe-1-pro`. +- Response: `text`, `duration`, `segments`, and, when the language is known, `language` and `language_code`. Speaker markers, `speaker_turns`, and `request_id` are `transcribe-1-pro` features. +- Errors: rely only on `status` and `message`. +- Limits: up to 50 MiB per request; keep MP3 and Opus files under 25 MiB. A request over the size limit returns 413 or 400. For recordings longer than a few minutes, use `transcribe-1-pro`. A long request can fail with 503 when it exceeds the processing-time limit; retry, and if it keeps failing, use `transcribe-1-pro` or split the audio. +- Formats: WAV, MP3, AAC (including M4A/MP4), FLAC, and Ogg (Opus or Vorbis). Convert WebM recordings (for example, from a browser's MediaRecorder) to Ogg/Opus, MP3, or WAV first, or use `transcribe-1-pro`. + ## Voice Design: `POST /v1/voice-design` Required headers: @@ -495,7 +613,7 @@ The S1 model uses `(parenthesis)` tags inside `text`, e.g. `(happy) What a day!` - Use `application/json` for normal TTS requests. It's the simplest and works for `reference_id` flows. - Use `application/msgpack` when you need to send raw audio bytes inline (inline `references`, or the WebSocket protocol). -- Use `multipart/form-data` for `/v1/asr` and `POST /model` because they upload files. +- Use `multipart/form-data` for `/v1/asr` and `POST /model` because they upload files. `/v1/asr` also accepts `application/msgpack`; it does not accept base64 audio in JSON. - All WebSocket frames are MessagePack binary, regardless of inner payload. ## Error handling checklist @@ -509,6 +627,13 @@ The S1 model uses `(parenthesis)` tags inside `text`, e.g. `(happy) What a day!` - Numeric param out of range (`temperature`, `top_p`, `chunk_length`, `min_chunk_length`, `prosody.speed`, `early_stop_threshold`). - `mp3_bitrate` / `opus_bitrate` set without matching `format`. - WebSocket: a `finish` event with `reason: "error"` means the server failed mid-stream. Surface the message and reconnect rather than retrying on the same socket. +- `POST /v1/asr` (`transcribe-1-pro`): branch on the HTTP status and on `code`, never on `message`. New `code` values may be added; handle unknown codes by HTTP status. With `transcribe-1`, rely only on `status` and `message`. + - 400 → fix the request; do not retry. Codes: `invalid_request` (unreadable body, missing or repeated `audio`), `invalid_parameter` (bad or conflicting field values), `invalid_audio` (undecodable or unsupported format), `audio_too_long` (over 60 minutes), `audio_too_short` (under about 0.08 s). + - 413 → request too large (`request_too_large`). It may come from the network edge without a JSON body. Send compressed audio. + - 415 → unsupported `Content-Type` (`unsupported_media_type`). Use `multipart/form-data` or `application/msgpack`. + - 429 → usually your account is at its concurrency limit, shared by all its API keys; each `/v1/asr` request holds a slot until its response returns. No `Retry-After` header is sent; retry with exponential backoff. + - 500 (`internal_error`), 502, 503 (`upstream_unavailable`, `upstream_timeout`, `excessive_repetition`, `diarization_failed`), 504 → temporary; retry with exponential backoff. + - Any other 4xx (`upstream_rejected`, rare) → do not retry. ## Decision shortcuts @@ -517,4 +642,4 @@ The S1 model uses `(parenthesis)` tags inside `text`, e.g. `(happy) What a day!` - User wants dialogue between multiple speakers → `POST /v1/tts` on `s2.1-pro` with `reference_id` array and `<|speaker:N|>` tags. - User is streaming tokens from an LLM and wants speech to play as it arrives → WebSocket `/v1/tts/live`. - User wants a persistent custom voice they can reuse → `POST /model` first, then reuse the returned `_id` as `reference_id`. -- User wants a transcript → `POST /v1/asr`. +- User wants a transcript → `POST /v1/asr` with the `model: transcribe-1-pro` header (send it on every request; without it, the request runs on `transcribe-1`). diff --git a/.mintlify/skills/fish-audio-sdk/SKILL.md b/.mintlify/skills/fish-audio-sdk/SKILL.md index 493a61d..5f342c8 100644 --- a/.mintlify/skills/fish-audio-sdk/SKILL.md +++ b/.mintlify/skills/fish-audio-sdk/SKILL.md @@ -18,7 +18,8 @@ If the user wants raw `curl` / HTTP / WebSocket without installing an SDK, use t - **Auth:** both SDKs read the API key from the `FISH_API_KEY` environment variable automatically. Get keys at `https://fish.audio/app/api-keys`. Never hardcode a key. - **Base URL:** `https://api.fish.audio` (override with `base_url=` in Python / `baseUrl:` in JS). -- **Models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument. +- **TTS models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument. +- **ASR models:** use `transcribe-1-pro` (recommended: speaker turns, long recordings, emotion cues). Neither SDK has an ASR `model` argument: send the `model: transcribe-1-pro` HTTP header on every request (Python `RequestOptions(additional_headers=...)`, JS `requestOptions.headers`). A request without it is served and billed as `transcribe-1`, the model for short recordings. See [references/speech-to-text.md](references/speech-to-text.md). - **Audio formats:** `mp3` (default), `wav`, `pcm`, `opus`. - **Playback in examples:** `play()` shells out to a system audio tool: Python uses **ffmpeg/ffplay** (or `mpv`), JS uses **ffplay**. It is for local/desktop use; in a server, `save()` to a file or stream the bytes instead. See [references/installation.md](references/installation.md). @@ -79,7 +80,7 @@ const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1"); | Install, auth, playback deps, verify a key | [references/installation.md](references/installation.md) | | Text-to-Speech (convert, stream, formats, prosody, model select) | [references/text-to-speech.md](references/text-to-speech.md) | | Voice cloning (instant references + persistent voice models) | [references/voice-cloning.md](references/voice-cloning.md) | -| Speech-to-Text (transcribe, segments, timestamps) | [references/speech-to-text.md](references/speech-to-text.md) | +| Speech-to-Text (models, timestamps, fields the SDK drops) | [references/speech-to-text.md](references/speech-to-text.md) | | Realtime WebSocket TTS (stream text → audio) | [references/websocket.md](references/websocket.md) | | Errors, retries, and timeouts (the **real** exception types) | [references/errors.md](references/errors.md) | @@ -101,6 +102,7 @@ The two SDKs do **not** use the same names. Use this map when porting code betwe | Credit balance | `client.account.get_credits()` | `client.user.get_api_credit()` | | Subscription package | `client.account.get_package()` | `client.user.get_package()` | | Choose model | `model="s2-pro"` keyword arg | positional `backend` arg, e.g. `convert(req, "s2-pro")` | +| ASR model (header) | `RequestOptions(additional_headers=...)` | `convert(req, { headers: { model: "transcribe-1-pro" } })` | ## Decision shortcuts @@ -109,11 +111,11 @@ The two SDKs do **not** use the same names. Use this map when porting code betwe - **Clone a voice instantly from a clip** → pass `references=[ReferenceAudio(audio=..., text=...)]` (Python) / `references: [{ audio, text }]` (JS). See [voice-cloning](references/voice-cloning.md). - **Persistent custom voice to reuse** → create a voice model, then use its `id` as `reference_id`. - **Stream tokens from an LLM and play speech as it arrives** → `tts.stream_websocket` (Python) / `textToSpeech.convertRealtime` (JS). See [websocket](references/websocket.md). -- **Transcribe audio** → `asr.transcribe` (Python) / `speechToText.convert` (JS). +- **Transcribe audio** → `asr.transcribe` (Python) / `speechToText.convert` (JS) with the `model: transcribe-1-pro` header (recommended; without it, the request runs on `transcribe-1`). Pro request fields, `speaker_turns`, `request_id`, and the language fields need raw HTTP in Python (`fish-audio-sdk` 1.3.0). See [speech-to-text](references/speech-to-text.md). ## Gotchas (verified against the SDK source) - Python `latency` accepts only **`"normal"` or `"balanced"`** (default `"balanced"`); there is no `"low"`. - The Python client has **no `max_retries`** and does **not** auto-retry; the JS client **does** auto-retry (configurable via per-call `requestOptions.maxRetries`). See [errors](references/errors.md). - Python defines a `ValidationError` class but **never raises it**, so don't catch it expecting validation failures; a 422 surfaces as `APIError`. The JS SDK throws `UnprocessableEntityError` on 422. -- ASR segment `start` / `end` are in **seconds**, but `duration` is in **milliseconds**. See [speech-to-text](references/speech-to-text.md). +- ASR segment `start` / `end` and `duration` are all in **seconds**. The Python SDK's `ASRResponse` docstring says milliseconds; that is wrong. See [speech-to-text](references/speech-to-text.md). diff --git a/.mintlify/skills/fish-audio-sdk/references/speech-to-text.md b/.mintlify/skills/fish-audio-sdk/references/speech-to-text.md index c62e401..18b98a5 100644 --- a/.mintlify/skills/fish-audio-sdk/references/speech-to-text.md +++ b/.mintlify/skills/fish-audio-sdk/references/speech-to-text.md @@ -1,14 +1,35 @@ # Speech-to-Text (ASR) +Both SDKs wrap `POST /v1/asr`: one audio file per request, and the response arrives when the whole file has been transcribed. Use `transcribe-1-pro`, the recommended model. Neither SDK has a `model` argument for ASR, so select it with the `model` HTTP header on every request. Full guide: `https://docs.fish.audio/features/speech-to-text`. + +## Choose a model + +- `transcribe-1-pro` (recommended): recordings up to 60 minutes, including multi-speaker conversations. `text` contains inline `<|speaker:N|>` markers and bracketed emotion or vocal-event cues such as `[laughter]` or `[高兴]`, and the API returns structured `speaker_turns` when you request timestamps. +- `transcribe-1`: general transcription of short recordings. It serves every request whose `model` header is missing or not an exact match. + +Select `transcribe-1-pro` with the `model` header: + +- Python: `request_options=RequestOptions(additional_headers={"model": "transcribe-1-pro"})`, with `from fishaudio.core import RequestOptions`. +- JavaScript: pass `{ headers: { model: "transcribe-1-pro" } }` as the second argument of `convert`. +- Write the value exactly, in lowercase. A missing or unrecognized value (for example `Transcribe-1-Pro` or `transcribe-1pro`) is served and billed as `transcribe-1`, and no error is returned. If you expected `transcribe-1-pro` but the transcript has no speaker markers, check the header. + ## Python: `client.asr.transcribe` ```python from fishaudio import FishAudio +from fishaudio.core import RequestOptions -client = FishAudio() +client = FishAudio() # reads FISH_API_KEY with open("audio.wav", "rb") as f: - result = client.asr.transcribe(audio=f.read(), language="en") + result = client.asr.transcribe( + audio=f.read(), + language="en", # optional hint; omit to auto-detect + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"}, # without this header: transcribe-1 + timeout=900, # long Pro recordings can take several minutes + ), + ) print(result.text) @@ -18,26 +39,88 @@ for seg in result.segments: Keyword params: -| Param | Type | Default | Notes | -| -------------------- | ------------------------ | ------------ | ----------------------------------------------------------------------------------------------------------------------- | -| `audio` | `bytes` | — (required) | Raw audio bytes. | -| `language` | `str` | auto-detect | Omit to auto-detect (e.g. `"en"`, `"zh"`, `"ja"`). | -| `include_timestamps` | `bool` | `True` | `False` omits per-segment timestamps (and `segments` is empty). Computing timestamps adds latency on clips under ~30 s. | -| `request_options` | `RequestOptions \| None` | `None` | Per-request timeout / headers. | +| Param | Type | Default | Notes | +| -------------------- | ------------------------ | ------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `audio` | `bytes` | — (required) | Raw file bytes; the format is detected from the content. The SDK sends them as MessagePack. | +| `language` | `str` | auto-detect | Optional hint, a lowercase ISO 639-1 code (`"en"`, `"zh"`, `"ja"`). Detection still runs, and the hint does not force the transcript language. `"en-US"` or `"English"` may be rejected with 400. | +| `include_timestamps` | `bool` | `True` | `True` (default) requests word-level timestamps, which adds processing time; pass `False` when you only need text (`segments` is then empty). | +| `request_options` | `RequestOptions \| None` | `None` | Per-request `timeout` and headers. Select the model here: `RequestOptions(additional_headers={"model": "transcribe-1-pro"})`. | ### Response shape (`ASRResponse`) ```python -result.text # str: full transcript -result.duration # float: total audio duration in MILLISECONDS -result.segments # list[ASRSegment] +result.text # str: full transcript (Pro: with <|speaker:N|> markers and [cues]) +result.duration # float: total audio duration in seconds, including silence +result.segments # list[ASRSegment]; [] when timestamps are off or no speech was found # each segment: -seg.text # str +seg.text # str: usually one word (one or a few characters in Chinese/Japanese) seg.start # float: seconds -seg.end # float: seconds +seg.end # float: seconds (can equal start) +``` + +`duration`, `start`, and `end` are all in **seconds**. The SDK's `ASRResponse` docstring says milliseconds; that is wrong, so do not divide by 1000. + +Segment text has no punctuation, speaker markers, or cues, and can be normalized (for example `35` for `3.5`), so it does not always match `text` character for character. Segments are not speaker turns. + +### Fields the Python SDK does not expose + +In `fish-audio-sdk` 1.3.0, `ASRResponse` keeps only `text`, `duration`, and `segments`. It silently drops these response fields: + +- `speaker_turns` (`transcribe-1-pro`, when timestamps are requested); +- `request_id` (`transcribe-1-pro`), also sent as the `x-request-id` header; +- `language` (English name, such as `English`) and `language_code` (ISO 639-1, such as `en`), returned when the language is known. + +`asr.transcribe()` also cannot send the `transcribe-1-pro` fields `tag_audio_events`, `diarize`, `num_speakers`, `min_speakers`, or `max_speakers`. For any of these, call the API directly. Field semantics are in the `fish-audio-api` skill (`POST /v1/asr`). + +```python +import os +import httpx + +with open("meeting.mp3", "rb") as f: + r = httpx.post( + "https://api.fish.audio/v1/asr", + headers={ + "Authorization": f"Bearer {os.environ['FISH_API_KEY']}", + "model": "transcribe-1-pro", + }, + files={"audio": f}, + data={"ignore_timestamps": "false"}, # word timestamps + speaker_turns + # Long Pro recordings can take several minutes; httpx defaults to 5 s. + timeout=httpx.Timeout(900.0, connect=10.0), + ) +r.raise_for_status() +result = r.json() +print(result.get("language_code"), result.get("request_id")) +for turn in result.get("speaker_turns", []): + print(f"{turn['speaker']} [{turn['start']:.2f}-{turn['end']:.2f}] {turn['text']}") +``` + +### Errors (Python) + +`asr.transcribe()` raises `APIError` (`RateLimitError` for 429, `ServerError` for 5xx) with `.status`, `.message`, and `.body` (the raw response text). On `transcribe-1-pro`, error bodies also carry `code` and `request_id`: + +```python +import json +from fishaudio.core import RequestOptions +from fishaudio.exceptions import APIError + +try: + result = client.asr.transcribe( + audio=audio_bytes, + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"} + ), + ) +except APIError as e: + try: + err = json.loads(e.body or "{}") + except ValueError: + err = {} # errors from the network edge may not be JSON + print(e.status, err.get("code"), err.get("request_id")) + raise ``` -> **Unit gotcha (verified in source):** segment `start` / `end` are in **seconds**, but `duration` is in **milliseconds**. Don't assume they share a unit. +Branch on `code` or the HTTP status, never on `message`. Retry 429 and 5xx with exponential backoff (the Python SDK does not retry); do not retry other 4xx responses. See the `fish-audio-api` skill for the status and `code` list. If you use `transcribe-1`, rely only on `status` and `message`. ## JavaScript: `client.speechToText.convert` @@ -48,11 +131,14 @@ import { readFile } from "node:fs/promises"; const client = new FishAudioClient(); const buf = await readFile("audio.wav"); -const result = await client.speechToText.convert({ - audio: new File([buf], "audio.wav"), - language: "en", // optional; omit to auto-detect - ignore_timestamps: false, // false → include per-segment timestamps -}); +const result = await client.speechToText.convert( + { + audio: new File([buf], "audio.wav"), + language: "en", // optional hint; omit to auto-detect + ignore_timestamps: false, // false → word-level segments (and speaker_turns on Pro) + }, + { headers: { model: "transcribe-1-pro" }, timeoutInSeconds: 900 } +); console.log(result.text); for (const seg of result.segments) { @@ -60,4 +146,43 @@ for (const seg of result.segments) { } ``` -`STTRequest` = `{ audio: File; language?: string; ignore_timestamps?: boolean }`. Note JS uses `ignore_timestamps` (the inverse of Python's `include_timestamps`). `STTResponse` mirrors the Python shape: `{ text, duration, segments }`. +`STTRequest` = `{ audio: File; language?: string; ignore_timestamps?: boolean }`, sent as multipart. JS uses `ignore_timestamps` (the inverse of Python's `include_timestamps`), and the server default is `true`, so you get no `segments` unless you pass `false`. + +In Node.js, the built-in `fetch` that the SDK uses stops waiting for response headers after 300 s, whatever `timeoutInSeconds` says (the call fails with `fetch failed`). For long Pro recordings, raise that limit once at startup with the `undici` package (`undici@7` runs on Node.js 20.18.1+; `undici@8` needs Node.js 22.19+ and fails at import on older versions): + +```ts +import { Agent, setGlobalDispatcher } from "undici"; // npm install undici@7 + +setGlobalDispatcher( + new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 }) +); +``` + +`STTResponse` types only `{ text, duration, segments }`; at runtime the body also has `language`, `language_code`, and on Pro `request_id` and `speaker_turns`. Widen the type to read them: + +```ts +type SpeakerTurn = { + speaker: string; + text: string; + start: number; + end: number; +}; +const body = result as typeof result & { + language?: string; + language_code?: string; + request_id?: string; + speaker_turns?: SpeakerTurn[]; +}; +for (const turn of body.speaker_turns ?? []) { + console.log(`${turn.speaker} [${turn.start}-${turn.end}] ${turn.text}`); +} +``` + +The JS SDK cannot send `tag_audio_events`, `diarize`, or the speaker counts; use `fetch` with `FormData` for those (see the `fish-audio-api` skill). + +## Limits and formats (both SDKs) + +- `transcribe-1-pro` accepts recordings up to 60 minutes (longer returns 400 `audio_too_long`). Send long recordings as compressed audio (MP3, Opus, or AAC); a request that is too large returns 413. Send a whole conversation as one file: speaker labels are consistent within one response, not across requests. +- Processing time grows with the length of the recording, and long `transcribe-1-pro` requests can take several minutes. Raise the SDK timeout for long recordings (the examples use 900 s; the defaults are in [errors](errors.md)), and in Node.js also raise the `fetch` header limit shown above. +- `transcribe-1-pro` accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file. AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM return 400. +- If you use `transcribe-1`: it is designed for short recordings, up to 50 MiB per request; keep MP3 and Opus files under 25 MiB. For recordings longer than a few minutes, use `transcribe-1-pro`. It accepts WAV, MP3, AAC (including M4A/MP4), FLAC, and Ogg (Opus or Vorbis); convert browser WebM recordings to Ogg/Opus, MP3, or WAV first, or use `transcribe-1-pro`. diff --git a/api-reference/endpoint/openapi-v1/speech-to-text.mdx b/api-reference/endpoint/openapi-v1/speech-to-text.mdx index 30f0751..1adc2c6 100644 --- a/api-reference/endpoint/openapi-v1/speech-to-text.mdx +++ b/api-reference/endpoint/openapi-v1/speech-to-text.mdx @@ -1,32 +1,59 @@ --- openapi: post /v1/asr title: "Speech to Text" -description: "Transcribe audio with transcribe-1 or transcribe-1-pro, with optional timestamps and emotion cues" +description: "Transcribe audio with transcribe-1-pro, the recommended speech-to-text model: request and response fields, speaker turns, limits, formats, and errors" icon: "microphone" iconType: "solid" --- - This BETA endpoint accepts audio only as `multipart/form-data` (a file upload) - or `application/msgpack`. JSON with base64-encoded audio is not supported. + This BETA endpoint transcribes one audio file per request. Send the audio as + `multipart/form-data` (a file upload) or `application/msgpack`. Base64-encoded + audio in a JSON body is not supported. ## Select a model -Send `POST https://api.fish.audio/v1/asr` with `Authorization: Bearer ` and an optional **`model` HTTP header**: +Send `POST https://api.fish.audio/v1/asr` with `Authorization: Bearer ` and a **`model` HTTP header** that selects the model. The request is synchronous: the response arrives when the whole file has been transcribed. -| Header value | Behavior | -| ------------------ | -------------------------------------------------------------------------------------------------- | -| `transcribe-1` | General transcription; the default if the header is omitted. | -| `transcribe-1-pro` | Transcription of multi-speaker conversations, preserving emotion and vocal-event cues in the text. | +| Header value | Use it for | +| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `transcribe-1-pro` | Recommended. Recordings up to 60 minutes, including multi-speaker conversations: speaker markers and speaker turns, with emotion and vocal-event cues preserved. | +| `transcribe-1` | General transcription of short recordings. Used when the `model` header is missing or not an exact match. | -The model belongs in the header for both multipart and MessagePack requests. The request body contains: +Send `model: transcribe-1-pro` with every request, written exactly, in lowercase. If the header is missing, or its value is not an exact match (for example `Transcribe-1-Pro` or `transcribe-1pro`), the request is served and billed as `transcribe-1`, and no error is returned. If you expected `transcribe-1-pro` but the transcript has no speaker markers, check the header. This applies to the API playground on this page too: set its `model` header to `transcribe-1-pro`. -| Field | Required | Default | Description | -| ------------------- | -------- | ------- | ----------------------------------------------------------------------------------------- | -| `audio` | Yes | — | Audio file upload for multipart, or raw audio bytes for MessagePack. | -| `language` | No | Unset | Optional language hint, such as `en`, `zh`, or `ja`. The language is still auto-detected. | -| `ignore_timestamps` | No | `true` | Set to `false` to request timestamped segments. Alignment adds processing time. | +The model belongs in the header for both multipart and MessagePack requests; a `model` field in the request body is ignored. See [Limits](#limits) for recording length and request size. + +To link the request to your own distributed trace, also send a W3C `traceparent` header. See [Tracing & Performance Analysis](/api-reference/observability). + +## Request fields + +Multipart and MessagePack bodies use the same field names. Multipart values are text; MessagePack values are typed. + +| Field | Required | Default | Description | +| ------------------------------ | -------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `audio` | Yes | — | Exactly one audio file per request: a file upload for multipart, or binary bytes for MessagePack. The format is detected from the file contents, not the file name. See [Supported audio formats](#supported-audio-formats). | +| `language` | No | Unset | Optional language hint, as a lowercase ISO 639-1 code such as `en`, `zh`, or `ja`. The language is detected automatically either way. An empty multipart value counts as unset. See [Language](#language). | +| `ignore_timestamps` | No | `true` | Default `true` (no timestamps). Set to `false` to get word-level `segments` and `speaker_turns`. Timestamps add processing time. | +| `tag_audio_events` | No | `true` | `true`: bracketed emotion and vocal-event cues such as `[laughter]` or `[高兴]` stay in `text` and in `speaker_turns`. `false`: the cues are removed. Timestamps, `duration`, and billing do not change. | +| `diarize` | No | `auto` | `auto` or `true`: return `speaker_turns` when timestamps are requested. `false`: omit `speaker_turns`. `false` does not change the transcript, and the speaker markers stay in `text`. | +| `num_speakers` | No | Unset | The expected number of speakers, an integer of 1 or more. Cannot be combined with `min_speakers` or `max_speakers`. | +| `min_speakers`, `max_speakers` | No | Unset | Lower and upper bounds on the number of speakers, integers of 1 or more. `min_speakers` must not exceed `max_speakers`. | + +In multipart forms, send `true` or `false` for `ignore_timestamps`. Any value other than `true` (case-insensitive), including `1` or an empty value, is read as `false` and turns timestamps on. In MessagePack, send a boolean; a string or `nil` returns 400. + +Speaker-count hints are applied on a best-effort basis: they guide speaker identification on longer recordings and may have no effect on short ones. They cannot be sent with `diarize=false`. + +`tag_audio_events`, `diarize`, and the speaker-count hints are validated strictly: + +- In multipart forms, `tag_audio_events` takes `true` or `false`, and `diarize` takes `auto`, `true`, or `false`, in any letter case. Speaker counts take digits only. Send each field at most once. +- In MessagePack, send `tag_audio_events` as a boolean, `diarize` as a boolean or one of the lowercase strings `auto`, `true`, or `false`, and speaker counts as integers, not floats or strings. `nil` means the default. +- Any other value, a repeated multipart field, or a conflicting combination returns 400 with code `invalid_parameter`. + +**If you use `transcribe-1`:** send `audio`, `language`, and `ignore_timestamps`. The other fields apply to `transcribe-1-pro`; send them only with `model: transcribe-1-pro`. + +### Examples ```bash curl --request POST https://api.fish.audio/v1/asr \ @@ -36,38 +63,203 @@ curl --request POST https://api.fish.audio/v1/asr \ --form ignore_timestamps=false ``` +This request returns word-level `segments` and `speaker_turns`. For a two-person interview, you can add a speaker-count hint, and remove emotion and vocal-event cues from the transcript: + +```bash +curl --request POST https://api.fish.audio/v1/asr \ + --header "Authorization: Bearer $FISH_API_KEY" \ + --header "model: transcribe-1-pro" \ + --form audio=@interview.mp3 \ + --form ignore_timestamps=false \ + --form num_speakers=2 \ + --form tag_audio_events=false +``` + Let your HTTP client set the multipart `Content-Type` and boundary. For MessagePack, set `Content-Type: application/msgpack` and encode `audio` as binary bytes. See the [Speech to Text guide](/features/speech-to-text#direct-api-messagepack) for an example. ## Read the response -| Field | Description | -| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `text` | Full transcript. Pro includes inline speaker markers such as `<\|speaker:0\|>` and can include emotion and vocal-event cues such as `[高兴]` (happy) or `[laughter]`. | -| `duration` | Audio duration in seconds. | -| `segments` | Array of `{ "text": string, "start": number, "end": number }`, with timestamps in seconds. Empty when timestamps are skipped; can also be empty if alignment is unavailable or no speech is detected. | -| `language_code` | Detected language code, such as `en` or `ja`, when available. Use this field in application logic. | -| `language` | Detected language name, such as `English`, when available. Intended for display. | +A successful response (HTTP 200) is a JSON object: + +| Field | Description | +| --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `text` | The full transcript, with inline speaker markers such as `<\|speaker:0\|>` and, unless `tag_audio_events=false`, bracketed emotion and vocal-event cues such as `[高兴]` (happy) or `[laughter]`. See [Multi-speaker response format](#multi-speaker-response-format). | +| `duration` | Length of the audio in seconds, including silence. | +| `segments` | Word-level timestamps: an array of `{ "text", "start", "end" }`, with times in seconds. Empty (`[]`) when `ignore_timestamps=true` (the default), when no speech was found, or when timing is temporarily unavailable. See [Segments](#segments). | +| `speaker_turns` | The transcript as timed speaker turns. Returned when `ignore_timestamps=false` and `diarize` is not `false`; otherwise absent. See [Speaker turns](#speaker-turns). | +| `language` | The detected language's English name, such as `English` or `Chinese`. Omitted when it cannot be determined. Intended for display. | +| `language_code` | Two-letter ISO 639-1 code for the language, such as `en`. Omitted when unknown. If the language cannot be determined and you sent a `language` hint, this reports your hint. See [Language](#language). | +| `request_id` | A unique ID for this request, also present in `transcribe-1-pro` error bodies and in the `x-request-id` response header. Platform and network-edge errors do not carry it. Include it when you contact support. | + +A field that is not returned is absent from the JSON, never `null`. Do not depend on the order of keys in the JSON. + +**If you use `transcribe-1`:** rely on `text`, `duration`, `segments`, `language`, and `language_code`. Speaker markers, `speaker_turns`, and `request_id` are `transcribe-1-pro` features. + + + Both SDKs select `transcribe-1-pro` through the `model` header. In Python, + pass + `request_options=RequestOptions(additional_headers={"model": "transcribe-1-pro"})` + to `asr.transcribe()`. In JavaScript, pass + `{ headers: { model: "transcribe-1-pro" } }` as the second argument to + `speechToText.convert()`. The Python SDK currently returns only `text`, + `duration`, and `segments`, and neither SDK can send `tag_audio_events`, + `diarize`, or the speaker-count hints. To use them, call the API directly, as + in the examples above. The JavaScript SDK returns the whole response body, + but its `STTResponse` type declares only `text`, `duration`, and `segments`. + ### Multi-speaker response format -Pro returns speaker turns **inside the `text` string**, using `<|speaker:N|>` markers. The numeric label `N` identifies the speaker for the text that follows, up to the next marker. A repeated label means that speaker is speaking again. Labels apply within the recording; they are not names or identities shared across requests. +`transcribe-1-pro` marks speaker changes **inside the `text` string** with `<|speaker:N|>` markers. The label `N` identifies the speaker for the text that follows, up to the next marker. A repeated label means that speaker is speaking again. -This illustrative response has two speakers and three turns. With `ignore_timestamps=true`, timestamps are skipped: +- Markers usually have a space on each side (`<|speaker:0|> 你好。 <|speaker:1|> ...`). Do not depend on exact spacing. +- Any text before the first marker belongs to the first turn. +- A transcript with no marker comes from a single speaker; treat it as speaker 0. +- Labels identify speakers within one response only. They are not names, and the same label in two requests is not the same person. Send a whole conversation as one file. + +This illustrative response has two speakers and three turns. With `ignore_timestamps=true` (the default), timestamps are skipped and `speaker_turns` is absent: ```json { - "text": "<|speaker:0|>你好。<|speaker:1|>[高兴]很开心认识你。<|speaker:0|>我也是。", + "text": "<|speaker:0|> 你好。 <|speaker:1|> [高兴]很开心认识你。 <|speaker:0|> 我也是。", "duration": 6.4, "segments": [], "language_code": "zh", - "language": "Chinese" + "language": "Chinese", + "request_id": "5f0c7e2a-9b1d-4c3e-8f6a-2d4b7c9e1a03" +} +``` + +Read this as speaker 0 saying `你好。`, speaker 1 saying `[高兴]很开心认识你。`, and speaker 0 saying `我也是。`. The `[高兴]` cue describes the delivery of speaker 1's speech. There is no separate emotion field; the cues are part of the text. + +### Speaker turns + +With `ignore_timestamps=false`, `transcribe-1-pro` also returns the turns as a structured `speaker_turns` array, unless you send `diarize=false`. When you request timestamps, prefer `speaker_turns` over parsing `text`. + +- `speaker_turns` lists the turns in the order they occur, and is `[]` when no speech was found. Each turn is `{ "speaker": "speaker:N", "text": "...", "start": , "end": }`. +- `speaker` is the string `speaker:N`, where `N` matches the `<|speaker:N|>` marker in `text`. Like the markers, labels identify speakers within one response only. +- `text` is that turn's speech without speaker markers. It keeps emotion and vocal-event cues unless `tag_audio_events=false`. +- `start` and `end` come from the word timestamps. If word timing is unavailable (`segments` is empty), turn times are approximate and can cover the whole recording. +- Consecutive turns can have the same speaker. Do not assume turns are contiguous or non-overlapping. + +This illustrative response is the same conversation with `ignore_timestamps=false`: + +```json +{ + "text": "<|speaker:0|> 你好。 <|speaker:1|> [高兴]很开心认识你。 <|speaker:0|> 我也是。", + "duration": 6.4, + "segments": [ + { "text": "你", "start": 0.24, "end": 0.52 }, + { "text": "好", "start": 0.52, "end": 0.8 }, + { "text": "很", "start": 1.6, "end": 1.84 }, + { "text": "开", "start": 1.84, "end": 2.1 }, + { "text": "心", "start": 2.1, "end": 2.36 }, + { "text": "认", "start": 2.36, "end": 2.7 }, + { "text": "识", "start": 2.7, "end": 2.98 }, + { "text": "你", "start": 2.98, "end": 3.3 }, + { "text": "我", "start": 4.5, "end": 4.78 }, + { "text": "也", "start": 4.78, "end": 5.02 }, + { "text": "是", "start": 5.02, "end": 5.4 } + ], + "speaker_turns": [ + { "speaker": "speaker:0", "text": "你好。", "start": 0.24, "end": 0.8 }, + { + "speaker": "speaker:1", + "text": "[高兴]很开心认识你。", + "start": 1.6, + "end": 3.3 + }, + { "speaker": "speaker:0", "text": "我也是。", "start": 4.5, "end": 5.4 } + ], + "language_code": "zh", + "language": "Chinese", + "request_id": "5f0c7e2a-9b1d-4c3e-8f6a-2d4b7c9e1a03" +} +``` + +See the [Speech to Text guide](/features/speech-to-text#multi-speaker-conversations) for parsing examples. + +### Segments + +A segment is usually one word; in Chinese and Japanese it is usually one character, or a few. Segment text has no punctuation and no speaker markers or cues, and can be normalized (for example `35` for `3.5`), so it does not always match `text` character for character. A segment can have `start` equal to `end`. Segments are not speaker turns and carry no speaker ID. + +### Language + +- `language` is optional. The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language. +- Use a lowercase ISO 639-1 code such as `en`, `zh`, or `ja`. Other forms, such as `en-US` or `English`, may be rejected with 400. +- If the language cannot be determined (for example, very short audio), `language_code` reports your hint, which is not checked against the audio, and `language` may be absent. +- Responses report one language. For recordings that switch languages, `language` and `language_code` do not list every language spoken. + +## Limits + +- **Audio length:** up to 60 minutes per request. Longer audio returns 400 with code `audio_too_long`. +- **Request size:** for long recordings, send compressed audio such as MP3, Opus, or AAC. An hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413. +- **One conversation per request:** send the whole recording in one request. Timestamps and speaker labels then cover the whole recording. Do not split a conversation, because labels are consistent only within one response. + +**If you use `transcribe-1`:** it is designed for short recordings; for recordings longer than a few minutes, use `transcribe-1-pro`. It accepts up to 50 MiB per request; keep MP3 and Opus files under 25 MiB. A request that exceeds the size limit returns 413 or 400. A long request can fail with 503 if it exceeds the processing-time limit. Retry, and if it keeps failing, use `transcribe-1-pro` or split the audio. + +### Processing time and timeouts + +- Processing time grows with the length of the recording. Long `transcribe-1-pro` recordings can take several minutes. +- Set your HTTP client's timeout accordingly. The official SDKs' default timeout can be too short for long recordings, and many HTTP libraries default to even less (httpx defaults to 5 seconds). For long recordings, use a generous timeout, for example 15 minutes: `httpx.Timeout(900.0, connect=10.0)` with httpx, `RequestOptions(timeout=900)` with the Python SDK, or `{ timeoutInSeconds: 900 }` with the JavaScript SDK. +- In Node.js, the built-in `fetch`, which the JavaScript SDK uses, also stops waiting for a response after 5 minutes, even if you set a longer timeout. To wait longer, call `setGlobalDispatcher(new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 }))` from the `undici` package once at startup. See [Processing time and timeouts](/features/speech-to-text#processing-time-and-timeouts). +- If a long request fails with a 5xx error or the connection drops, retry it. +- Each request holds one of your account's [concurrent request slots](/developer-guide/models-pricing/pricing-and-rate-limits#concurrent-request-limits) until its response is returned. + +## Supported audio formats + +- `transcribe-1-pro` accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings. For a video file, it uses the first audio track. Send the original file bytes; no conversion is needed. +- Not supported: AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM. These return 400. + +**If you use `transcribe-1`:** it accepts WAV, MP3, AAC (including M4A/MP4), FLAC, and Ogg (Opus or Vorbis). Browser WebM recordings (for example, from `MediaRecorder`) may not be accepted; convert them to Ogg/Opus, MP3, or WAV first. + +## Errors + +Error responses have one of these shapes: + +- **`transcribe-1-pro` errors** are JSON with `status`, `message`, `code`, and `request_id`. Branch on `code`, not on `message`. +- **Platform errors**, such as authentication, credit, and concurrency errors, are returned before the request reaches a model. They are JSON with exactly `status` and `message`. +- **Errors from the network edge** in front of the API, such as some 413 and 5xx responses, may not have a JSON body. Handle them by HTTP status. + +This illustrative `transcribe-1-pro` error body shows the shape: + +```json +{ + "status": 400, + "message": "A human-readable description of the problem.", + "code": "invalid_parameter", + "request_id": "5f0c7e2a-9b1d-4c3e-8f6a-2d4b7c9e1a03" } ``` -Read this as speaker 0 saying `你好。`, speaker 1 saying `[高兴]很开心认识你。`, and speaker 0 saying `我也是。`. The `[高兴]` cue describes the delivery of speaker 1's speech. +| Status | `code` | Meaning | Retry? | +| ------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------- | ----------------- | +| 400 | `invalid_request` | Malformed request: the body could not be parsed, `audio` is missing, or more than one `audio` was sent. | No | +| 400 | `invalid_parameter` | A parameter has a value it does not accept, or parameters conflict. | No | +| 400 | `invalid_audio` | The audio could not be decoded, or the format is not supported. | No | +| 400 | `audio_too_long` | The audio is longer than 60 minutes. | No | +| 400 | `audio_too_short` | The audio is too short to transcribe (under about 0.08 seconds). | No | +| 400 | — | The request body could not be read. | No | +| 401 | — | Missing or invalid API key. | No | +| 402 | — | Insufficient API credit. | No | +| 413 | `request_too_large` | The request is too large. It may come from the network edge without a JSON body. | No | +| 415 | `unsupported_media_type` | Unsupported `Content-Type`. | No | +| 429 | — (rarely `upstream_rejected`) | Usually your account is at its concurrency limit. No `Retry-After` header is sent. | Yes, with backoff | +| Other 4xx (not 429) | `upstream_rejected` | Rare: the request was rejected during transcription. | No | +| 500 | `internal_error` | Internal error. | Yes, with backoff | +| 502 | — | The service could not be reached. | Yes, with backoff | +| 503 | `upstream_unavailable` | The transcription service is temporarily unavailable. | Yes, with backoff | +| 503 | `upstream_timeout` | Processing took too long. | Yes, with backoff | +| 503 | `excessive_repetition` | The transcription produced unusable repeated output. | Yes, with backoff | +| 503 | `diarization_failed` | Speaker identification failed for this recording. | Yes, with backoff | +| 504 | — | Timed out connecting to the service. | Yes, with backoff | -With `ignore_timestamps=false` (as in the curl example above), the same markers remain in `text`. The `segments` array contains timed speech with `text`, `start`, and `end`; alignment excludes speaker, emotion, and event markers. Segments are not speaker turns and do not contain a `speaker_id` field. +- A dash means the error has no `code`; a 503 without a `code` also means the service is temporarily unavailable. +- New `code` values may be added; handle unknown codes by HTTP status. +- Retry 429 and 5xx responses with exponential backoff. Do not retry other 4xx responses; change the request first. +- Requests that return an error response are not billed. +- With the Python SDK, the raw error body is in the exception's `body` attribute; parse it with `json.loads` to read `code` and `request_id`. -The API returns no separate speaker list or emotion field. Parse the inline markers in `text` when your application needs speaker turns or emotion cues. See the [speaker parsing example](/features/speech-to-text#multi-speaker-conversations). +**If you use `transcribe-1`:** its errors are JSON with `status` and `message`. Rely only on these two fields and on the HTTP status. -See the [model capabilities and SDK examples](/features/speech-to-text) for more detail. +See [Errors](/api-reference/errors) for retry and SDK exception examples. diff --git a/api-reference/errors.mdx b/api-reference/errors.mdx index 861905c..b311945 100644 --- a/api-reference/errors.mdx +++ b/api-reference/errors.mdx @@ -4,7 +4,7 @@ description: "HTTP status codes, the error response shape, and how to handle the icon: "triangle-exclamation" --- -Every Fish Audio error comes back as JSON with a `message` and a `status`: +Fish Audio errors come back as JSON with a `message` and a `status`: ```json { "message": "Invalid Token", "status": 401 } @@ -12,6 +12,8 @@ Every Fish Audio error comes back as JSON with a `message` and a `status`: (A request whose body can't be parsed returns a plain-text parse error instead.) +Some endpoints add fields. For example, `transcribe-1-pro` speech-to-text errors also include `code` and `request_id`; see [Speech to Text errors](/api-reference/endpoint/openapi-v1/speech-to-text#errors). Errors returned by the network edge in front of the API, such as some `413` or `5xx` responses, may not be JSON. + ## Status codes | Status | Meaning | What to do | @@ -21,7 +23,7 @@ Every Fish Audio error comes back as JSON with a `message` and a `status`: | `402` | Insufficient credits | Top up on [API Billing](https://fish.audio/app/developers/billing/). | | `403` | Not permitted for this key/resource | Check the key's scope and the resource owner. | | `404` | Model or voice not found | Verify the `model_id` / `reference_id`. | -| `429` | Rate limit exceeded | Back off and retry (see below). | +| `429` | Too many requests, for example when your account is at its [concurrency limit](/developer-guide/models-pricing/pricing-and-rate-limits#concurrent-request-limits) | Back off and retry (see below). | | `5xx` | Server error | Retry with backoff; if it persists, contact support. | ## Retries diff --git a/api-reference/introduction.mdx b/api-reference/introduction.mdx index d2d41cb..a24b7b9 100644 --- a/api-reference/introduction.mdx +++ b/api-reference/introduction.mdx @@ -15,7 +15,7 @@ See our [Quick Start](/developer-guide/getting-started/quickstart) guide to gene ## Errors -Every error returns a JSON body with a `message` and a `status`. See [Errors](/api-reference/errors) for the full status-code table, retry guidance, and SDK exception handling. +Errors return a JSON body with a `message` and a `status`; some endpoints add fields. Errors from the network edge in front of the API, such as some `413` or `5xx` responses, may not have a JSON body. See [Errors](/api-reference/errors) for the full status-code table, retry guidance, and SDK exception handling. ## OpenAPI Schema diff --git a/api-reference/observability.mdx b/api-reference/observability.mdx index a820542..ea1eafd 100644 --- a/api-reference/observability.mdx +++ b/api-reference/observability.mdx @@ -174,6 +174,15 @@ const ws = new WebSocket("wss://api.fish.audio/v1/tts/live", { where your frontend can send the header through `fetch`. +## Speech-to-Text Request IDs + +`transcribe-1-pro` responses from `POST /v1/asr` include a `request_id` field +in the JSON body, on success and on its own error responses, and the same value +in the `x-request-id` response header. Platform errors, such as authentication, +credit, and concurrency errors, and errors from the network edge do not carry +it. Include it when you contact support about a transcription request. See +[Speech to Text](/api-reference/endpoint/openapi-v1/speech-to-text#read-the-response). + ## Enterprise Performance Analysis diff --git a/api-reference/sdk/javascript/api-reference.mdx b/api-reference/sdk/javascript/api-reference.mdx index 7b6a2e9..2d29394 100644 --- a/api-reference/sdk/javascript/api-reference.mdx +++ b/api-reference/sdk/javascript/api-reference.mdx @@ -43,13 +43,18 @@ Returns: `RealtimeConnection` (`EventEmitter`-like connection) emitting `Realtim Transcribe audio to text. ```typescript -const res = await fishAudio.speechToText.convert({ audio: myAudio }); +const res = await fishAudio.speechToText.convert( + { audio: myAudio }, + { headers: { model: "transcribe-1-pro" } } +); console.log(res.text); ``` -Parameters: `request` (STTRequest)
+Parameters: `request` (STTRequest), `requestOptions?` (for example `headers` and `timeoutInSeconds`)
Returns: `STTResponse` +Select the model with the `model` header in `requestOptions.headers`. Without it, the request is served and billed as `transcribe-1`. See [Speech to Text](/features/speech-to-text). + ## Voices ### search() diff --git a/api-reference/sdk/python/overview.mdx b/api-reference/sdk/python/overview.mdx index 69030c0..62f3d88 100644 --- a/api-reference/sdk/python/overview.mdx +++ b/api-reference/sdk/python/overview.mdx @@ -143,9 +143,18 @@ audio = client.tts.stream(text="Hello!").collect() ### Speech-to-Text ```python -# Transcribe audio +from fishaudio.core import RequestOptions + +# Transcribe audio with transcribe-1-pro (recommended). +# Without the model header, the request runs on transcribe-1. with open("audio.wav", "rb") as f: - result = client.asr.transcribe(audio=f.read(), language="en") + result = client.asr.transcribe( + audio=f.read(), + language="en", + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"} + ), + ) print(result.text) diff --git a/archive/python-sdk-legacy/migration-guide.mdx b/archive/python-sdk-legacy/migration-guide.mdx index 647c8d5..6812ac5 100644 --- a/archive/python-sdk-legacy/migration-guide.mdx +++ b/archive/python-sdk-legacy/migration-guide.mdx @@ -168,7 +168,8 @@ session = Session("your_api_key") with open("audio.mp3", "rb") as f: response = session.asr(ASRRequest( audio=f.read(), - language="en" + language="en", + ignore_timestamps=False )) print(response.text) @@ -191,16 +192,12 @@ with open("audio.mp3", "rb") as f: print(result.text) -# Timestamps in MILLISECONDS +# Timestamps in SECONDS (unchanged) for segment in result.segments: - print(f"[{segment.start}ms - {segment.end}ms]") + print(f"[{segment.start}s - {segment.end}s]") ``` - -ASR timestamps changed from seconds to milliseconds. Divide by 1000 to convert: `seconds = segment.start / 1000` - - ## WebSocket Streaming Migration @@ -322,14 +319,6 @@ audio = client.tts.convert(text="...", format="mp3") Pass parameters directly to methods. - -**Before:** `segment.start` in seconds (e.g., 1.5) - -**After:** `segment.start` in milliseconds (e.g., 1500) - -Convert: `seconds = segment.start / 1000` - - - `session.create_model()` → `client.voices.create()` - `session.list_models()` → `client.voices.list()` @@ -369,13 +358,6 @@ ws_session.tts(TTSRequest(text=""), text_stream()) client.tts.stream_websocket(text_stream()) ``` - - -New SDK uses milliseconds instead of seconds: -```python -seconds = segment.start / 1000 -``` - ## Next Steps diff --git a/archive/python-sdk-legacy/speech-to-text.mdx b/archive/python-sdk-legacy/speech-to-text.mdx index eac060e..5e31f70 100644 --- a/archive/python-sdk-legacy/speech-to-text.mdx +++ b/archive/python-sdk-legacy/speech-to-text.mdx @@ -36,35 +36,38 @@ with open("audio.mp3", "rb") as f: # Transcribe response = session.asr(ASRRequest( - audio=audio_data + audio=audio_data, + ignore_timestamps=True )) print(response.text) -print(f"Duration: {response.duration}ms") +print(f"Duration: {response.duration}s") ``` ## Language Specification -Improve accuracy by specifying the language: +Pass a language hint when you know the source language: ```python # English transcription response = session.asr(ASRRequest( audio=audio_data, - language="en" + language="en", + ignore_timestamps=True )) # Chinese transcription response = session.asr(ASRRequest( audio=audio_data, - language="zh" + language="zh", + ignore_timestamps=True )) ``` Common language codes: `en` (English), `zh` (Chinese), `es` (Spanish), `fr` (French), `de` (German), `ja` (Japanese), `ko` (Korean), `pt` (Portuguese) -Automatic language detection works well, but specifying the language improves accuracy and speed. +The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language. ## Working with Segments @@ -73,7 +76,8 @@ Get detailed timing for each segment: ```python response = session.asr(ASRRequest( - audio=audio_data + audio=audio_data, + ignore_timestamps=False )) # Full transcription @@ -89,7 +93,7 @@ for segment in response.segments: Control timestamp generation: ```python -# Include timestamps (default) +# Include timestamps response = session.asr(ASRRequest( audio=audio_data, ignore_timestamps=False # False = include timestamps @@ -103,7 +107,7 @@ response = session.asr(ASRRequest( ``` -`ignore_timestamps=False` (default) includes segment timestamps. Set to `True` to skip timestamp processing for faster transcription when you only need the text. +Always set `ignore_timestamps`. The legacy SDK sends `None` when you leave it out, and the API rejects that with a 400 error. `ignore_timestamps=False` includes segment timestamps. Set it to `True` to skip timestamp processing for faster transcription when you only need the text. ## Audio Formats @@ -117,8 +121,7 @@ Supported audio formats: - AAC File requirements: -- Maximum size: 100MB -- Maximum duration: 60 minutes +- Size and duration limits depend on the model; see [Speech to Text limits](/features/speech-to-text#limits) - Sample rate: 16kHz or higher recommended ## Transcribing TTS Output @@ -137,7 +140,8 @@ for chunk in session.tts(TTSRequest( # Transcribe it response = session.asr(ASRRequest( - audio=bytes(audio_buffer) + audio=bytes(audio_buffer), + ignore_timestamps=True )) print(response.text) @@ -152,12 +156,13 @@ from fish_audio_sdk.exceptions import HttpCodeErr try: response = session.asr(ASRRequest( - audio=audio_data + audio=audio_data, + ignore_timestamps=True )) except HttpCodeErr as e: - if e.status_code == 413: - print("Audio file too large (max 100MB)") - elif e.status_code == 400: + if e.status == 413: + print("Audio file too large") + elif e.status == 400: print("Invalid audio format") else: raise e @@ -170,7 +175,7 @@ The ASR response includes: | Field | Type | Description | |------------|---------------------|--------------------------------| | `text` | str | Complete transcription | -| `duration` | float | Audio duration (milliseconds) | +| `duration` | float | Audio duration (seconds) | | `segments` | list[ASRSegment] | Timestamped text segments | Segment structure: @@ -180,14 +185,10 @@ Segment structure: | `start` | float | Start time (seconds) | | `end` | float | End time (seconds) | - -Note the timing units: `duration` is in milliseconds while segment `start`/`end` are in seconds. - - ## Request Parameters | Parameter | Type | Description | Default | |---------------------|-------|----------------------------|--------------------| | `audio` | bytes | Audio data to transcribe | Required | | `language` | str | Language code (e.g., "en") | None (auto-detect) | -| `ignore_timestamps` | bool | Skip timestamp processing | False | \ No newline at end of file +| `ignore_timestamps` | bool | Skip timestamp processing | None (must be set) | \ No newline at end of file diff --git a/developer-guide/compat/capabilities.mdx b/developer-guide/compat/capabilities.mdx index 9e63fdc..95a5e3a 100644 --- a/developer-guide/compat/capabilities.mdx +++ b/developer-guide/compat/capabilities.mdx @@ -97,6 +97,11 @@ in-band error event): | OpenAI chat | `n ≠ 1`; audio input combined with audio output in one call | | OpenAI Realtime | `audio/pcmu`, `audio/pcma`; any `turn_detection` with a type (`server_vad`, `semantic_vad`, anything else; only `null` and `"none"` pass); `output_modalities` that omits `"audio"`; a second concurrent response (`response_already_active`) | +ElevenLabs STT refuses `diarize` and `tag_audio_events`. For speaker turns and +event tags, use the native +[Speech to Text API](/features/speech-to-text#multi-speaker-conversations) with +`transcribe-1-pro`. + On OpenAI transcription, `languages`, `keywords`, `prompt`, and `temperature` are in the [ignored tier](#how-parameters-are-handled). diff --git a/developer-guide/compat/migrate-from-elevenlabs.mdx b/developer-guide/compat/migrate-from-elevenlabs.mdx index ba1c567..bef7d1b 100644 --- a/developer-guide/compat/migrate-from-elevenlabs.mdx +++ b/developer-guide/compat/migrate-from-elevenlabs.mdx @@ -256,19 +256,23 @@ rather than silently ignored: | Option | Why it's refused | | ---------------------------------- | ------------------------------------------- | -| `diarize` | speaker attribution is not available | +| `diarize` | speaker attribution is not available here | | `use_multi_channel` | per-channel transcription is not available | | `entity_detection` | entities would be missing from the response | | `entity_redaction` | text would come back unredacted | | `entity_redaction_mode` | same | | `no_verbatim` | the transcript would still be verbatim | -| `tag_audio_events` | the transcript would carry no event tags | +| `tag_audio_events` | event tags are not available here | | `additional_formats` | the extra formats would be absent | | `timestamps_granularity=character` | only word granularity exists | | `webhook`, `webhook_id` | there is no delivery callback | | `source_url`, `cloud_storage_url` | remote audio would not be fetched | | `enable_logging=false` | zero-retention mode cannot be honored | +For speaker turns and event tags, use the native +[Speech to Text API](/features/speech-to-text#multi-speaker-conversations) with +`transcribe-1-pro` instead. + **`tag_audio_events` is the one that bites on migration.** ElevenLabs defaults it to `true` and their examples pass it explicitly, so code copied from an diff --git a/developer-guide/compat/migrate-from-openai.mdx b/developer-guide/compat/migrate-from-openai.mdx index 8b6cdad..33c59b1 100644 --- a/developer-guide/compat/migrate-from-openai.mdx +++ b/developer-guide/compat/migrate-from-openai.mdx @@ -188,6 +188,10 @@ Transcription model names carry over the same way: `whisper-1`, `gpt-4o-transcribe`, and `gpt-4o-mini-transcribe` are accepted as aliases for `fish-audio/transcribe-1`; an unrecognized name is refused with the same 400. +For speaker turns, long recordings, and inline emotion cues, use +`transcribe-1-pro` on the native +[Speech to Text API](/features/speech-to-text) instead. + In `verbose_json`, the Whisper-engine internals (`avg_logprob`, `no_speech_prob`, `compression_ratio`, `temperature`) are neutral constants; diff --git a/developer-guide/compat/migrate-from-openrouter.mdx b/developer-guide/compat/migrate-from-openrouter.mdx index b51825f..9caaf72 100644 --- a/developer-guide/compat/migrate-from-openrouter.mdx +++ b/developer-guide/compat/migrate-from-openrouter.mdx @@ -187,6 +187,10 @@ JSON bodies with base64 `input_audio` also work, matching OpenRouter's schema. For word-level timing, ask for `response_format="verbose_json"` together with `timestamp_granularities=["word"]`. +For speaker turns, long recordings, and inline emotion cues, use +`transcribe-1-pro` on the native +[Speech to Text API](/features/speech-to-text) instead. + ## Model catalog diff --git a/developer-guide/getting-started/migration.mdx b/developer-guide/getting-started/migration.mdx index 1f5cdc7..ccae6b1 100644 --- a/developer-guide/getting-started/migration.mdx +++ b/developer-guide/getting-started/migration.mdx @@ -369,6 +369,10 @@ response echoes what you sent: empty when you sent none, never a detection result (the [ElevenLabs protocol](/developer-guide/compat/migrate-from-elevenlabs#speech-to-text) echoes `language_code` the same way). +For speaker turns, long recordings, and inline emotion cues, use +`transcribe-1-pro` on the native +[Speech to Text API](/features/speech-to-text). + ## Voices The examples above use the model's default voice. To pick a specific one, pass diff --git a/developer-guide/models-pricing/models-overview.mdx b/developer-guide/models-pricing/models-overview.mdx index 941ff6d..9a7cf25 100644 --- a/developer-guide/models-pricing/models-overview.mdx +++ b/developer-guide/models-pricing/models-overview.mdx @@ -67,14 +67,20 @@ We recommend using `s2.1-pro` for production projects. Use `s2.1-pro-free` when ## Speech-to-Text Models -Both ASR models use `POST /v1/asr`. Select one with the `model` HTTP header. +We recommend `transcribe-1-pro` for speech to text. Both ASR models use `POST /v1/asr`, and you select one with the `model` HTTP header, written exactly as shown. Send `model: transcribe-1-pro` with every request: if the header is missing or its value is not an exact match, the request is served and billed as `transcribe-1`, and no error is returned. -| Model | Capabilities | -| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `transcribe-1` | General transcription with optional timestamps. Default when the `model` header is omitted. | -| `transcribe-1-pro` | Multi-speaker conversation transcription, with inline emotion and vocal-event cues such as `[高兴]` (happy) and `[laughter]`. Supports optional timestamps. | +| Model | Capabilities | +| -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `transcribe-1-pro` (recommended) | Anything from short clips to long recordings (up to 60 minutes), including multi-speaker conversations. Speaker markers in `text`, structured `speaker_turns` with start and end times, speaker-count hints, and inline emotion and vocal-event cues such as `[高兴]` (happy) and `[laughter]` (optional via `tag_audio_events`). Supports optional timestamps. | +| `transcribe-1` | General transcription of short recordings, with optional timestamps. Serves requests whose `model` header is missing or not an exact match. | -Pro returns speaker turns as inline `<|speaker:N|>` markers in `text`, alongside emotion cues. Repeated labels identify the same speaker within a recording. Timestamped `segments` contain spoken words without these markers and do not expose per-segment speaker IDs. See the [multi-speaker response example](/features/speech-to-text#multi-speaker-conversations) and [ASR pricing](/developer-guide/models-pricing/pricing-and-rate-limits#automatic-speech-recognition-asr-models). +Pro labels speakers with inline `<|speaker:N|>` markers in `text`, alongside emotion cues. A label identifies the same speaker within one response only. When you request timestamps (`ignore_timestamps=false`), Pro also returns `speaker_turns`, a structured list of turns with each turn's speaker, text, and start and end times, unless you send `diarize=false`. Timestamped `segments` are word-level, contain spoken words without these markers, and do not expose per-segment speaker IDs. See the [multi-speaker response example](/features/speech-to-text#multi-speaker-conversations) and [Speaker turns](/features/speech-to-text#speaker-turns). + +Pro has its own optional request fields: `tag_audio_events`, `diarize`, and the speaker-count hints `num_speakers`, `min_speakers`, and `max_speakers`. Speaker-count hints are applied on a best-effort basis and may have no effect on short recordings. For supported audio formats, file size, and length, see [Input audio](/features/speech-to-text#input-audio) and [Limits](/features/speech-to-text#limits). For costs, see [ASR pricing](/developer-guide/models-pricing/pricing-and-rate-limits#automatic-speech-recognition-asr-models). + + +**If you use `transcribe-1`:** it is designed for short recordings; for recordings longer than a few minutes, use `transcribe-1-pro`. Speaker markers, speaker turns, and emotion and vocal-event cues are `transcribe-1-pro` features, so send the Pro request fields above only with `model: transcribe-1-pro`. + ## Model Specifications diff --git a/developer-guide/models-pricing/pricing-and-rate-limits.mdx b/developer-guide/models-pricing/pricing-and-rate-limits.mdx index cdc1b9d..00a03f4 100644 --- a/developer-guide/models-pricing/pricing-and-rate-limits.mdx +++ b/developer-guide/models-pricing/pricing-and-rate-limits.mdx @@ -35,17 +35,19 @@ TTS pricing is based on the size of input text, measured in millions of UTF-8 by ### Automatic Speech Recognition (ASR) Models -| Model Name | Price (USD) | -| ------------------ | ------------------ | -| `transcribe-1` | $0.36 / audio hour | -| `transcribe-1-pro` | $0.36 / audio hour | +| Model Name | Price (USD) | +| -------------------------------- | ------------------ | +| `transcribe-1-pro` (recommended) | $0.36 / audio hour | +| `transcribe-1` | $0.36 / audio hour | -Both models use the [Speech to Text API](/api-reference/endpoint/openapi-v1/speech-to-text). Select `transcribe-1-pro` with the `model` HTTP header for multi-speaker conversation transcription and inline emotion cues. See the [model comparison](/features/speech-to-text#choose-an-asr-model). +We recommend `transcribe-1-pro`, which handles multi-speaker conversations (speaker markers and speaker turns), long recordings, and inline emotion cues. Both models cost the same and use the [Speech to Text API](/api-reference/endpoint/openapi-v1/speech-to-text). Send the `model: transcribe-1-pro` HTTP header to select Pro; a request whose `model` header is missing or not an exact match is served and billed as `transcribe-1`. See the [model comparison](/features/speech-to-text#choose-an-asr-model). **How ASR billing works:** - Charges are based on the duration of audio processed - Duration is rounded up to the nearest second +- The full duration of the recording is billed, including silence, once per successful request +- Requests that return an error response are not billed ### Voice Design @@ -79,6 +81,15 @@ These limits help us ensure fair usage and maintain service quality for all user tier immediately. +Speech-to-text requests count against the same concurrency limit as your other +API requests, shared by all API keys on your account. Each `/v1/asr` request +holds one slot until its response is returned, which for long +`transcribe-1-pro` recordings can be several minutes. + +When you are at the limit, the [native API](/api-reference/introduction), +including `POST /v1/asr`, returns 429 without a `Retry-After` header. Retry +with exponential backoff. + ### Convert Concurrency to QPS or QPM Fish Audio rate limits are based on **concurrent requests**: the number of diff --git a/developer-guide/resources/agent-quickstart.mdx b/developer-guide/resources/agent-quickstart.mdx index 3936242..e5d071c 100644 --- a/developer-guide/resources/agent-quickstart.mdx +++ b/developer-guide/resources/agent-quickstart.mdx @@ -135,6 +135,7 @@ Not a coding agent installing a skill, but an autonomous agent, RAG pipeline, or - Base API URL: `https://api.fish.audio` - Authentication: `Authorization: Bearer ` - TTS model selection: optional `model` header (`s1`, `s2-pro`, `s2.1-pro`, `s2.1-pro-free`). Recommended: `s2.1-pro` for production, `s2.1-pro-free` for the free developer tier. If omitted, the server uses `s2.1-pro` + - ASR model selection: `model` header on `POST /v1/asr` (`transcribe-1-pro`, `transcribe-1`). Recommended: `transcribe-1-pro` (speaker turns, long recordings, emotion cues); send it explicitly. If the header is missing or not an exact match, the request is served and billed as `transcribe-1`. The `/v1/asr` entry in `openapi.json` does not yet list the `transcribe-1-pro` fields, `request_id`, `speaker_turns`, or the error `code`; for those, follow the [Speech to Text API reference](https://docs.fish.audio/api-reference/endpoint/openapi-v1/speech-to-text.md) - Main REST endpoints: - `POST /v1/tts` - `POST /v1/asr` @@ -194,15 +195,15 @@ Not a coding agent installing a skill, but an autonomous agent, RAG pipeline, or - **Generate speech** → Quick Start, the Text to Speech guide, and `POST - /v1/tts`. - **Transcribe audio** → the Speech to Text guide and `POST - /v1/asr`. - **Clone or manage voices** → Creating Voice Models and the - `/model` endpoints. - **Stream audio in real time** → AsyncAPI, WebSocket TTS - Streaming, and the realtime guides. - **Pick a model or estimate cost** → - Models Overview and Pricing & Rate Limits. + /v1/tts`. - **Transcribe audio** → the Speech to Text guide and `POST /v1/asr` + with `model: transcribe-1-pro`. - **Clone or manage voices** → Creating Voice + Models and the `/model` endpoints. - **Stream audio in real time** → AsyncAPI, + WebSocket TTS Streaming, and the realtime guides. - **Pick a model or estimate + cost** → Models Overview and Pricing & Rate Limits. - - Prefer `openapi.json` and `asyncapi.yml` for machine-readable schemas. + - Prefer `openapi.json` and `asyncapi.yml` for machine-readable schemas, except for the `/v1/asr` gaps listed under Canonical API facts. - Append `.md` to any page URL to fetch the human-authored page as plain Markdown. - Some richer pages use interactive MDX widgets. If a fetched page contains UI or component noise, fall back to `llms.txt`, `llms-full.txt`, or the API spec files. diff --git a/developer-guide/sdk-guide/cookbook/batch-transcribe-with-language-hint.mdx b/developer-guide/sdk-guide/cookbook/batch-transcribe-with-language-hint.mdx index 5e6ac38..140e1d3 100644 --- a/developer-guide/sdk-guide/cookbook/batch-transcribe-with-language-hint.mdx +++ b/developer-guide/sdk-guide/cookbook/batch-transcribe-with-language-hint.mdx @@ -12,58 +12,111 @@ import Prerequisites from "/snippets/prerequisites.mdx"; ## Recipe -Read each file's bytes from disk and pass them to [`asr.transcribe()`](/api-reference/sdk/python/resources#transcribe) with an explicit `language`. A language hint is more reliable than auto-detection when you already know the source language, especially for phonetically similar languages. Collect one result row per file as you go. +This recipe uses `transcribe-1-pro`, which handles recordings up to 60 minutes long. Every tab selects it with the `model` request header; without that header, the request is served and billed as `transcribe-1`. + +Read each file's bytes from disk and pass them to [`asr.transcribe()`](/api-reference/sdk/python/resources#transcribe) with `language` set to a lowercase ISO 639-1 code such as `en`, `zh`, or `ja`. Other forms, such as `en-US` or `English`, may be rejected with 400. Pass `language` when you know the source language. Detection still runs, and the hint does not force the transcript into that language; the hint is reported as `language_code` when the language cannot be determined. See [Language](/features/speech-to-text#language) for details. + +Collect one result row per file as you go. If a file cannot be read, its request fails with a network error such as a timeout, or the API returns an error for it, record the error in that file's row and move on to the next file. + ```python Synchronous +import httpx from fishaudio import FishAudio +from fishaudio.core import RequestOptions +from fishaudio.exceptions import APIError client = FishAudio() paths = ["speech.wav"] # add more file paths here language = "en" +options = RequestOptions( + timeout=900, # long recordings can take several minutes + additional_headers={"model": "transcribe-1-pro"}, +) results = [] for path in paths: - with open(path, "rb") as f: - audio = f.read() - - transcript = client.asr.transcribe(audio=audio, language=language) + try: + with open(path, "rb") as f: + audio = f.read() + transcript = client.asr.transcribe( + audio=audio, + language=language, + include_timestamps=False, + request_options=options, + ) + except (OSError, httpx.HTTPError, APIError) as e: + # Record the failure and continue with the next file. + results.append({"file": path, "error": f"{type(e).__name__}: {e}"}) + continue results.append({ "file": path, "text": transcript.text, "duration": transcript.duration, # seconds }) +failed = [row for row in results if "error" in row] for row in results: - print(f"{row['file']} ({row['duration']:.1f}s): {row['text']}") + if "error" in row: + print(f"{row['file']}: FAILED: {row['error']}") + else: + print(f"{row['file']} ({row['duration']:.1f}s): {row['text']}") + +if failed: + raise SystemExit(f"{len(failed)} of {len(results)} files failed") ``` ```python Asynchronous import asyncio + +import httpx from fishaudio import AsyncFishAudio +from fishaudio.core import RequestOptions +from fishaudio.exceptions import APIError async def main(): async with AsyncFishAudio() as client: paths = ["speech.wav"] # add more file paths here language = "en" + options = RequestOptions( + timeout=900, # long recordings can take several minutes + additional_headers={"model": "transcribe-1-pro"}, + ) results = [] for path in paths: - with open(path, "rb") as f: - audio = f.read() - - transcript = await client.asr.transcribe(audio=audio, language=language) + try: + with open(path, "rb") as f: + audio = f.read() + transcript = await client.asr.transcribe( + audio=audio, + language=language, + include_timestamps=False, + request_options=options, + ) + except (OSError, httpx.HTTPError, APIError) as e: + # Record the failure and continue with the next file. + results.append({"file": path, "error": f"{type(e).__name__}: {e}"}) + continue results.append({ "file": path, "text": transcript.text, "duration": transcript.duration, # seconds }) + return results - for row in results: +results = asyncio.run(main()) + +failed = [row for row in results if "error" in row] +for row in results: + if "error" in row: + print(f"{row['file']}: FAILED: {row['error']}") + else: print(f"{row['file']} ({row['duration']:.1f}s): {row['text']}") -asyncio.run(main()) +if failed: + raise SystemExit(f"{len(failed)} of {len(results)} files failed") ``` ```javascript JavaScript @@ -74,30 +127,68 @@ const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY }); const paths = ["speech.wav"]; // add more file paths here const language = "en"; +const options = { + headers: { model: "transcribe-1-pro" }, + timeoutInSeconds: 900, // long recordings can take several minutes +}; const results = []; for (const path of paths) { - const audio = new File([await readFile(path)], path); - - const transcript = await client.speechToText.convert({ audio, language }); - results.push({ - file: path, - text: transcript.text, - duration: transcript.duration, // seconds - }); + try { + const audio = new File([await readFile(path)], path); + const transcript = await client.speechToText.convert( + { audio, language }, + options + ); + results.push({ + file: path, + text: transcript.text, + duration: transcript.duration, // seconds + }); + } catch (err) { + // Record the failure and continue with the next file. + const error = err.statusCode + ? `HTTP ${err.statusCode}: ${err.body?.message ?? "(no JSON body)"}` + : err.message; + results.push({ file: path, error }); + } } +const failed = results.filter(row => row.error); for (const row of results) { - console.log(`${row.file} (${row.duration.toFixed(1)}s): ${row.text}`); + if (row.error) { + console.log(`${row.file}: FAILED: ${row.error}`); + } else { + console.log(`${row.file} (${row.duration.toFixed(1)}s): ${row.text}`); + } +} + +if (failed.length > 0) { + console.error(`${failed.length} of ${results.length} files failed`); + // Exit now: a request that failed with a network error can leave its timeout running. + process.exit(1); } ``` + -Each call returns an [`ASRResponse`](/api-reference/sdk/python/types#asrresponse-objects) with `.text`, a `.duration` in seconds, and per-phrase `.segments`. The loop keeps files independent, so one bad file does not block the rest of the batch. +Each call returns an [`ASRResponse`](/api-reference/sdk/python/types#asrresponse-objects) with `.text` and `.duration` (seconds). With `transcribe-1-pro`, `text` can contain inline speaker markers such as `<|speaker:0|>` and cues such as `[laughter]`; see [Speaker markers](/features/speech-to-text#speaker-markers) to split it into turns. This recipe skips timestamps; pass `include_timestamps=True` (Python) or `ignore_timestamps: false` (JavaScript) for word-level `.segments`. The Python SDK currently returns only `text`, `duration`, and `segments`; to read `language_code`, call the API directly (see [Direct API](/features/speech-to-text#direct-api-messagepack)). + +A file that fails is recorded with its error, and the loop continues. After the summary, the script exits with a non-zero status if any file failed, so a scheduled job does not report success. A file that failed with 429, a 5xx status, or a network error such as a timeout is worth retrying later with backoff; for other 4xx errors, fix the file or the request first. + +Long recordings can take several minutes to transcribe, so every tab sets a 15-minute request timeout, longer than the SDKs' default; in Node.js, also see the note below. + + + In Node.js, the built-in `fetch` that the JavaScript SDK uses stops waiting + for a response after 5 minutes, even with a longer `timeoutInSeconds`. To + wait longer, install `undici@7` and, once at startup, call its + `setGlobalDispatcher(new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 }))`. + See [Processing time and + timeouts](/features/speech-to-text#processing-time-and-timeouts). + - Use one `language` per batch; split mixed-language files into separate - lists. + Use one `language` per batch; split mixed-language files into separate lists. ## Related diff --git a/developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx b/developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx index d137d2f..a73299f 100644 --- a/developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx +++ b/developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx @@ -1,6 +1,6 @@ --- title: "Transcribe audio to SRT/VTT captions" -description: "Transcribe audio with timestamps and write valid SRT and WebVTT caption files from the segments" +description: "Transcribe audio with word-level timestamps, group the words into caption cues, and write valid SRT and WebVTT files" icon: "closed-captioning" --- @@ -12,43 +12,104 @@ import Prerequisites from "/snippets/prerequisites.mdx"; ## Recipe -Call [`asr.transcribe()`](/api-reference/sdk/python/resources#transcribe) with `include_timestamps=True`, then turn each [`ASRSegment`](/api-reference/sdk/python/types#asrsegment-objects) into a numbered cue. Segment `start` / `end` are in **seconds**, so the only real work is formatting them: SRT wants `HH:MM:SS,mmm` (comma), WebVTT wants `HH:MM:SS.mmm` (dot). +This recipe uses `transcribe-1-pro`, which handles recordings up to 60 minutes long. Both tabs select it with the `model` request header; without that header, the request is served and billed as `transcribe-1`. + +Call [`asr.transcribe()`](/api-reference/sdk/python/resources#transcribe) with `include_timestamps=True` (in JavaScript, `ignore_timestamps: false`) to get timed segments ([`ASRSegment`](/api-reference/sdk/python/types#asrsegment-objects)). Segments are word-level (one character, or a few, in Chinese and Japanese) and carry no punctuation, so group consecutive segments into caption-sized cues before you format them. Segment `start` / `end` are in **seconds**; SRT wants `HH:MM:SS,mmm` (comma), WebVTT wants `HH:MM:SS.mmm` (dot). + +The recipe starts a new cue after a pause longer than 0.6 seconds, or when the cue would grow past 42 characters or 6 seconds. It joins words with a space, except between Chinese or Japanese characters, which are written without spaces. A segment can have `start` equal to `end`, so a cue with no duration gets a short one that ends no later than the next cue starts. Every cue therefore ends after it starts, as SRT and WebVTT players expect. + ```python Python +import re + from fishaudio import FishAudio +from fishaudio.core import RequestOptions client = FishAudio() - -def to_srt_timestamp(seconds: float) -> str: - """Format a time in seconds as an SRT timestamp: HH:MM:SS,mmm.""" - millis = round(seconds * 1000) - hours, millis = divmod(millis, 3_600_000) - minutes, millis = divmod(millis, 60_000) - secs, millis = divmod(millis, 1000) - return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}" +MAX_CHARS = 42 # longest cue text, in characters +MAX_MS = 6000 # longest cue, in milliseconds +MAX_GAP_MS = 600 # a longer pause between words starts a new cue +MIN_MS = 500 # display time for a cue whose words have no duration + +# Chinese and Japanese are written without spaces between words. +NO_SPACE = re.compile(r"[\u3040-\u30ff\u3400-\u4dbf\u4e00-\u9fff]") + + +def join_words(left: str, right: str) -> str: + if NO_SPACE.match(left[-1]) and NO_SPACE.match(right[0]): + return left + right + return f"{left} {right}" + + +def group_cues(segments) -> list[dict]: + """Group word-level segments into cues, with start and end in milliseconds.""" + cues = [] + for segment in segments: + word = segment.text.strip() + if not word: + continue + start, end = round(segment.start * 1000), round(segment.end * 1000) + if cues: + cue = cues[-1] + text = join_words(cue["text"], word) + if ( + start - cue["end"] <= MAX_GAP_MS + and len(text) <= MAX_CHARS + and end - cue["start"] <= MAX_MS + ): + cue["text"], cue["end"] = text, max(cue["end"], end) + continue + cues.append({"start": start, "end": end, "text": word}) + + # A segment can have start == end. Give a cue with no duration a short one, + # ending no later than the next cue starts, so every cue ends after it starts. + for i, cue in enumerate(cues): + if cue["end"] <= cue["start"]: + end = cue["start"] + MIN_MS + if i + 1 < len(cues): + end = min(end, cues[i + 1]["start"]) + cue["end"] = max(end, cue["start"] + 1) + return cues + + +def format_timestamp(ms: int, separator: str) -> str: + """HH:MM:SS,mmm for SRT (separator ","), HH:MM:SS.mmm for WebVTT (".").""" + hours, ms = divmod(ms, 3_600_000) + minutes, ms = divmod(ms, 60_000) + seconds, ms = divmod(ms, 1000) + return f"{hours:02d}:{minutes:02d}:{seconds:02d}{separator}{ms:03d}" with open("speech.wav", "rb") as f: - result = client.asr.transcribe(audio=f.read(), include_timestamps=True) + result = client.asr.transcribe( + audio=f.read(), + include_timestamps=True, + request_options=RequestOptions( + timeout=900, # long recordings can take several minutes + additional_headers={"model": "transcribe-1-pro"}, + ), + ) + +cues = group_cues(result.segments) # SRT: 1-based index, comma decimal separator, blank line between cues. with open("captions.srt", "w", encoding="utf-8") as srt: - for i, segment in enumerate(result.segments, start=1): - start = to_srt_timestamp(segment.start) - end = to_srt_timestamp(segment.end) - srt.write(f"{i}\n{start} --> {end}\n{segment.text.strip()}\n\n") + for i, cue in enumerate(cues, start=1): + start = format_timestamp(cue["start"], ",") + end = format_timestamp(cue["end"], ",") + srt.write(f"{i}\n{start} --> {end}\n{cue['text']}\n\n") # WebVTT: same cues, "WEBVTT" header, dot decimal separator. with open("captions.vtt", "w", encoding="utf-8") as vtt: vtt.write("WEBVTT\n\n") - for segment in result.segments: - start = to_srt_timestamp(segment.start).replace(",", ".") - end = to_srt_timestamp(segment.end).replace(",", ".") - vtt.write(f"{start} --> {end}\n{segment.text.strip()}\n\n") + for cue in cues: + start = format_timestamp(cue["start"], ".") + end = format_timestamp(cue["end"], ".") + vtt.write(f"{start} --> {end}\n{cue['text']}\n\n") -print(f"Wrote {len(result.segments)} cues to captions.srt and captions.vtt") +print(f"Wrote {len(cues)} cues to captions.srt and captions.vtt") ``` ```javascript JavaScript @@ -57,46 +118,119 @@ import { readFile, writeFile } from "fs/promises"; const client = new FishAudioClient({ apiKey: process.env.FISH_API_KEY }); -// Format a time in seconds as an SRT timestamp: HH:MM:SS,mmm. -function toSrtTimestamp(seconds) { - let millis = Math.round(seconds * 1000); - const hours = Math.floor(millis / 3_600_000); - millis -= hours * 3_600_000; - const minutes = Math.floor(millis / 60_000); - millis -= minutes * 60_000; - const secs = Math.floor(millis / 1000); - millis -= secs * 1000; +const MAX_CHARS = 42; // longest cue text, in characters +const MAX_MS = 6000; // longest cue, in milliseconds +const MAX_GAP_MS = 600; // a longer pause between words starts a new cue +const MIN_MS = 500; // display time for a cue whose words have no duration + +// Chinese and Japanese are written without spaces between words. +const NO_SPACE = /[\u3040-\u30ff\u3400-\u4dbf\u4e00-\u9fff]/; + +function joinWords(left, right) { + if (NO_SPACE.test(left.at(-1)) && NO_SPACE.test(right[0])) { + return left + right; + } + return `${left} ${right}`; +} + +// Group word-level segments into cues, with start and end in milliseconds. +function groupCues(segments) { + const cues = []; + for (const segment of segments) { + const word = segment.text.trim(); + if (!word) continue; + const start = Math.round(segment.start * 1000); + const end = Math.round(segment.end * 1000); + const cue = cues.at(-1); + if (cue) { + const text = joinWords(cue.text, word); + if ( + start - cue.end <= MAX_GAP_MS && + text.length <= MAX_CHARS && + end - cue.start <= MAX_MS + ) { + cue.text = text; + cue.end = Math.max(cue.end, end); + continue; + } + } + cues.push({ start, end, text: word }); + } + + // A segment can have start == end. Give a cue with no duration a short one, + // ending no later than the next cue starts, so every cue ends after it starts. + cues.forEach((cue, i) => { + if (cue.end <= cue.start) { + let end = cue.start + MIN_MS; + if (i + 1 < cues.length) end = Math.min(end, cues[i + 1].start); + cue.end = Math.max(end, cue.start + 1); + } + }); + return cues; +} + +// HH:MM:SS,mmm for SRT (separator ","), HH:MM:SS.mmm for WebVTT ("."). +function formatTimestamp(ms, separator) { const pad = (n, width) => String(n).padStart(width, "0"); - return `${pad(hours, 2)}:${pad(minutes, 2)}:${pad(secs, 2)},${pad(millis, 3)}`; + const hours = Math.floor(ms / 3_600_000); + const minutes = Math.floor((ms % 3_600_000) / 60_000); + const seconds = Math.floor((ms % 60_000) / 1000); + const millis = ms % 1000; + return `${pad(hours, 2)}:${pad(minutes, 2)}:${pad(seconds, 2)}${separator}${pad(millis, 3)}`; } -const result = await client.speechToText.convert({ - audio: new File([await readFile("speech.wav")], "speech.wav"), - language: "en", - ignore_timestamps: false, -}); +const result = await client.speechToText.convert( + { + audio: new File([await readFile("speech.wav")], "speech.wav"), + ignore_timestamps: false, + }, + { + headers: { model: "transcribe-1-pro" }, + timeoutInSeconds: 900, // long recordings can take several minutes + } +); + +const cues = groupCues(result.segments); +const timing = (cue, separator) => + `${formatTimestamp(cue.start, separator)} --> ${formatTimestamp(cue.end, separator)}`; // SRT: 1-based index, comma decimal separator, blank line between cues. -const cues = result.segments.map((segment, i) => { - const start = toSrtTimestamp(segment.start); - const end = toSrtTimestamp(segment.end); - return `${i + 1}\n${start} --> ${end}\n${segment.text.trim()}\n`; -}); -await writeFile("captions.srt", cues.join("\n"), "utf-8"); - -console.log(`Wrote ${result.segments.length} cues to captions.srt`); +const srt = cues.map( + (cue, i) => `${i + 1}\n${timing(cue, ",")}\n${cue.text}\n\n` +); +await writeFile("captions.srt", srt.join(""), "utf-8"); + +// WebVTT: same cues, "WEBVTT" header, dot decimal separator. +const vtt = cues.map(cue => `${timing(cue, ".")}\n${cue.text}\n\n`); +await writeFile("captions.vtt", "WEBVTT\n\n" + vtt.join(""), "utf-8"); + +console.log(`Wrote ${cues.length} cues to captions.srt and captions.vtt`); ``` + -Both files share one timestamp helper; WebVTT just swaps `,` for `.`. +Both tabs send the same request and write the same `captions.srt` and `captions.vtt`. Long recordings can take several minutes to transcribe, so both tabs set a 15-minute request timeout, longer than the SDKs' default; in Node.js, also see the note below. Tune `MAX_CHARS`, `MAX_MS`, and `MAX_GAP_MS` for your player and audience. `segments` is empty when no speech was found or word timing was temporarily unavailable, and the files then contain no cues, so check the cue count before you publish them. + + + In Node.js, the built-in `fetch` that the JavaScript SDK uses stops waiting + for a response after 5 minutes, even with a longer `timeoutInSeconds`. To + wait longer, install `undici@7` and, once at startup, call its + `setGlobalDispatcher(new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 }))`. + See [Processing time and + timeouts](/features/speech-to-text#processing-time-and-timeouts). + - Pass `language=` (for example `"en"` or `"zh"`) when you know it. Explicit - language selection sharpens segment boundaries, which keeps your cue timing - tight. + Captions built from `segments` have no punctuation. Because this recipe + requests timestamps from `transcribe-1-pro`, the response also includes + `speaker_turns`, with each turn's punctuated text and its start and end time. + To keep speakers in separate cues, start a new cue at each turn boundary. In + JavaScript, read `result.speaker_turns`; it is not in the SDK's TypeScript + type. The Python SDK does not return `speaker_turns`; to read it from the API, + see [Speaker turns](/features/speech-to-text#speaker-turns). ## Related - [Speech-to-Text guide](/features/speech-to-text) -- [ASR Types Reference](/api-reference/sdk/python/types#asr) +- [ASR Types Reference](/api-reference/sdk/python/types#fishaudio.types.asr) diff --git a/developer-guide/sdk-guide/cookbook/voice-agent-loop.mdx b/developer-guide/sdk-guide/cookbook/voice-agent-loop.mdx index 5de90a2..63d5873 100644 --- a/developer-guide/sdk-guide/cookbook/voice-agent-loop.mdx +++ b/developer-guide/sdk-guide/cookbook/voice-agent-loop.mdx @@ -14,9 +14,15 @@ import Prerequisites from "/snippets/prerequisites.mdx"; A voice agent is three stages chained together: [`asr.transcribe()`](/api-reference/sdk/python/resources#transcribe) turns the caller's audio into text, your own LLM turns that text into a reply, and [`tts.stream()`](/api-reference/sdk/python/resources#stream) turns the reply back into speech. The transcript and the reply are just strings, so the only Fish Audio-specific parts are the first and last calls. Streaming the reply lets you start writing (or forwarding) audio before the whole sentence is synthesized. +The recipe transcribes with `transcribe-1-pro`. Every tab selects it with the `model` request header; without that header, the request is served and billed as `transcribe-1`. + + ```python Synchronous +import re + from fishaudio import FishAudio +from fishaudio.core import RequestOptions from fishaudio.utils import save client = FishAudio() @@ -29,9 +35,17 @@ def reply_from_llm(text: str) -> str: def voice_agent_turn(audio_path: str, out_path: str) -> str: with open(audio_path, "rb") as f: - heard = client.asr.transcribe(audio=f.read()) - - reply = reply_from_llm(heard.text) + heard = client.asr.transcribe( + audio=f.read(), + include_timestamps=False, + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"} + ), + ) + + # Drop speaker markers such as <|speaker:0|>; keep cues such as [laughter]. + said = re.sub(r"\s*<\|speaker:\d+\|>\s*", " ", heard.text).strip() + reply = reply_from_llm(said) audio_stream = client.tts.stream(text=reply, reference_id="") save(audio_stream, out_path) # writes chunks as they arrive @@ -43,7 +57,10 @@ print("Agent:", reply) ```python Asynchronous import asyncio +import re + from fishaudio import AsyncFishAudio +from fishaudio.core import RequestOptions from fishaudio.utils import save def reply_from_llm(text: str) -> str: @@ -54,9 +71,17 @@ def reply_from_llm(text: str) -> str: async def main(): async with AsyncFishAudio() as client: with open("speech.wav", "rb") as f: - heard = await client.asr.transcribe(audio=f.read()) - - reply = reply_from_llm(heard.text) + heard = await client.asr.transcribe( + audio=f.read(), + include_timestamps=False, + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"} + ), + ) + + # Drop speaker markers such as <|speaker:0|>; keep cues such as [laughter]. + said = re.sub(r"\s*<\|speaker:\d+\|>\s*", " ", heard.text).strip() + reply = reply_from_llm(said) audio_stream = await client.tts.stream(text=reply, reference_id="") with open("reply.mp3", "wb") as out: @@ -81,12 +106,17 @@ function replyFromLlm(text) { } async function voiceAgentTurn(audioPath, outPath) { - const heard = await client.speechToText.convert({ - audio: new File([await readFile(audioPath)], audioPath), - language: "en", - }); + const heard = await client.speechToText.convert( + { + audio: new File([await readFile(audioPath)], audioPath), + language: "en", + }, + { headers: { model: "transcribe-1-pro" } } + ); - const reply = replyFromLlm(heard.text); + // Drop speaker markers such as <|speaker:0|>; keep cues such as [laughter]. + const said = heard.text.replace(/\s*<\|speaker:\d+\|>\s*/g, " ").trim(); + const reply = replyFromLlm(said); const stream = await client.textToSpeech.convert( { text: reply, reference_id: "", format: "mp3" }, @@ -101,15 +131,16 @@ async function voiceAgentTurn(audioPath, outPath) { const reply = await voiceAgentTurn("speech.wav", "reply.mp3"); console.log("Agent:", reply); ``` + -`heard` is an [`ASRResponse`](/api-reference/sdk/python/types#asrresponse-objects): `heard.text` is the full transcript and `heard.duration` is the clip length in seconds. Pass `language="en"` to `transcribe()` to skip auto-detection when you already know the input language. +`heard` is an [`ASRResponse`](/api-reference/sdk/python/types#asrresponse-objects): `heard.text` is the full transcript and `heard.duration` is the clip length in seconds. With `transcribe-1-pro`, `heard.text` can contain inline speaker markers such as `<|speaker:0|>` and cues such as `[laughter]`. The recipe removes the speaker markers before it calls your LLM and keeps the cues as context; see [Speaker markers](/features/speech-to-text#speaker-markers). The Python calls pass `include_timestamps=False` because the loop needs only the text, and word timestamps add processing time. Pass `language="en"` as a hint when you know the caller's language; see [Language](/features/speech-to-text#language). For the lowest latency, feed your LLM's token stream straight into [`stream_websocket()`](/api-reference/sdk/python/resources#stream_websocket) - instead of waiting for the full reply string. See - [Realtime: LLM tokens → speech](/developer-guide/sdk-guide/cookbook/realtime-llm-to-speech). + instead of waiting for the full reply string. See [Realtime: LLM tokens → + speech](/developer-guide/sdk-guide/cookbook/realtime-llm-to-speech). ## Reply in the caller's voice diff --git a/features/speech-to-text.mdx b/features/speech-to-text.mdx index c8119e2..c4f9c25 100644 --- a/features/speech-to-text.mdx +++ b/features/speech-to-text.mdx @@ -1,10 +1,10 @@ --- title: "Speech to Text" -description: "Transcribe audio with transcribe-1 or transcribe-1-pro, including multi-speaker conversations and emotion cues" +description: "Transcribe audio with transcribe-1-pro, including word-level timestamps, speaker turns for multi-speaker conversations, long recordings, and emotion cues" icon: "waveform" --- -Turn spoken audio into text using Fish Audio's automatic speech recognition (ASR) models. Send an audio file to `POST /v1/asr` to receive a transcript, its duration, and optional timestamped segments. Use `transcribe-1-pro` for multi-speaker conversations and transcripts that preserve emotion and vocal-event cues. +Turn spoken audio into text with `transcribe-1-pro`, Fish Audio's recommended automatic speech recognition (ASR) model. Send an audio file to `POST /v1/asr` with the header `model: transcribe-1-pro` to receive a transcript, its duration, and optional word-level timestamps. Pro handles multi-speaker conversations and recordings up to 60 minutes: it marks speakers inline in the transcript, returns a structured list of speaker turns when you request timestamps, and keeps emotion and vocal-event cues. ` markers in the response's `text` string**, where `N` is a numeric speaker label. A marker assigns the following text to that speaker until the next marker. The same label can appear again when that speaker resumes talking. - -For example, this illustrative response contains three turns from two speakers. It uses `ignore_timestamps=true`, so `segments` is empty: - -```json -{ - "text": "<|speaker:0|>你好。<|speaker:1|>[高兴]很开心认识你。<|speaker:0|>我也是。", - "duration": 6.4, - "segments": [], - "language_code": "zh", - "language": "Chinese" -} -``` - -| Marker | How to read this example | -| ----------------- | ----------------------------------------------------------------------- | -| `<\|speaker:0\|>` | Speaker 0 says `你好。`, then later `我也是。`. | -| `<\|speaker:1\|>` | Speaker 1 says `很开心认识你。`, with the emotion cue `[高兴]` (happy). | - -Treat speaker labels as identifiers within that recording, not as names or identities you can match across separate requests. The response contains one annotated `text` string, not a JSON array of speaker turns. To display turns separately, split the string at speaker markers, preserving any emotion cues in each turn. For example, after decoding the JSON response into `result`: - -```python -import re - -pattern = r"<\|speaker:(\d+)\|>(.*?)(?=<\|speaker:\d+\|>|$)" -for turn in re.finditer(pattern, result["text"], re.DOTALL): - print(f"Speaker {turn.group(1)}: {turn.group(2).strip()}") -``` - -```text -Speaker 0: 你好。 -Speaker 1: [高兴]很开心认识你。 -Speaker 0: 我也是。 -``` - -With `ignore_timestamps=false`, `text` keeps the speaker markers, while `segments` contains aligned speech with `text`, `start`, and `end`. Speaker markers and emotion cues are excluded from timestamp alignment. There is no separate speaker list or per-segment `speaker_id`, so use the markers in `text` to read speaker turns; do not assume each segment corresponds to a turn. - -### Emotion and vocal-event cues - -Pro can retain cues about how speech sounds alongside the spoken words. These appear as inline bracketed text, such as `[高兴]` (happy) or `[laughter]`, in the response's `text` field. For example, a transcript might contain: - -```text -你好,[高兴]很开心认识你 -``` - -These are annotations inferred from the audio, not words you need to add to the request. The API has no separate `emotion` field or fixed emotion enum; preserve the returned text when your application needs these cues. - -Timestamp alignment uses the spoken words with speaker, emotion, and event markers removed. Read `text` for the annotated transcript and `segments` for timed speech; the segment text may therefore differ from the full transcript. +If you use `transcribe-1`, select it with `model: transcribe-1` and send only `audio`, `language`, and `ignore_timestamps`. The other fields apply to `transcribe-1-pro` only. ## When to use it - - Timed segments map straight to SRT/VTT cues. + + Group word-level timestamps into SRT/VTT cues. Use Pro to transcribe multi-speaker recordings for summaries and search. @@ -113,7 +77,7 @@ Timestamp alignment uses the spoken words with speaker, emotion, and event marke Create an [API key](/developer-guide/getting-started/api-key) and set `FISH_API_KEY` in your environment. For the Python examples, [install the Fish Audio SDK](/developer-guide/sdk-guide/quickstart). -The examples below select `transcribe-1-pro` and request timestamps. Set the header to `transcribe-1` to use the standard model. +Every example on this page selects `transcribe-1-pro`. The examples below also request timestamps. @@ -128,7 +92,8 @@ with open("speech.wav", "rb") as f: audio=f.read(), include_timestamps=True, request_options=RequestOptions( - additional_headers={"model": "transcribe-1-pro"} + timeout=900, # long recordings can take several minutes + additional_headers={"model": "transcribe-1-pro"}, ), ) @@ -167,18 +132,25 @@ console.log(result.text); -The response gives you the full `text`, the audio `duration` in seconds, and `segments`. When available, it also includes the detected `language_code` and display name `language`. +The response contains `text`, `duration` (seconds), `segments`, and `request_id`, plus [`speaker_turns`](#speaker-turns) when timestamps are requested, and `language_code` and `language` when the language is known. Do not depend on the order of keys in the JSON. If you use `transcribe-1`, rely on `text`, `duration`, `segments`, and the language fields. + +The Python SDK currently returns only `text`, `duration`, and `segments`. To read the other fields, call the API directly (see [Direct API](#direct-api-messagepack)). The Python SDK selects Pro through `RequestOptions.additional_headers`; - `asr.transcribe()` has no `model` argument. The JavaScript example calls the + `asr.transcribe()` has no `model` argument, and it cannot send the + `transcribe-1-pro` fields such as `diarize`. The JavaScript example calls the REST API directly so the header is explicit. For multipart uploads, let your HTTP client set `Content-Type` and its boundary. ## Read the timestamps -Each segment carries `text`, `start`, and `end`, with times in **seconds**. With the API, request timestamps with `ignore_timestamps=false`. The default is `true`, which returns an empty `segments` array. Alignment adds processing time, and `segments` can also be empty when alignment is unavailable or no speech is detected. +Request timestamps with `ignore_timestamps=false`. The default is `true`, which skips them. Timestamps add processing time. In multipart forms, send `true` or `false`: any other value, including `1` or an empty value, is read as `false` and turns timestamps on. + +Each segment is `{ "text", "start", "end" }`, with times in **seconds**. Segments are word-level: a segment is usually one word, or in Chinese and Japanese usually one character or a few. Segment text has no punctuation and no speaker markers or cues, and can be normalized (for example `35` for `3.5`), so it does not always match `text` character for character. A segment can have `start` equal to `end`. + +`segments` is an empty array, never omitted, when you do not request timestamps, when no speech is found, or when timing is temporarily unavailable. Segments are not speaker turns; to see who spoke when, use [`speaker_turns`](#speaker-turns). @@ -209,26 +181,202 @@ curl --request POST https://api.fish.audio/v1/asr \ +For captions, group consecutive segments into caption-sized cues instead of writing one cue per word. The [captions cookbook](/developer-guide/sdk-guide/cookbook/transcribe-to-captions) shows how. + In the Python SDK, segment timestamps are **on by default**. Pass `include_timestamps=False` to skip them. That's the *inverse* of the API/JavaScript flag `ignore_timestamps`. +## Multi-speaker conversations + +`transcribe-1-pro` supports recordings with multiple speakers, such as interviews, meetings, and calls. Send the recording as one audio file. Speaker labels are consistent within one response, not across requests, so do not split a conversation into several requests. + +### Speaker markers + +Pro identifies speakers with **inline `<|speaker:N|>` markers in the response's `text` string**, where `N` is a numeric speaker label. A marker assigns the following text to that speaker until the next marker. The same label can appear again when that speaker resumes talking. + +Markers usually have a space on each side (`<|speaker:0|> 你好。 <|speaker:1|> ...`); do not rely on exact spacing. Any text before the first marker belongs to the first turn. A transcript with no marker comes from a single speaker; treat it as speaker 0. + +For example, this illustrative response contains three turns from two speakers. It uses `ignore_timestamps=true`, so `segments` is empty: + +```json +{ + "text": "<|speaker:0|> 你好。 <|speaker:1|> [高兴]很开心认识你。 <|speaker:0|> 我也是。", + "duration": 6.4, + "segments": [], + "language_code": "zh", + "language": "Chinese", + "request_id": "0b6f4c1e-7d2a-4e8b-9f3c-5a1d2e3f4b6c" +} +``` + +| Marker | How to read this example | +| ----------------- | ----------------------------------------------------------------------- | +| `<\|speaker:0\|>` | Speaker 0 says `你好。`, then later `我也是。`. | +| `<\|speaker:1\|>` | Speaker 1 says `很开心认识你。`, with the emotion cue `[高兴]` (happy). | + +Treat speaker labels as identifiers within that response, not as names or identities you can match across separate requests. To display turns without requesting timestamps, split `text` at the speaker markers, keeping any emotion cues in each turn. For example, after decoding the JSON response into `result`: + +```python +import re + +parts = re.split(r"<\|speaker:(\d+)\|>", result["text"]) +lead, rest = parts[0].strip(), parts[1:] +turns = [(rest[i], rest[i + 1].strip()) for i in range(0, len(rest), 2)] +if not turns: + turns = [("0", lead)] # no marker: a single speaker +elif lead: + turns[0] = (turns[0][0], f"{lead} {turns[0][1]}".strip()) +for speaker, text in turns: + if text: + print(f"Speaker {speaker}: {text}") +``` + +```text +Speaker 0: 你好。 +Speaker 1: [高兴]很开心认识你。 +Speaker 0: 我也是。 +``` + +When you request timestamps, prefer `speaker_turns` over parsing `text`. + +### Speaker turns + +With `ignore_timestamps=false`, `transcribe-1-pro` also returns `speaker_turns`, a list of who spoke when. The field is present when `ignore_timestamps=false` and `diarize` is not `false`; otherwise it is absent. Speaker identification is a `transcribe-1-pro` feature. + +This illustrative response is the same conversation with timestamps requested: + +```json +{ + "text": "<|speaker:0|> 你好。 <|speaker:1|> [高兴]很开心认识你。 <|speaker:0|> 我也是。", + "duration": 6.4, + "segments": [ + { "text": "你", "start": 0.32, "end": 0.56 }, + { "text": "好", "start": 0.56, "end": 0.88 }, + { "text": "很", "start": 1.84, "end": 2.04 }, + { "text": "开", "start": 2.04, "end": 2.24 }, + { "text": "心", "start": 2.24, "end": 2.48 }, + { "text": "认", "start": 2.48, "end": 2.68 }, + { "text": "识", "start": 2.68, "end": 2.88 }, + { "text": "你", "start": 2.88, "end": 3.2 }, + { "text": "我", "start": 4.56, "end": 4.8 }, + { "text": "也", "start": 4.8, "end": 5.0 }, + { "text": "是", "start": 5.0, "end": 5.36 } + ], + "speaker_turns": [ + { "speaker": "speaker:0", "text": "你好。", "start": 0.32, "end": 0.88 }, + { + "speaker": "speaker:1", + "text": "[高兴]很开心认识你。", + "start": 1.84, + "end": 3.2 + }, + { "speaker": "speaker:0", "text": "我也是。", "start": 4.56, "end": 5.36 } + ], + "language_code": "zh", + "language": "Chinese", + "request_id": "0b6f4c1e-7d2a-4e8b-9f3c-5a1d2e3f4b6c" +} +``` + +- `speaker_turns` lists the turns in the order they occur, and is `[]` when no speech was found. Each turn is `{ "speaker", "text", "start", "end" }`. +- `speaker` is the string `speaker:N`, where `N` matches the `<|speaker:N|>` marker in `text`. Labels identify speakers within one response only. They are not names, and the same label in two requests is not the same person. +- `text` is that turn's speech without speaker markers. It keeps emotion and event cues unless `tag_audio_events=false`. +- `start` and `end` are in seconds and come from the word timestamps. If word timing is unavailable (`segments` is empty), turn times are approximate and can cover the whole recording. +- Consecutive turns can have the same speaker. Do not assume turns are contiguous or non-overlapping. + +### Speaker options + +These `transcribe-1-pro` fields control speaker identification: + +- `diarize`: `auto` (default) or `true` returns `speaker_turns` when timestamps are requested. `false` omits `speaker_turns`. It does not change the transcript, and the speaker markers stay in `text`. +- `num_speakers`: the expected number of speakers. This is a hint, applied on a best-effort basis; it guides speaker identification on longer recordings and may have no effect on short ones. It cannot be combined with `min_speakers` or `max_speakers`. +- `min_speakers`, `max_speakers`: bounds on the number of speakers, with the same best-effort rule. `min_speakers` must not exceed `max_speakers`. + +Speaker counts are integers of 1 or more. Speaker-count hints cannot be sent with `diarize=false`. Invalid values return 400 with the code `invalid_parameter`. + +For example, to transcribe an interview with two speakers and print each turn: + + + +```python Python (httpx) +import os + +import httpx + +with open("interview.mp3", "rb") as f: + resp = httpx.post( + "https://api.fish.audio/v1/asr", + headers={ + "Authorization": f"Bearer {os.environ['FISH_API_KEY']}", + "model": "transcribe-1-pro", + }, + files={"audio": ("interview.mp3", f)}, + data={"ignore_timestamps": "false", "num_speakers": "2"}, + # Long recordings can take several minutes to process. + timeout=httpx.Timeout(900.0, connect=10.0), + ) +resp.raise_for_status() +result = resp.json() + +for turn in result.get("speaker_turns", []): + print(f"[{turn['start']:7.2f}s] {turn['speaker']}: {turn['text']}") +``` + +```bash API (curl) +curl --request POST https://api.fish.audio/v1/asr \ + --header "Authorization: Bearer $FISH_API_KEY" \ + --header "model: transcribe-1-pro" \ + --form audio=@interview.mp3 \ + --form ignore_timestamps=false \ + --form num_speakers=2 | jq '.speaker_turns' +``` + + + +## Emotion and vocal-event cues + +Pro can retain cues about how speech sounds alongside the spoken words. These appear as inline bracketed text, such as `[高兴]` (happy) or `[laughter]`, in the response's `text` field and in the turn text of `speaker_turns`. For example, a transcript might contain: + +```text +你好,[高兴]很开心认识你 +``` + +These are annotations inferred from the audio, not words you need to add to the request. The API has no separate `emotion` field or fixed emotion enum; preserve the returned text when your application needs these cues. + +To get a transcript without cues, send `tag_audio_events=false` (`transcribe-1-pro`). Cues are then removed from `text` and `speaker_turns`; timestamps and billing do not change. + +Word timestamps ignore speaker markers and cues. Read `text` for the annotated transcript and `segments` for timed speech; the segment text may therefore differ from the full transcript. + ## Implementation details ### Language -`language` is an optional hint, such as `en`, `zh`, or `ja`. Language detection still runs when you provide a hint; it does not force the returned language. Use the response's `language_code` in application logic and `language` for display when those fields are available. +`language` is optional. The language is detected automatically whether or not you send a hint, and a hint does not force the transcript into that language. With `transcribe-1-pro`, the hint does not change the transcript; it matters only when the language cannot be determined, as described below. + +Use a lowercase ISO 639-1 code such as `en`, `zh`, or `ja`. Other forms, such as `en-US` or `English`, may be rejected with 400. + +The response reports the language in two fields, which are omitted when unknown (never `null`): + +- `language_code`: the two-letter ISO 639-1 code, such as `en`. Use it in application logic. +- `language`: the language's English name, such as `English` or `Chinese`. Use it for display. + +If the language cannot be determined (for example, very short audio), `language_code` reports your hint, which is not checked against the audio, and `language` may be absent. Responses report one language: for recordings that switch languages, `language` and `language_code` do not list every language spoken. The Python SDK does not return these fields; call the API directly to read them. ```python Python +pro = RequestOptions(additional_headers={"model": "transcribe-1-pro"}) + # Auto-detect -result = client.asr.transcribe(audio=audio_bytes) +result = client.asr.transcribe(audio=audio_bytes, request_options=pro) # Provide a language hint -result = client.asr.transcribe(audio=audio_bytes, language="zh") +result = client.asr.transcribe( + audio=audio_bytes, language="zh", request_options=pro +) ``` @@ -236,6 +384,7 @@ result = client.asr.transcribe(audio=audio_bytes, language="zh") # Omit the form field to auto-detect, or set it explicitly: curl --request POST https://api.fish.audio/v1/asr \ --header "Authorization: Bearer $FISH_API_KEY" \ + --header "model: transcribe-1-pro" \ --form audio=@speech.wav \ --form language=zh ``` @@ -244,11 +393,120 @@ curl --request POST https://api.fish.audio/v1/asr \ ### Input audio -Common formats work directly: `wav`, `mp3`, `opus`, and more. Send the raw file bytes; no pre-processing required. The endpoint accepts `multipart/form-data` (shown above) or `application/msgpack`. +`transcribe-1-pro` accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file. Send the original file bytes; no conversion is needed. + +Not supported: AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM. These return 400. + +Send audio as `multipart/form-data` (a file upload, shown above) or `application/msgpack` (see [Direct API](#direct-api-messagepack)). Base64-encoded audio in a JSON body is not supported. + +If you use `transcribe-1`, send WAV, MP3, AAC (including M4A/MP4), FLAC, or Ogg (Opus or Vorbis), and convert WebM recordings (for example, from a browser's MediaRecorder) to Ogg/Opus, MP3, or WAV first, or use `transcribe-1-pro`. + +### Limits + +These limits apply to `transcribe-1-pro`: + +- Recordings up to 60 minutes long. Longer audio returns 400 `audio_too_long`. +- For long recordings, send compressed audio such as MP3, Opus, or AAC. An hour of 128 kbps MP3 is about 55 MiB. A request that is too large returns 413. +- Very short clips (under about 0.08 seconds) return 400 `audio_too_short`. + +If you use `transcribe-1`, send up to 50 MiB per request and keep MP3 and Opus files under 25 MiB; a request that exceeds the size limit returns 413 or 400. `transcribe-1` is designed for short recordings; for recordings longer than a few minutes, use `transcribe-1-pro`. A long request can fail with 503 if it exceeds the processing-time limit. Retry, and if it keeps failing, use `transcribe-1-pro` or split the audio. ### Long recordings -One request transcribes one audio file. For long recordings, split the audio into shorter clips and transcribe each, then offset each chunk's `start`/`end` by where it began in the full recording. Check that each response covers the complete clip before combining transcripts. +With `transcribe-1-pro`, send the whole recording, up to 60 minutes, in one request. Timestamps and speaker labels cover the whole recording. Do not split conversations: speaker labels are consistent only within one response. Long recordings can take several minutes, so raise your [client timeout](#processing-time-and-timeouts) and retry on 5xx errors. You are billed for the full duration of the recording once. + +If you use `transcribe-1`, keep each request short. For a long recording, use `transcribe-1-pro`, or split the audio into shorter clips, transcribe each, and offset each clip's `start`/`end` by where it began in the full recording. Check that each response covers the complete clip before combining transcripts. + +### Processing time and timeouts + +Each request returns only when the whole file has been transcribed. Processing time grows with the length of the recording, and timestamps add to it. Long `transcribe-1-pro` recordings can take several minutes. + +Set your HTTP client's timeout accordingly. The official SDKs' default timeout can be too short for long recordings, and many HTTP libraries default to even less. To wait up to 15 minutes, as several examples on this page do: + +- Python SDK: `RequestOptions(timeout=900)` for one request, or `FishAudio(timeout=900)` for the client. +- httpx: `timeout=httpx.Timeout(900.0, connect=10.0)`. Without it, httpx waits only 5 seconds. +- Node.js: the built-in `fetch` stops waiting for a response after 5 minutes, even if you pass a longer `AbortSignal`. For longer requests, use `fetch` from the `undici` package with a dispatcher that waits longer. `undici` 7 runs on Node.js 20.18.1 and later; `undici` 8 requires Node.js 22.19 or later. + +```javascript +import { readFile } from "node:fs/promises"; +import { Agent, FormData, fetch } from "undici"; // npm install undici@7 + +// Node's built-in fetch stops waiting for a response after 5 minutes. +// This dispatcher waits up to 15 minutes. +const dispatcher = new Agent({ + headersTimeout: 900_000, + bodyTimeout: 900_000, +}); + +const form = new FormData(); +form.append("audio", new Blob([await readFile("meeting.mp3")]), "meeting.mp3"); +form.append("ignore_timestamps", "false"); + +const response = await fetch("https://api.fish.audio/v1/asr", { + method: "POST", + headers: { + Authorization: `Bearer ${process.env.FISH_API_KEY}`, + model: "transcribe-1-pro", + }, + body: form, + dispatcher, +}); + +if (!response.ok) throw new Error(await response.text()); +const result = await response.json(); +console.log(result.text); +``` + +These values are examples, not a server guarantee. If a long request fails with a 5xx error or the connection drops, retry it. Each request also holds one of your account's concurrent request slots until its response is returned; see [concurrent request limits](/developer-guide/models-pricing/pricing-and-rate-limits#concurrent-request-limits). + +### Request IDs + +`transcribe-1-pro` returns a unique `request_id` in the response body, both on success and on its own errors, and the same value in the `x-request-id` response header. Platform and network-edge errors do not carry it. Log it, and include it when you contact support. The Python SDK does not return it on success; read it with a [direct API call](#direct-api-messagepack), or from the error body as shown in [Errors](#errors). + +### Errors + +The error body depends on where the error comes from: + +- `transcribe-1-pro` errors are JSON with `status`, `message`, `code`, and `request_id`. Branch on `code`, not on `message`. +- Platform errors, such as a missing or invalid API key (401), insufficient API credit (402), or the concurrency limit (429), are JSON with `status` and `message`. +- Errors from the network edge in front of the API, such as some 413 and 5xx responses, may not have a JSON body. + +| Status | `code` | What to do | +| ------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | +| 400 | `invalid_request`, `invalid_parameter` | Fix the request or the parameter value. | +| 400 | `invalid_audio` | Convert the audio to a [supported format](#input-audio). | +| 400 | `audio_too_long`, `audio_too_short` | Send audio within the [limits](#limits). | +| 413 | `request_too_large` | Send a smaller file, for example compressed audio. | +| 415 | `unsupported_media_type` | Send `multipart/form-data` or `application/msgpack`. | +| 429 | | Usually your account is at its concurrency limit. Retry with backoff. | +| 5xx | For example `upstream_unavailable`, `upstream_timeout`, `diarization_failed` | The service is temporarily unavailable or could not finish. Retry with backoff. | + +Retry 429 and 5xx responses with exponential backoff; 429 responses have no `Retry-After` header. Do not retry other 4xx responses; change the request first. New `code` values may be added; handle unknown codes by HTTP status. See the [API reference](/api-reference/endpoint/openapi-v1/speech-to-text#errors) for every status and code. + +With the Python SDK, a failed request raises `APIError` (`RateLimitError` for 429, `ServerError` for 5xx). `e.status` is the HTTP status, and `e.body` is the raw response body, where `transcribe-1-pro` errors carry `code` and `request_id`: + +```python +import json + +from fishaudio.exceptions import APIError + +try: + result = client.asr.transcribe( + audio=audio_bytes, + request_options=RequestOptions( + additional_headers={"model": "transcribe-1-pro"} + ), + ) +except APIError as e: + try: + error = json.loads(e.body or "{}") + except ValueError: + error = {} # for example, an error without a JSON body + print(e.status, error.get("code"), error.get("request_id")) + raise +``` + +If you use `transcribe-1`, its errors are JSON with `status` and `message`; rely only on these two fields. ### Async transcription @@ -265,7 +523,8 @@ async def main(): result = await client.asr.transcribe( audio=f.read(), request_options=RequestOptions( - additional_headers={"model": "transcribe-1-pro"} + timeout=900, + additional_headers={"model": "transcribe-1-pro"}, ), ) print(result.text) @@ -278,12 +537,16 @@ To run several files in parallel, gather the coroutines: ```python import asyncio from fishaudio import AsyncFishAudio +from fishaudio.core import RequestOptions async def transcribe_all(paths): client = AsyncFishAudio() + pro = RequestOptions( + timeout=900, additional_headers={"model": "transcribe-1-pro"} + ) clips = [open(p, "rb").read() for p in paths] return await asyncio.gather(*[ - client.asr.transcribe(audio=clip, language="en") for clip in clips + client.asr.transcribe(audio=clip, request_options=pro) for clip in clips ]) for result in asyncio.run(transcribe_all(["speech.wav"])): @@ -292,7 +555,7 @@ for result in asyncio.run(transcribe_all(["speech.wav"])): ### Direct API (MessagePack) -`POST /v1/asr` also accepts a [MessagePack](https://msgpack.org) body instead of multipart form data, the same path the API reference links to for low-overhead, server-side calls. Pack the audio bytes and options into one payload and set `Content-Type: application/msgpack`: +`POST /v1/asr` also accepts a [MessagePack](https://msgpack.org) body instead of multipart form data, the same path the API reference links to for low-overhead, server-side calls. Calling the API directly also gives you the response fields the Python SDK does not return. Pack the audio bytes and options into one payload and set `Content-Type: application/msgpack`: ```python import os @@ -312,13 +575,18 @@ resp = httpx.post( "Content-Type": "application/msgpack", "model": "transcribe-1-pro", }, + # httpx waits 5 seconds by default; long recordings can take minutes. + timeout=httpx.Timeout(900.0, connect=10.0), ) resp.raise_for_status() result = resp.json() print(result["text"]) +print(result.get("language_code"), result.get("request_id")) +for turn in result.get("speaker_turns", []): + print(turn["speaker"], turn["start"], turn["end"], turn["text"]) ``` -The response shape is identical to the multipart path: `text`, `duration` (seconds), `segments`, and optional detected-language fields. Model selection stays in the HTTP header for both formats. +The response is the same as for multipart: `text`, `duration` (seconds), `segments`, and `request_id`, plus `speaker_turns` when timestamps are requested, and `language_code` and `language` when the language is known. Model selection stays in the HTTP header for both formats. MessagePack values are typed: send booleans for `ignore_timestamps` and `tag_audio_events`, integers for speaker counts, and a boolean or `"auto"` for `diarize`. ## Going further diff --git a/llms.txt b/llms.txt index 8f63a68..a7d132d 100644 --- a/llms.txt +++ b/llms.txt @@ -17,7 +17,7 @@ - [API Introduction](https://docs.fish.audio/api-reference/introduction.md): How to use the Fish Audio API. - [Text to Speech Endpoint](https://docs.fish.audio/api-reference/endpoint/openapi-v1/text-to-speech.md): Convert text to speech. -- [Speech to Text Endpoint](https://docs.fish.audio/api-reference/endpoint/openapi-v1/speech-to-text.md): Select transcribe-1 or transcribe-1-pro with the model header; request timestamps and read transcript and language fields. +- [Speech to Text Endpoint](https://docs.fish.audio/api-reference/endpoint/openapi-v1/speech-to-text.md): Transcribe audio with transcribe-1-pro (recommended; select it with the model header) or transcribe-1; request fields, timestamps, speaker turns, limits, formats and errors. - [Voice Design Endpoint](https://docs.fish.audio/api-reference/endpoint/openapi-v1/voice-design.md): Generate candidate voices from a prompt. - [List Models](https://docs.fish.audio/api-reference/endpoint/model/list-models.md): Get a list of all models. - [Create Model](https://docs.fish.audio/api-reference/endpoint/model/create-model.md): Create a new voice model. @@ -42,7 +42,7 @@ ## Product Guides - [Text to Speech Guide](https://docs.fish.audio/developer-guide/core-features/text-to-speech.md): Convert text to natural-sounding speech with Fish Audio. -- [Speech to Text Guide](https://docs.fish.audio/features/speech-to-text.md): Transcribe audio, including multi-speaker conversations and inline emotion cues with transcribe-1-pro. +- [Speech to Text Guide](https://docs.fish.audio/features/speech-to-text.md): Transcribe audio with transcribe-1-pro, including multi-speaker conversations with speaker turns, long recordings and inline emotion cues. - [Voice Design Guide](https://docs.fish.audio/features/voice-design.md): Generate candidate voices from natural-language prompts. - [Creating Voice Models](https://docs.fish.audio/developer-guide/core-features/creating-models.md): Learn how to create custom voice models with Fish Audio. - [Emotion Control](https://docs.fish.audio/developer-guide/core-features/emotions.md): Add natural emotions and expressions to your AI-generated speech. @@ -88,7 +88,7 @@ ## Operational Docs -- [Models Overview](https://docs.fish.audio/developer-guide/models-pricing/models-overview.md): Explore Fish Audio's voice generation models and their capabilities. +- [Models Overview](https://docs.fish.audio/developer-guide/models-pricing/models-overview.md): Explore Fish Audio's speech generation and transcription models and their capabilities. - [Choosing a Model](https://docs.fish.audio/developer-guide/models-pricing/choosing-a-model.md): Select the right Fish Audio model for your use case and requirements. - [Pricing And Rate Limits](https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits.md): Understand pricing, usage costs, and API rate limits. - [Model Deprecations](https://docs.fish.audio/developer-guide/models-pricing/deprecations.md): Track deprecated models and migration timelines. diff --git a/overview/capabilities.mdx b/overview/capabilities.mdx index 3702526..1f56f25 100644 --- a/overview/capabilities.mdx +++ b/overview/capabilities.mdx @@ -15,7 +15,7 @@ Fish Audio is a voice AI platform. Every core feature is available three ways: i - Transcribe audio with optional timestamps. Use `transcribe-1-pro` for multi-speaker conversations and emotion cues. + Transcribe audio with `transcribe-1-pro`: multi-speaker conversations, long recordings, emotion cues, and optional timestamps. @@ -62,9 +62,12 @@ These text-to-speech models power most capabilities: - **`s2-pro`**: the previous-generation S2 model, with multi-speaker and natural-language expression control. - **`s1`**: the previous generation, with `(parenthesis)` emotion tags. -For speech to text, use **`transcribe-1`** for general transcription or **`transcribe-1-pro`** for multi-speaker conversations and inline emotion cues. Select either model with the `model` header on `POST /v1/asr`. +For speech to text, select the model with the `model` header on `POST /v1/asr`: -See [Models Overview](/developer-guide/models-pricing/models-overview) and [Choosing a Model](/developer-guide/models-pricing/choosing-a-model) for the full lineup, languages, and limits. +- **`transcribe-1-pro`**: the recommended model, for everything from short clips to multi-speaker conversations (speaker markers and speaker turns) and long recordings, with inline emotion cues. Send `model: transcribe-1-pro` explicitly. +- **`transcribe-1`**: general transcription of short recordings. A request whose `model` header is missing or not an exact match is served by this model. + +See [Models Overview](/developer-guide/models-pricing/models-overview) and [Choosing a Model](/developer-guide/models-pricing/choosing-a-model) for the full lineup and languages, and [Speech to Text](/features/speech-to-text#input-audio) for transcription formats and limits. ## Pick your path diff --git a/tests/cookbooks/specs.py b/tests/cookbooks/specs.py index 3998a61..d58b7be 100644 --- a/tests/cookbooks/specs.py +++ b/tests/cookbooks/specs.py @@ -47,17 +47,32 @@ ], }, # ---- recipes authored by the cookbook workflow (one live-tested primary block each) ---- + # The speech-to-text recipes select transcribe-1-pro with the `model` header. { "slug": "transcribe-to-captions", "path": f"{COOKBOOK}/transcribe-to-captions.mdx", - "cases": [{"name": "SRT/VTT captions", "block": 0, "file": ("captions.srt", "srt")}], + "cases": [ + { + "name": "SRT/VTT captions", + "block": 0, + "file": ("captions.srt", "srt"), + # Cues are grouped words, each ending after it starts; the VTT file is written too. + "postamble": ( + "assert cues and all(c['end'] > c['start'] for c in cues), cues\n" + "assert open('captions.vtt', encoding='utf-8').read().startswith('WEBVTT\\n')" + ), + } + ], }, { "slug": "batch-transcribe-with-language-hint", "path": f"{COOKBOOK}/batch-transcribe-with-language-hint.mdx", "cases": [ - {"name": "batch transcribe (sync)", "block": 0, "truthy": "results"}, - {"name": "batch transcribe (async)", "block": 1}, # runs to completion = pass + # The recipes record per-file errors and exit non-zero; also fail on any recorded error. + {"name": "batch transcribe (sync)", "block": 0, "truthy": "results", + "postamble": "assert all('error' not in r for r in results), results"}, + {"name": "batch transcribe (async)", "block": 1, "truthy": "results", + "postamble": "assert all('error' not in r for r in results), results"}, ], }, { @@ -100,8 +115,10 @@ "slug": "voice-agent-loop", "path": f"{COOKBOOK}/voice-agent-loop.mdx", "cases": [ + # The recipe strips transcribe-1-pro speaker markers before the LLM call. {"name": "asr -> reply -> tts (sync)", "block": 0, "file": ("reply.mp3", "mp3"), - "subs": {'""': f'"{PUBLIC_VOICE}"'}}, + "subs": {'""': f'"{PUBLIC_VOICE}"'}, + "postamble": "assert '<|speaker' not in reply, reply"}, {"name": "asr -> reply -> tts (async)", "block": 1, "file": ("reply.mp3", "mp3"), "subs": {'""': f'"{PUBLIC_VOICE}"'}}, ], diff --git a/tests/js/specs.mjs b/tests/js/specs.mjs index 10ab859..9037e7f 100644 --- a/tests/js/specs.mjs +++ b/tests/js/specs.mjs @@ -7,7 +7,7 @@ export const SPECS = [ { slug: "realtime-streaming", mdx: "features/realtime-streaming.mdx", cases: [{ name: "primary", block: 0, file: ["out.mp3", "mp3"] }] }, { slug: "streaming-to-file", mdx: "developer-guide/sdk-guide/cookbook/streaming-to-file.mdx", cases: [{ name: "primary", block: 0, file: ["output.mp3", "mp3"] }] }, { slug: "instant-voice-cloning", mdx: "developer-guide/sdk-guide/cookbook/instant-voice-cloning.mdx", cases: [{ name: "primary", block: 0, file: ["cloned.mp3", "mp3"] }] }, - { slug: "transcribe-to-captions", mdx: "developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx", cases: [{ name: "primary", block: 0, file: ["captions.srt", "srt"] }] }, + { slug: "transcribe-to-captions", mdx: "developer-guide/sdk-guide/cookbook/transcribe-to-captions.mdx", cases: [{ name: "primary", block: 0, file: ["captions.srt", "srt"] }, { name: "vtt", block: 0, file: ["captions.vtt", "srt"] }] }, { slug: "batch-transcribe-with-language-hint", mdx: "developer-guide/sdk-guide/cookbook/batch-transcribe-with-language-hint.mdx", cases: [{ name: "primary", block: 0 }] }, { slug: "telephony-8khz-audio", mdx: "developer-guide/sdk-guide/cookbook/telephony-8khz-audio.mdx", cases: [{ name: "primary", block: 0, file: ["out.wav", "wav"] }] }, { slug: "developer-guide/sdk-guide/cookbook/clone-and-wait-until-ready", mdx: "developer-guide/sdk-guide/cookbook/clone-and-wait-until-ready.mdx", cases: [{ name: "primary", block: 0, file: ["out.mp3", "mp3"] }] },