Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
161 changes: 143 additions & 18 deletions .mintlify/skills/fish-audio-api/SKILL.md

Large diffs are not rendered by default.

10 changes: 6 additions & 4 deletions .mintlify/skills/fish-audio-sdk/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,8 @@ If the user wants raw `curl` / HTTP / WebSocket without installing an SDK, use t

- **Auth:** both SDKs read the API key from the `FISH_API_KEY` environment variable automatically. Get keys at `https://fish.audio/app/api-keys`. Never hardcode a key.
- **Base URL:** `https://api.fish.audio` (override with `base_url=` in Python / `baseUrl:` in JS).
- **Models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument.
- **TTS models:** the API supports `s1`, `s2-pro`, `s2.1-pro` (recommended for production), and `s2.1-pro-free` (free tier), but the SDK type definitions currently list only `s1` and `s2-pro` (`s2-pro` = SDK default). Both SDKs forward the model value without runtime validation, so `"s2.1-pro"` works over the wire. Static type checkers will flag it, so add `# type: ignore` (Python) / an `as` cast (TS), or use the `fish-audio-api` skill for raw calls. `speech-1.5` / `speech-1.6` are **deprecated**. In Python pass `model="s2-pro"` (keyword); in JS pass the **positional** `backend` argument.
- **ASR models:** use `transcribe-1-pro` (recommended: speaker turns, long recordings, emotion cues). Neither SDK has an ASR `model` argument: send the `model: transcribe-1-pro` HTTP header on every request (Python `RequestOptions(additional_headers=...)`, JS `requestOptions.headers`). A request without it is served and billed as `transcribe-1`, the model for short recordings. See [references/speech-to-text.md](references/speech-to-text.md).
- **Audio formats:** `mp3` (default), `wav`, `pcm`, `opus`.
- **Playback in examples:** `play()` shells out to a system audio tool: Python uses **ffmpeg/ffplay** (or `mpv`), JS uses **ffplay**. It is for local/desktop use; in a server, `save()` to a file or stream the bytes instead. See [references/installation.md](references/installation.md).

Expand Down Expand Up @@ -79,7 +80,7 @@ const audio = await client.textToSpeech.convert({ text: "Hi" }, "s1");
| Install, auth, playback deps, verify a key | [references/installation.md](references/installation.md) |
| Text-to-Speech (convert, stream, formats, prosody, model select) | [references/text-to-speech.md](references/text-to-speech.md) |
| Voice cloning (instant references + persistent voice models) | [references/voice-cloning.md](references/voice-cloning.md) |
| Speech-to-Text (transcribe, segments, timestamps) | [references/speech-to-text.md](references/speech-to-text.md) |
| Speech-to-Text (models, timestamps, fields the SDK drops) | [references/speech-to-text.md](references/speech-to-text.md) |
| Realtime WebSocket TTS (stream text → audio) | [references/websocket.md](references/websocket.md) |
| Errors, retries, and timeouts (the **real** exception types) | [references/errors.md](references/errors.md) |

Expand All @@ -101,6 +102,7 @@ The two SDKs do **not** use the same names. Use this map when porting code betwe
| Credit balance | `client.account.get_credits()` | `client.user.get_api_credit()` |
| Subscription package | `client.account.get_package()` | `client.user.get_package()` |
| Choose model | `model="s2-pro"` keyword arg | positional `backend` arg, e.g. `convert(req, "s2-pro")` |
| ASR model (header) | `RequestOptions(additional_headers=...)` | `convert(req, { headers: { model: "transcribe-1-pro" } })` |

## Decision shortcuts

Expand All @@ -109,11 +111,11 @@ The two SDKs do **not** use the same names. Use this map when porting code betwe
- **Clone a voice instantly from a clip** → pass `references=[ReferenceAudio(audio=..., text=...)]` (Python) / `references: [{ audio, text }]` (JS). See [voice-cloning](references/voice-cloning.md).
- **Persistent custom voice to reuse** → create a voice model, then use its `id` as `reference_id`.
- **Stream tokens from an LLM and play speech as it arrives** → `tts.stream_websocket` (Python) / `textToSpeech.convertRealtime` (JS). See [websocket](references/websocket.md).
- **Transcribe audio** → `asr.transcribe` (Python) / `speechToText.convert` (JS).
- **Transcribe audio** → `asr.transcribe` (Python) / `speechToText.convert` (JS) with the `model: transcribe-1-pro` header (recommended; without it, the request runs on `transcribe-1`). Pro request fields, `speaker_turns`, `request_id`, and the language fields need raw HTTP in Python (`fish-audio-sdk` 1.3.0). See [speech-to-text](references/speech-to-text.md).

## Gotchas (verified against the SDK source)

- Python `latency` accepts only **`"normal"` or `"balanced"`** (default `"balanced"`); there is no `"low"`.
- The Python client has **no `max_retries`** and does **not** auto-retry; the JS client **does** auto-retry (configurable via per-call `requestOptions.maxRetries`). See [errors](references/errors.md).
- Python defines a `ValidationError` class but **never raises it**, so don't catch it expecting validation failures; a 422 surfaces as `APIError`. The JS SDK throws `UnprocessableEntityError` on 422.
- ASR segment `start` / `end` are in **seconds**, but `duration` is in **milliseconds**. See [speech-to-text](references/speech-to-text.md).
- ASR segment `start` / `end` and `duration` are all in **seconds**. The Python SDK's `ASRResponse` docstring says milliseconds; that is wrong. See [speech-to-text](references/speech-to-text.md).
165 changes: 145 additions & 20 deletions .mintlify/skills/fish-audio-sdk/references/speech-to-text.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,35 @@
# Speech-to-Text (ASR)

Both SDKs wrap `POST /v1/asr`: one audio file per request, and the response arrives when the whole file has been transcribed. Use `transcribe-1-pro`, the recommended model. Neither SDK has a `model` argument for ASR, so select it with the `model` HTTP header on every request. Full guide: `https://docs.fish.audio/features/speech-to-text`.

## Choose a model

- `transcribe-1-pro` (recommended): recordings up to 60 minutes, including multi-speaker conversations. `text` contains inline `<|speaker:N|>` markers and bracketed emotion or vocal-event cues such as `[laughter]` or `[高兴]`, and the API returns structured `speaker_turns` when you request timestamps.
- `transcribe-1`: general transcription of short recordings. It serves every request whose `model` header is missing or not an exact match.

Select `transcribe-1-pro` with the `model` header:

- Python: `request_options=RequestOptions(additional_headers={"model": "transcribe-1-pro"})`, with `from fishaudio.core import RequestOptions`.
- JavaScript: pass `{ headers: { model: "transcribe-1-pro" } }` as the second argument of `convert`.
- Write the value exactly, in lowercase. A missing or unrecognized value (for example `Transcribe-1-Pro` or `transcribe-1pro`) is served and billed as `transcribe-1`, and no error is returned. If you expected `transcribe-1-pro` but the transcript has no speaker markers, check the header.

## Python: `client.asr.transcribe`

```python
from fishaudio import FishAudio
from fishaudio.core import RequestOptions

client = FishAudio()
client = FishAudio() # reads FISH_API_KEY

with open("audio.wav", "rb") as f:
result = client.asr.transcribe(audio=f.read(), language="en")
result = client.asr.transcribe(
audio=f.read(),
language="en", # optional hint; omit to auto-detect
request_options=RequestOptions(
additional_headers={"model": "transcribe-1-pro"}, # without this header: transcribe-1
timeout=900, # long Pro recordings can take several minutes
),
)

print(result.text)

Expand All @@ -18,26 +39,88 @@ for seg in result.segments:

Keyword params:

| Param | Type | Default | Notes |
| -------------------- | ------------------------ | ------------ | ----------------------------------------------------------------------------------------------------------------------- |
| `audio` | `bytes` | — (required) | Raw audio bytes. |
| `language` | `str` | auto-detect | Omit to auto-detect (e.g. `"en"`, `"zh"`, `"ja"`). |
| `include_timestamps` | `bool` | `True` | `False` omits per-segment timestamps (and `segments` is empty). Computing timestamps adds latency on clips under ~30 s. |
| `request_options` | `RequestOptions \| None` | `None` | Per-request timeout / headers. |
| Param | Type | Default | Notes |
| -------------------- | ------------------------ | ------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `audio` | `bytes` | — (required) | Raw file bytes; the format is detected from the content. The SDK sends them as MessagePack. |
| `language` | `str` | auto-detect | Optional hint, a lowercase ISO 639-1 code (`"en"`, `"zh"`, `"ja"`). Detection still runs, and the hint does not force the transcript language. `"en-US"` or `"English"` may be rejected with 400. |
| `include_timestamps` | `bool` | `True` | `True` (default) requests word-level timestamps, which adds processing time; pass `False` when you only need text (`segments` is then empty). |
| `request_options` | `RequestOptions \| None` | `None` | Per-request `timeout` and headers. Select the model here: `RequestOptions(additional_headers={"model": "transcribe-1-pro"})`. |

### Response shape (`ASRResponse`)

```python
result.text # str: full transcript
result.duration # float: total audio duration in MILLISECONDS
result.segments # list[ASRSegment]
result.text # str: full transcript (Pro: with <|speaker:N|> markers and [cues])
result.duration # float: total audio duration in seconds, including silence
result.segments # list[ASRSegment]; [] when timestamps are off or no speech was found
# each segment:
seg.text # str
seg.text # str: usually one word (one or a few characters in Chinese/Japanese)
seg.start # float: seconds
seg.end # float: seconds
seg.end # float: seconds (can equal start)
```

`duration`, `start`, and `end` are all in **seconds**. The SDK's `ASRResponse` docstring says milliseconds; that is wrong, so do not divide by 1000.

Segment text has no punctuation, speaker markers, or cues, and can be normalized (for example `35` for `3.5`), so it does not always match `text` character for character. Segments are not speaker turns.

### Fields the Python SDK does not expose

In `fish-audio-sdk` 1.3.0, `ASRResponse` keeps only `text`, `duration`, and `segments`. It silently drops these response fields:

- `speaker_turns` (`transcribe-1-pro`, when timestamps are requested);
- `request_id` (`transcribe-1-pro`), also sent as the `x-request-id` header;
- `language` (English name, such as `English`) and `language_code` (ISO 639-1, such as `en`), returned when the language is known.

`asr.transcribe()` also cannot send the `transcribe-1-pro` fields `tag_audio_events`, `diarize`, `num_speakers`, `min_speakers`, or `max_speakers`. For any of these, call the API directly. Field semantics are in the `fish-audio-api` skill (`POST /v1/asr`).

```python
import os
import httpx

with open("meeting.mp3", "rb") as f:
r = httpx.post(
"https://api.fish.audio/v1/asr",
headers={
"Authorization": f"Bearer {os.environ['FISH_API_KEY']}",
"model": "transcribe-1-pro",
},
files={"audio": f},
data={"ignore_timestamps": "false"}, # word timestamps + speaker_turns
# Long Pro recordings can take several minutes; httpx defaults to 5 s.
timeout=httpx.Timeout(900.0, connect=10.0),
)
r.raise_for_status()
result = r.json()
print(result.get("language_code"), result.get("request_id"))
for turn in result.get("speaker_turns", []):
print(f"{turn['speaker']} [{turn['start']:.2f}-{turn['end']:.2f}] {turn['text']}")
```

### Errors (Python)

`asr.transcribe()` raises `APIError` (`RateLimitError` for 429, `ServerError` for 5xx) with `.status`, `.message`, and `.body` (the raw response text). On `transcribe-1-pro`, error bodies also carry `code` and `request_id`:

```python
import json
from fishaudio.core import RequestOptions
from fishaudio.exceptions import APIError

try:
result = client.asr.transcribe(
audio=audio_bytes,
request_options=RequestOptions(
additional_headers={"model": "transcribe-1-pro"}
),
)
except APIError as e:
try:
err = json.loads(e.body or "{}")
except ValueError:
err = {} # errors from the network edge may not be JSON
print(e.status, err.get("code"), err.get("request_id"))
raise
```

> **Unit gotcha (verified in source):** segment `start` / `end` are in **seconds**, but `duration` is in **milliseconds**. Don't assume they share a unit.
Branch on `code` or the HTTP status, never on `message`. Retry 429 and 5xx with exponential backoff (the Python SDK does not retry); do not retry other 4xx responses. See the `fish-audio-api` skill for the status and `code` list. If you use `transcribe-1`, rely only on `status` and `message`.

## JavaScript: `client.speechToText.convert`

Expand All @@ -48,16 +131,58 @@ import { readFile } from "node:fs/promises";
const client = new FishAudioClient();

const buf = await readFile("audio.wav");
const result = await client.speechToText.convert({
audio: new File([buf], "audio.wav"),
language: "en", // optional; omit to auto-detect
ignore_timestamps: false, // false → include per-segment timestamps
});
const result = await client.speechToText.convert(
{
audio: new File([buf], "audio.wav"),
language: "en", // optional hint; omit to auto-detect
ignore_timestamps: false, // false → word-level segments (and speaker_turns on Pro)
},
{ headers: { model: "transcribe-1-pro" }, timeoutInSeconds: 900 }
);

console.log(result.text);
for (const seg of result.segments) {
console.log(`[${seg.start}-${seg.end}] ${seg.text}`);
}
```

`STTRequest` = `{ audio: File; language?: string; ignore_timestamps?: boolean }`. Note JS uses `ignore_timestamps` (the inverse of Python's `include_timestamps`). `STTResponse` mirrors the Python shape: `{ text, duration, segments }`.
`STTRequest` = `{ audio: File; language?: string; ignore_timestamps?: boolean }`, sent as multipart. JS uses `ignore_timestamps` (the inverse of Python's `include_timestamps`), and the server default is `true`, so you get no `segments` unless you pass `false`.

In Node.js, the built-in `fetch` that the SDK uses stops waiting for response headers after 300 s, whatever `timeoutInSeconds` says (the call fails with `fetch failed`). For long Pro recordings, raise that limit once at startup with the `undici` package (`undici@7` runs on Node.js 20.18.1+; `undici@8` needs Node.js 22.19+ and fails at import on older versions):

```ts
import { Agent, setGlobalDispatcher } from "undici"; // npm install undici@7

setGlobalDispatcher(
new Agent({ headersTimeout: 900_000, bodyTimeout: 900_000 })
);
```

`STTResponse` types only `{ text, duration, segments }`; at runtime the body also has `language`, `language_code`, and on Pro `request_id` and `speaker_turns`. Widen the type to read them:

```ts
type SpeakerTurn = {
speaker: string;
text: string;
start: number;
end: number;
};
const body = result as typeof result & {
language?: string;
language_code?: string;
request_id?: string;
speaker_turns?: SpeakerTurn[];
};
for (const turn of body.speaker_turns ?? []) {
console.log(`${turn.speaker} [${turn.start}-${turn.end}] ${turn.text}`);
}
```

The JS SDK cannot send `tag_audio_events`, `diarize`, or the speaker counts; use `fetch` with `FormData` for those (see the `fish-audio-api` skill).

## Limits and formats (both SDKs)

- `transcribe-1-pro` accepts recordings up to 60 minutes (longer returns 400 `audio_too_long`). Send long recordings as compressed audio (MP3, Opus, or AAC); a request that is too large returns 413. Send a whole conversation as one file: speaker labels are consistent within one response, not across requests.
- Processing time grows with the length of the recording, and long `transcribe-1-pro` requests can take several minutes. Raise the SDK timeout for long recordings (the examples use 900 s; the defaults are in [errors](errors.md)), and in Node.js also raise the `fetch` header limit shown above.
- `transcribe-1-pro` accepts WAV, MP3, AAC (including M4A/MP4), FLAC, Ogg (Opus or Vorbis), WebM/Matroska, and MOV, including browser recordings, and uses the first audio track of a video file. AIFF, CAF, WMA, AMR, AC-3, and raw (headerless) PCM return 400.
- If you use `transcribe-1`: it is designed for short recordings, up to 50 MiB per request; keep MP3 and Opus files under 25 MiB. For recordings longer than a few minutes, use `transcribe-1-pro`. It accepts WAV, MP3, AAC (including M4A/MP4), FLAC, and Ogg (Opus or Vorbis); convert browser WebM recordings to Ogg/Opus, MP3, or WAV first, or use `transcribe-1-pro`.
Loading
Loading