Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,15 @@ node_modules/
.terraform/
plugins/plugins/*/.env
plugins/plugins/*/.env.*
plugins/plugins/*/.venv
plugins/plugins/*/token.json
external/vibeVoice/.venv
# Local server config (API keys); keep *example* files tracked
config.toml

# Python virtualenvs and bytecode (external sidecars, plugins)
.venv/
__pycache__/
*.pyc

# Go build / test artifacts
*.exe
*.exe~
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,7 @@ StreamCore starts one layer below prompt-and-tool frameworks: the media path. Yo

Details and code: [Bring your own agent](./docs/bring-your-own-agent.md) · [Agent runtime](./docs/agent-runtime.md).

Providers: Deepgram, AssemblyAI, OpenAI, Cartesia, ElevenLabs, MiniMax, Speechify, Telnyx, Ollama, VibeVoice (local), xAI Grok Voice (speech-to-speech), pgvector/Supabase for retrieval. See [Providers](./docs/providers.md).
Providers: Deepgram, AssemblyAI, OpenAI, Cartesia, ElevenLabs, MiniMax, Speechify, Telnyx, Ollama, Moonshine (local), VibeVoice (local), xAI Grok Voice (speech-to-speech), pgvector/Supabase for retrieval. See [Providers](./docs/providers.md).

OpenAI STT supports `whisper-1`, `gpt-4o-transcribe`, and `gpt-4o-mini-transcribe` through the independent `openai.stt_model` setting.

Expand All @@ -128,7 +128,7 @@ Telnyx STT fronts a dozen engines behind one key (`telnyx.transcription_engine`,
| [Bring your own agent](./docs/bring-your-own-agent.md) | Five ways to own the intelligence, including the HTTP agent endpoint and the `llm.Client` interface |
| [Agent runtime](./docs/agent-runtime.md) | Plugins, skills, RAG, document ingestion |
| [Developer agent](./docs/developer-agent.md) | Optional GitHub App and Codex integrations: CI investigation, isolated worktrees, confirmation-gated pull requests |
| [Providers](./docs/providers.md) | Grok speech-to-speech, MiniMax, local VibeVoice, per-provider caveats |
| [Providers](./docs/providers.md) | Grok speech-to-speech, MiniMax, local Moonshine and VibeVoice, per-provider caveats |
| [Configuration](./docs/configuration.md) | Full annotated `config.toml` reference |
| [Protocol](./docs/protocol.md) | WHIP signaling, DataChannel events, auth |
| [Architecture](./docs/architecture.md) | Media flow, why Go, package layout |
Expand Down
4 changes: 2 additions & 2 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,7 +113,7 @@ StreamCore 位于「提示词 + 工具」类框架的下一层:媒体链路。

详情与代码:[接入你自己的智能体](./docs/bring-your-own-agent.zh-CN.md) · [智能体运行时](./docs/agent-runtime.zh-CN.md)。

服务商:Deepgram、AssemblyAI、OpenAI、Cartesia、ElevenLabs、MiniMax、Speechify、Telnyx、Ollama、VibeVoice(本地)、xAI Grok Voice(语音到语音),检索支持 pgvector / Supabase。见[服务商](./docs/providers.zh-CN.md)。
服务商:Deepgram、AssemblyAI、OpenAI、Cartesia、ElevenLabs、MiniMax、Speechify、Telnyx、Ollama、Moonshine(本地)、VibeVoice(本地)、xAI Grok Voice(语音到语音),检索支持 pgvector / Supabase。见[服务商](./docs/providers.zh-CN.md)。

OpenAI STT 可通过独立的 `openai.stt_model` 配置选择 `whisper-1`、`gpt-4o-transcribe` 或 `gpt-4o-mini-transcribe`。

Expand All @@ -128,7 +128,7 @@ Telnyx STT 用一个 key 前置十余种引擎(`telnyx.transcription_engine`
| [接入你自己的智能体](./docs/bring-your-own-agent.zh-CN.md) | 掌控智能体的五种方式,含 HTTP 智能体端点与 `llm.Client` 接口 |
| [智能体运行时](./docs/agent-runtime.zh-CN.md) | 插件、技能、RAG、文档入库 |
| [开发者智能体](./docs/developer-agent.zh-CN.md) | 可选的 GitHub App 与 Codex 集成:CI 排查、隔离 worktree、需确认的 Pull Request |
| [服务商](./docs/providers.zh-CN.md) | Grok 语音到语音、MiniMax、本地 VibeVoice 及各服务商注意事项 |
| [服务商](./docs/providers.zh-CN.md) | Grok 语音到语音、MiniMax、本地 Moonshine 与 VibeVoice 及各服务商注意事项 |
| [配置](./docs/configuration.zh-CN.md) | 完整带注释的 `config.toml` 参考 |
| [协议](./docs/protocol.zh-CN.md) | WHIP 信令、DataChannel 事件、鉴权 |
| [架构](./docs/architecture.zh-CN.md) | 媒体流转、为什么用 Go、包结构 |
Expand Down
16 changes: 14 additions & 2 deletions config.toml.example
Original file line number Diff line number Diff line change
Expand Up @@ -110,13 +110,13 @@ turn_merge_ms = 350 # Debounce window for merging finals into one turn.
provider = "" # Supported: grok (empty = classic pipeline)

[stt]
provider = "deepgram" # Supported: aliyun, assemblyai, deepgram, openai, telnyx, vibevoice, volcengine
provider = "deepgram" # Supported: aliyun, assemblyai, deepgram, moonshine, openai, telnyx, vibevoice, volcengine

[llm]
provider = "openai" # Supported: openai, ollama, agent

[tts]
provider = "cartesia" # Supported: aliyun, cartesia, deepgram, elevenlabs, mimo, minimax, speechify, telnyx, vibevoice, volcengine
provider = "cartesia" # Supported: aliyun, cartesia, deepgram, elevenlabs, mimo, minimax, moonshine, speechify, telnyx, vibevoice, volcengine

# Provider credentials

Expand Down Expand Up @@ -248,6 +248,18 @@ asr_url = "ws://127.0.0.1:8200" # WebSocket URL for VibeVoice ASR server
tts_url = "http://127.0.0.1:8300" # HTTP URL for VibeVoice TTS server
voice = "en-Emma_woman" # TTS voice name

# Moonshine — local STT and TTS via external Python services, no API key.
# Start the servers first:
# python external/moonshine/moonshineStt/server.py (default port 8210)
# python external/moonshine/moonshineTts/server.py (default port 8310)
# Language and model selection are sidecar flags, not config keys: both are
# fixed when the model loads rather than chosen per request.
[moonshine]
stt_url = "ws://127.0.0.1:8210" # WebSocket URL for the Moonshine STT server
tts_url = "http://127.0.0.1:8310" # HTTP URL for the Moonshine TTS server
voice = "kokoro_af_heart" # TTS voice id. The prefix picks the vocoder:
# kokoro_, piper_, or zipvoice_

# Alibaba Cloud Model Studio (DashScope). One endpoint and one key serve both
# streaming ASR (stt.provider = "aliyun") and streaming TTS
# (tts.provider = "aliyun"). Get an API key at https://bailian.console.aliyun.com/
Expand Down
4 changes: 2 additions & 2 deletions docs/capabilities.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,9 +114,9 @@ StreamCore can run a complete speech-to-agent-to-speech pipeline, but that is on

| AI integration | Providers |
|----------------|-----------|
| Streaming STT | Deepgram, AssemblyAI, OpenAI, VibeVoice (local) |
| Streaming STT | Deepgram, AssemblyAI, OpenAI, Moonshine (local), VibeVoice (local) |
| LLM | OpenAI, Ollama (local or self-hosted), or your own HTTP agent endpoint (`agent`) |
| Streaming TTS | Cartesia, Deepgram, ElevenLabs, MiniMax, Speechify, VibeVoice (local) |
| Streaming TTS | Cartesia, Deepgram, ElevenLabs, MiniMax, Moonshine (local), Speechify, VibeVoice (local) |
| Speech-to-speech | xAI Grok Voice (replaces STT + LLM + TTS in one model) |
| Retrieval | pgvector, Supabase |
| Custom tools | Python / TypeScript / JavaScript plugins, native Go tools |
Expand Down
9 changes: 7 additions & 2 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,13 +73,13 @@ turn_merge_ms = 350 # Debounce window for merging finals into o
provider = "" # "grok", or empty for the classic pipeline

[stt]
provider = "deepgram" # aliyun | assemblyai | deepgram | openai | telnyx | vibevoice | volcengine
provider = "deepgram" # aliyun | assemblyai | deepgram | moonshine | openai | telnyx | vibevoice | volcengine

[llm]
provider = "openai" # openai | ollama | agent

[tts]
provider = "cartesia" # cartesia | deepgram | elevenlabs | mimo | minimax | speechify | telnyx | vibevoice
provider = "cartesia" # cartesia | deepgram | elevenlabs | mimo | minimax | moonshine | speechify | telnyx | vibevoice

# [grok] # Used when realtime.provider = "grok"
# api_key = ""
Expand Down Expand Up @@ -173,6 +173,11 @@ asr_url = "ws://127.0.0.1:8200"
tts_url = "http://127.0.0.1:8300"
voice = "en-Emma_woman"

[moonshine]
stt_url = "ws://127.0.0.1:8210"
tts_url = "http://127.0.0.1:8310"
voice = "kokoro_af_heart" # Prefix picks the vocoder: kokoro_, piper_, or zipvoice_

# RAG is optional — omit the [rag] section to disable it entirely.
# [rag]
# provider = "supabase" # "pgvector" or "supabase"
Expand Down
9 changes: 7 additions & 2 deletions docs/configuration.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,13 +59,13 @@ turn_merge_ms = 350 # Debounce window for merging finals into o
provider = "" # "grok", or empty for the classic pipeline

[stt]
provider = "deepgram" # aliyun | assemblyai | deepgram | openai | telnyx | vibevoice | volcengine
provider = "deepgram" # aliyun | assemblyai | deepgram | moonshine | openai | telnyx | vibevoice | volcengine

[llm]
provider = "openai" # openai | ollama | agent

[tts]
provider = "cartesia" # cartesia | deepgram | elevenlabs | mimo | minimax | speechify | telnyx | vibevoice
provider = "cartesia" # cartesia | deepgram | elevenlabs | mimo | minimax | moonshine | speechify | telnyx | vibevoice

# [grok] # Used when realtime.provider = "grok"
# api_key = ""
Expand Down Expand Up @@ -158,6 +158,11 @@ asr_url = "ws://127.0.0.1:8200"
tts_url = "http://127.0.0.1:8300"
voice = "en-Emma_woman"

[moonshine]
stt_url = "ws://127.0.0.1:8210"
tts_url = "http://127.0.0.1:8310"
voice = "kokoro_af_heart" # Prefix picks the vocoder: kokoro_, piper_, or zipvoice_

# RAG is optional — omit the [rag] section to disable it entirely.
# [rag]
# provider = "supabase" # "pgvector" or "supabase"
Expand Down
40 changes: 36 additions & 4 deletions docs/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@

| Role | Providers | Required credentials |
|------|-----------|----------------------|
| STT | `aliyun`, `assemblyai`, `deepgram`, `openai`, `telnyx`, `vibevoice`, `volcengine` | Matching provider API key, or a local VibeVoice ASR server |
| STT | `aliyun`, `assemblyai`, `deepgram`, `moonshine`, `openai`, `telnyx`, `vibevoice`, `volcengine` | Matching provider API key, or a local Moonshine / VibeVoice ASR server |
| LLM | `openai`, `ollama`, `agent` | OpenAI API key, an Ollama instance you control, or your own HTTP agent endpoint |
| TTS | `cartesia`, `deepgram`, `elevenlabs`, `mimo`, `minimax`, `speechify`, `telnyx`, `vibevoice` | Matching provider API key, or a local VibeVoice TTS server |
| TTS | `cartesia`, `deepgram`, `elevenlabs`, `mimo`, `minimax`, `moonshine`, `speechify`, `telnyx`, `vibevoice` | Matching provider API key, or a local Moonshine / VibeVoice TTS server |
| Speech-to-speech | `grok` | xAI API key — replaces STT, LLM, and TTS together |
| RAG (optional) | `pgvector`, `supabase` | Postgres connection string or Supabase URL + key, plus an OpenAI key for embeddings |

Expand All @@ -17,6 +17,7 @@ Notes:
- `llm.provider = "ollama"` targets any Ollama-compatible endpoint via `base_url` — local or on your own infrastructure.
- `llm.provider = "agent"` POSTs each turn to an HTTP endpoint you host; your agent owns memory, prompting, and tools, and replies stream back as SSE, chunked text, or JSON. See [Bring your own agent](./bring-your-own-agent.md).
- `stt.provider = "vibevoice"` and `tts.provider = "vibevoice"` use local models; start the Python sidecars first.
- `stt.provider = "moonshine"` and `tts.provider = "moonshine"` are the other fully local pair, also behind Python sidecars. The STT sidecar answers in Deepgram's wire format, so it runs through the same transcript handling as Deepgram itself. See [Local Moonshine setup](#local-moonshine-setup).
- `tts.provider = "minimax"` covers 40+ languages and is the strongest option for Mandarin. See [MiniMax TTS](#minimax-tts) for the region and model-plan caveats.
- `tts.provider = "telnyx"` is Telnyx hosted synthesis over a per-utterance WebSocket; voice availability varies by account. See [Telnyx TTS](#telnyx-tts) for the connection model and the voice catalog.
- `tts.provider = "mimo"` is Xiaomi's MiMo TTS, with Chinese and English voices and optional voice cloning on the paid models.
Expand Down Expand Up @@ -176,9 +177,9 @@ VibeVoice provides fully local STT and TTS with no API keys, using [VibeVoice-AS

```bash
# Apple Silicon (MLX)
pip install mlx-audio numpy websockets fastapi uvicorn
pip install mlx-audio numpy websockets onnxruntime requests fastapi uvicorn
# OR PyTorch (Linux / CUDA)
pip install torch transformers librosa numpy websockets fastapi uvicorn
pip install torch "transformers>=5.3.0" accelerate librosa numpy websockets onnxruntime requests fastapi uvicorn

python external/vibeVoice/vibeVoiceAsr/server.py # ws://127.0.0.1:8200
python external/vibeVoice/vibeVoiceTTS/server.py # http://127.0.0.1:8300
Expand All @@ -198,3 +199,34 @@ voice = "en-Emma_woman"
```

The ASR server accepts live PCM over WebSocket and emits JSON transcript events. The TTS server accepts HTTP POST and returns raw PCM.

## Local Moonshine setup

[Moonshine](https://moonshine.ai) is the other fully local pair: streaming STT and TTS from one pip package, no API key and no account. Models are downloaded on first run and cached. English STT models are MIT; other languages load under the non-commercial [Moonshine Community License](https://www.moonshine.ai/license).

```bash
pip install -r external/moonshine/moonshineStt/requirements.txt
pip install -r external/moonshine/moonshineTts/requirements.txt

python external/moonshine/moonshineStt/server.py # ws://127.0.0.1:8210
python external/moonshine/moonshineTts/server.py # http://127.0.0.1:8310
```

```toml
[stt]
provider = "moonshine"

[tts]
provider = "moonshine"

[moonshine]
stt_url = "ws://127.0.0.1:8210"
tts_url = "http://127.0.0.1:8310"
voice = "kokoro_af_heart"
```

Language and model selection are sidecar flags rather than config keys, because both are fixed when the model loads rather than chosen per request: `--language`, `--model-arch` for STT, `--language` and `--voice` for TTS. English resolves to `medium-streaming` by default; `--model-arch tiny-streaming` trades accuracy for a much smaller footprint. The first request for a TTS voice downloads it, so `--preload` is worth setting for anything but a first try.

Both sidecars speak Deepgram's wire format. The STT server sends `Results`, `SpeechStarted` and `UtteranceEnd` frames, which the server decodes into Deepgram's own types and routes through the same accumulator — so overlapping finals are merged and an immediate repeat is suppressed exactly as they are for Deepgram. It reports no confidence, and the frames leave the field out rather than inventing one; the pipeline reads the absent value as unknown, not as low. The TTS server mirrors `/v1/speak`, writing raw headerless PCM as it synthesizes, so playback starts on the first clause.

Measured on an M-series Mac: first TTS chunk at ~130ms with synthesis running about 9x faster than playback.
40 changes: 36 additions & 4 deletions docs/providers.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@

| 角色 | 服务商 | 所需凭据 |
|------|-----------|----------------------|
| STT | `aliyun`、`assemblyai`、`deepgram`、`openai`、`telnyx`、`vibevoice`、`volcengine` | 对应服务商的 API key,或一个本地 VibeVoice ASR 服务 |
| STT | `aliyun`、`assemblyai`、`deepgram`、`moonshine`、`openai`、`telnyx`、`vibevoice`、`volcengine` | 对应服务商的 API key,或一个本地 Moonshine / VibeVoice ASR 服务 |
| LLM | `openai`、`ollama`、`agent` | OpenAI API key、你自己掌控的 Ollama 实例,或你自己的 HTTP 智能体端点 |
| TTS | `cartesia`、`deepgram`、`elevenlabs`、`mimo`、`minimax`、`speechify`、`telnyx`、`vibevoice` | 对应服务商的 API key,或一个本地 VibeVoice TTS 服务 |
| TTS | `cartesia`、`deepgram`、`elevenlabs`、`mimo`、`minimax`、`moonshine`、`speechify`、`telnyx`、`vibevoice` | 对应服务商的 API key,或一个本地 Moonshine / VibeVoice TTS 服务 |
| 语音到语音 | `grok` | xAI API key —— 一并取代 STT、LLM 与 TTS |
| RAG(可选) | `pgvector`、`supabase` | Postgres 连接串或 Supabase URL + key,另需 OpenAI key 用于 embedding |

Expand All @@ -17,6 +17,7 @@
- `llm.provider = "ollama"` 通过 `base_url` 指向任何兼容 Ollama 的端点 —— 本地或你自己的基础设施均可。
- `llm.provider = "agent"` 把每一轮对话 POST 到你托管的 HTTP 端点;记忆、提示词与工具都由你的智能体掌控,回复以 SSE、分块文本或 JSON 流式返回。见[接入你自己的智能体](./bring-your-own-agent.zh-CN.md)。
- `stt.provider = "vibevoice"` 与 `tts.provider = "vibevoice"` 使用本地模型;请先启动 Python 边车进程。
- `stt.provider = "moonshine"` 与 `tts.provider = "moonshine"` 是另一组完全本地的方案,同样基于 Python 边车进程。STT 边车按 Deepgram 的线格式应答,因此走的是与 Deepgram 相同的转写处理路径。见[本地 Moonshine 配置](#本地-moonshine-配置)。
- `tts.provider = "minimax"` 覆盖 40+ 语言,是中文场景下最强的选项。区域与套餐相关的坑见 [MiniMax TTS](#minimax-tts)。
- `tts.provider = "telnyx"` 是 Telnyx 托管合成,每个话语一条 WebSocket 连接;音色可用性因账号而异。连接模型与音色目录见 [Telnyx TTS](#telnyx-tts)。
- `tts.provider = "mimo"` 是小米 MiMo TTS,中英文音色齐备,付费模型还支持声音克隆。
Expand Down Expand Up @@ -176,9 +177,9 @@ VibeVoice 提供完全本地、无需 API key 的 STT 与 TTS:识别用 [VibeV

```bash
# Apple Silicon (MLX)
pip install mlx-audio numpy websockets fastapi uvicorn
pip install mlx-audio numpy websockets onnxruntime requests fastapi uvicorn
# 或 PyTorch(Linux / CUDA)
pip install torch transformers librosa numpy websockets fastapi uvicorn
pip install torch "transformers>=5.3.0" accelerate librosa numpy websockets onnxruntime requests fastapi uvicorn

python external/vibeVoice/vibeVoiceAsr/server.py # ws://127.0.0.1:8200
python external/vibeVoice/vibeVoiceTTS/server.py # http://127.0.0.1:8300
Expand All @@ -198,3 +199,34 @@ voice = "en-Emma_woman"
```

ASR 服务通过 WebSocket 接收实时 PCM 并输出 JSON 转写事件。TTS 服务接收 HTTP POST 并返回裸 PCM。

## 本地 Moonshine 配置

[Moonshine](https://moonshine.ai) 是另一组完全本地的方案:流式 STT 与 TTS 来自同一个 pip 包,无需 API key,也无需账号。模型首次运行时下载并缓存。英语 STT 模型为 MIT 许可,其他语言使用非商业的 [Moonshine Community License](https://www.moonshine.ai/license)。

```bash
pip install -r external/moonshine/moonshineStt/requirements.txt
pip install -r external/moonshine/moonshineTts/requirements.txt

python external/moonshine/moonshineStt/server.py # ws://127.0.0.1:8210
python external/moonshine/moonshineTts/server.py # http://127.0.0.1:8310
```

```toml
[stt]
provider = "moonshine"

[tts]
provider = "moonshine"

[moonshine]
stt_url = "ws://127.0.0.1:8210"
tts_url = "http://127.0.0.1:8310"
voice = "kokoro_af_heart"
```

语言与模型的选择放在边车进程的命令行参数里,而不是配置项里,因为两者都在模型加载时就固定下来,并非按请求选择:STT 用 `--language`、`--model-arch`,TTS 用 `--language`、`--voice`。英语默认解析到 `medium-streaming`;`--model-arch tiny-streaming` 以准确率换取小得多的占用。某个 TTS 音色的首次请求会触发下载,因此除了初次尝试,建议加上 `--preload`。

两个边车都使用 Deepgram 的线格式。STT 服务发送 `Results`、`SpeechStarted` 与 `UtteranceEnd` 帧,服务端将其解析为 Deepgram 自己的类型并交给同一个累加器 —— 重叠的 final 会被合并,紧随其后的重复会被抑制,与 Deepgram 完全一致。它不提供 confidence,帧里也就不带这个字段,而不是编造一个;管线把缺失值理解为「未知」而非「低置信度」。TTS 服务对齐 `/v1/speak`,边合成边写出无头 PCM,因此播放可以从第一个短句开始。

在 M 系列 Mac 上实测:首个 TTS 分块约 130 ms,合成速度约为播放速度的 9 倍。
Loading
Loading