Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,7 @@ Providers: Deepgram, AssemblyAI, OpenAI, Cartesia, ElevenLabs, MiniMax, Speechif

OpenAI STT supports `whisper-1`, `gpt-4o-transcribe`, and `gpt-4o-mini-transcribe` through the independent `openai.stt_model` setting.

Telnyx STT fronts a dozen engines behind one key (`telnyx.transcription_engine`, default `Deepgram` so barge-in and live captions work out of the box); the in-house `Telnyx` engine is finals-only, so both are off when it is selected, and a startup log line says so.
Telnyx STT fronts a dozen engines behind one key (`telnyx.transcription_engine`, default `Deepgram` so barge-in and live captions work out of the box); the in-house `Telnyx` engine is finals-only, so live captions show finals only and barge-in waits out the full backchannel window on VAD alone, and a startup log line says so.

## Documentation

Expand Down
2 changes: 1 addition & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,7 @@ StreamCore 位于「提示词 + 工具」类框架的下一层:媒体链路。

OpenAI STT 可通过独立的 `openai.stt_model` 配置选择 `whisper-1`、`gpt-4o-transcribe` 或 `gpt-4o-mini-transcribe`。

Telnyx STT 用一个 key 前置十余种引擎(`telnyx.transcription_engine`,默认 `Deepgram`,打断与实时字幕开箱即用);自研 `Telnyx` 引擎只出最终结果,选中它时两者关闭,启动日志会写明。
Telnyx STT 用一个 key 前置十余种引擎(`telnyx.transcription_engine`,默认 `Deepgram`,打断与实时字幕开箱即用);自研 `Telnyx` 引擎只出最终结果,实时字幕只显示最终结果,打断降级为等满整个回应窗口后仅凭 VAD 触发,启动日志会写明。

## 文档

Expand Down
6 changes: 3 additions & 3 deletions config.toml.example
Original file line number Diff line number Diff line change
Expand Up @@ -221,9 +221,9 @@ voice = "Telnyx.Qwen3TTS.d9348e0d-988a-42cc-a64e-18093fe45c03" # Any catalog vo
voice_speed = 1.0 # Playback-rate multiplier, clamped to 0.8-1.2 like the per-utterance delivery tags
transcription_engine = "Deepgram" # STT. Transcription engine when stt.provider = "telnyx".
# Verified values: "Deepgram" (streams partials, so barge-in and live captions
# work) or "Telnyx" (in-house, finals-only, so barge-in and live captions are
# off and a startup log line says so). Case-sensitive, sent verbatim; other
# hosted engines the endpoint fronts pass through untested
# work) or "Telnyx" (in-house, finals-only: live captions show finals only and
# barge-in waits out the backchannel window on VAD alone; a startup log line
# says so). Case-sensitive, sent verbatim; other hosted engines pass through untested

[mimo]
api_key = "" # Required if tts.provider = "mimo"
Expand Down
4 changes: 2 additions & 2 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,8 +159,8 @@ api_key = ""
voice = "Telnyx.Qwen3TTS.d9348e0d-988a-42cc-a64e-18093fe45c03" # Any catalog voice from GET /v2/text-to-speech/voices; availability varies by account
voice_speed = 1.0 # Playback-rate multiplier, clamped to 0.8-1.2
transcription_engine = "Deepgram" # STT engine, verified values: "Deepgram" (partials, barge-in works; the default) or
# "Telnyx" (in-house, finals-only: barge-in and live captions are off, and a startup
# log line says so). Case-sensitive, sent verbatim
# "Telnyx" (in-house, finals-only: live captions show finals only and barge-in waits
# out the backchannel window on VAD alone; a startup log line says so). Case-sensitive, sent verbatim

[minimax]
api_key = ""
Expand Down
2 changes: 1 addition & 1 deletion docs/configuration.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ api_key = ""
voice = "Telnyx.Qwen3TTS.d9348e0d-988a-42cc-a64e-18093fe45c03" # GET /v2/text-to-speech/voices 目录中的任意音色;可用性因账号而异
voice_speed = 1.0 # 播放速率倍数,限制在 0.8-1.2
transcription_engine = "Deepgram" # STT 引擎,已验证取值:"Deepgram"(有中间结果,打断可用;默认值)或
# "Telnyx"(自研,只出最终结果:打断与实时字幕关闭,启动日志会写明)。大小写敏感,原样透传
# "Telnyx"(自研,只出最终结果:实时字幕只显示最终结果,打断降级为等满回应窗口后仅凭 VAD 触发;启动日志会写明)。大小写敏感,原样透传

[minimax]
api_key = ""
Expand Down
8 changes: 4 additions & 4 deletions docs/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@

Notes:

- `stt.provider = "openai"` uses batch final transcription instead of streaming partials, so barge-in and live captions do not work, since both depend on partials; choose `whisper-1`, `gpt-4o-transcribe`, or `gpt-4o-mini-transcribe` with `openai.stt_model`.
- `stt.provider = "telnyx"` fronts Telnyx's in-house and a dozen hosted transcription engines over one WebSocket and one key. `transcription_engine` defaults to `Deepgram`, so barge-in and live captions work out of the box; the in-house `Telnyx` engine is finals-only, so both are off when it is selected. See [Telnyx STT](#telnyx-stt).
- `stt.provider = "openai"` uses batch final transcription instead of streaming partials, so live captions show finals only and barge-in runs in a degraded mode: the interrupt fires on VAD alone once the full 600ms backchannel window has elapsed, since there is no text to classify a short burst with. Choose `whisper-1`, `gpt-4o-transcribe`, or `gpt-4o-mini-transcribe` with `openai.stt_model`.
- `stt.provider = "telnyx"` fronts Telnyx's in-house and a dozen hosted transcription engines over one WebSocket and one key. `transcription_engine` defaults to `Deepgram`, so barge-in and live captions work out of the box; the in-house `Telnyx` engine is finals-only, so live captions show finals only and barge-in waits out the backchannel window on VAD alone. See [Telnyx STT](#telnyx-stt).
- `llm.provider = "ollama"` targets any Ollama-compatible endpoint via `base_url` — local or on your own infrastructure.
- `llm.provider = "agent"` POSTs each turn to an HTTP endpoint you host; your agent owns memory, prompting, and tools, and replies stream back as SSE, chunked text, or JSON. See [Bring your own agent](./bring-your-own-agent.md).
- `stt.provider = "vibevoice"` and `tts.provider = "vibevoice"` use local models; start the Python sidecars first.
Expand Down Expand Up @@ -164,8 +164,8 @@ transcription_engine = "Deepgram" # verified: "Deepgram" (partial results, bar

Three things to know:

- **`Deepgram` is the default.** It streams interim results exactly like the built-in Deepgram provider, so barge-in and live captions work out of the box. A finals-only default would silently switch both off for anyone who just sets the provider and starts talking.
- **The in-house `Telnyx` engine is finals-only.** It emits exactly one final per utterance, only after the caller stops speaking: no interims, no timestamps, confidence `null`. Barge-in and live captions depend on interim results (the pipeline gates interruption on partial text), so **neither works with this engine**: interruptions never fire, and the client transcript shows nothing until the final lands. Because the engine answers only after audio stops, the client runs its own endpointing (`internal/vad`, the same detector the pipeline uses, with the silence timeout `openai.go` uses) and opens one socket per utterance, closing it once the final lands, since the server keeps it open. A startup log line says that barge-in and live captions are off for the session, so the operator learns it from the log rather than from a caller talking over the agent with nothing happening. See the [design discussion](https://github.com/streamcoreai/streamcore-server/issues/75) for the trade-offs.
- **`Deepgram` is the default.** It streams interim results exactly like the built-in Deepgram provider, so barge-in and live captions work out of the box. A finals-only default would silently degrade both for anyone who just sets the provider and starts talking: barge-in to VAD-only, live captions to finals only.
- **The in-house `Telnyx` engine is finals-only.** It emits exactly one final per utterance, only after the caller stops speaking: no interims, no timestamps, confidence `null`. Live captions therefore show nothing until the final lands. Barge-in cannot confirm an interruption from partial text, so it degrades instead of switching off: speech over the agent opens the suppression window on VAD alone, and the interrupt fires once the speech has lasted past the full 600ms backchannel window. A burst that ends inside the window is treated as backchannel — with no text, it cannot be told apart from "mm-hm" — so sustained interruptions work and quick ones do not. Because the engine answers only after audio stops, the client runs its own endpointing (`internal/vad`, the same detector the pipeline uses, with the silence timeout `openai.go` uses) and opens one socket per utterance, closing it once the final lands, since the server keeps it open. A startup log line says the session runs barge-in VAD-only on this engine, so the operator learns it from the log. See the [design discussion](https://github.com/streamcoreai/streamcore-server/issues/75) for the trade-offs.
- **The engine name is case-sensitive and sent verbatim.** `telnyx` is rejected with a structured error frame that lists the supported engines. Only `Deepgram` and `Telnyx` are verified here; the other hosted engines the endpoint fronts (AssemblyAI, Azure, and the rest of that list) pass through untested.

Confidence arrives as `null` from the in-house engine and a 0-1 float from hosted ones; the pipeline treats `null` as unknown rather than low.
Expand Down
8 changes: 4 additions & 4 deletions docs/providers.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@

注意:

- `stt.provider = "openai"` 使用批量最终转写而不是流式中间结果,打断(barge-in)与实时字幕都依赖中间结果,因此都不工作;可通过 `openai.stt_model` 选择 `whisper-1`、`gpt-4o-transcribe` 或 `gpt-4o-mini-transcribe`。
- `stt.provider = "telnyx"` 通过一条 WebSocket、一个 key 前置自研与十余种托管转写引擎。`transcription_engine` 默认 `Deepgram`,打断与实时字幕开箱即用;自研 `Telnyx` 引擎只出最终结果,选中它时两者关闭。见 [Telnyx STT](#telnyx-stt)。
- `stt.provider = "openai"` 使用批量最终转写而不是流式中间结果,因此实时字幕只显示最终结果,打断降级为等满 600ms 回应窗口后仅凭 VAD 触发(没有文本,无法对短促语做分类);可通过 `openai.stt_model` 选择 `whisper-1`、`gpt-4o-transcribe` 或 `gpt-4o-mini-transcribe`。
- `stt.provider = "telnyx"` 通过一条 WebSocket、一个 key 前置自研与十余种托管转写引擎。`transcription_engine` 默认 `Deepgram`,打断与实时字幕开箱即用;自研 `Telnyx` 引擎只出最终结果,实时字幕只显示最终结果,打断降级为等满回应窗口后仅凭 VAD 触发。见 [Telnyx STT](#telnyx-stt)。
- `llm.provider = "ollama"` 通过 `base_url` 指向任何兼容 Ollama 的端点 —— 本地或你自己的基础设施均可。
- `llm.provider = "agent"` 把每一轮对话 POST 到你托管的 HTTP 端点;记忆、提示词与工具都由你的智能体掌控,回复以 SSE、分块文本或 JSON 流式返回。见[接入你自己的智能体](./bring-your-own-agent.zh-CN.md)。
- `stt.provider = "vibevoice"` 与 `tts.provider = "vibevoice"` 使用本地模型;请先启动 Python 边车进程。
Expand Down Expand Up @@ -164,8 +164,8 @@ transcription_engine = "Deepgram" # 已验证取值:"Deepgram"(有中间

有三件事必须弄对:

- **默认是 `Deepgram`。** 它像内置的 Deepgram 服务商一样流式输出中间结果,打断(barge-in)与实时字幕开箱即用。一个只出最终结果的默认引擎,会让任何只改了 provider 就开始说话的人悄无声息地失去这两项能力。
- **自研 `Telnyx` 引擎只出最终结果。** 它在来电者停止说话后才发出唯一一帧 final —— 没有中间结果、没有时间戳、置信度为 `null`。打断与实时字幕依赖中间结果(流水线以部分文本来判定打断),因此**这两项在该引擎下不工作**:打断永远不会触发,客户端字幕也要等到 final 落地才有内容。因为该引擎只在整个话语结束后才应答,客户端自己做端点检测(`internal/vad`,即流水线在用的同一个检测器,静音窗口与 `openai.go` 相同),每个话语开一条新连接,final 落地后由客户端关闭(服务端会一直握着连接不放)。启动时日志里会写明本会话的打断与实时字幕已关闭,运维从日志就能知道,而不用等到来电者对着智能体说话却毫无反应。取舍讨论见[设计讨论](https://github.com/streamcoreai/streamcore-server/issues/75)。
- **默认是 `Deepgram`。** 它像内置的 Deepgram 服务商一样流式输出中间结果,打断(barge-in)与实时字幕开箱即用。一个只出最终结果的默认引擎,会让任何只改了 provider 就开始说话的人悄无声息地让这两项能力降级:打断退化为仅凭 VAD,实时字幕只剩最终结果。
- **自研 `Telnyx` 引擎只出最终结果。** 它在来电者停止说话后才发出唯一一帧 final —— 没有中间结果、没有时间戳、置信度为 `null`。客户端字幕要等到 final 落地才有内容。打断没有中间文本可判定,因此是降级而不是关闭:来电者压过智能体说话时,抑制窗口仅凭 VAD 打开,只有说话持续超过整个 600ms 回应窗口后才确认打断;窗口内结束的短促语一律按回应词处理 —— 没有文本就无法把它和 "嗯嗯" 区分开。持续抢话可以打断,短促抢话不能。因为该引擎只在整个话语结束后才应答,客户端自己做端点检测(`internal/vad`,即流水线在用的同一个检测器,静音窗口与 `openai.go` 相同),每个话语开一条新连接,final 落地后由客户端关闭(服务端会一直握着连接不放)。启动时日志里会写明本会话的打断以仅凭 VAD 的降级模式运行,运维从日志就能知道。取舍讨论见[设计讨论](https://github.com/streamcoreai/streamcore-server/issues/75)。
- **引擎名大小写敏感,原样透传。** `telnyx` 会被一帧结构化错误拒绝,错误里列出支持的引擎。此处只验证了 `Deepgram` 与 `Telnyx` 两个取值;该端点前置的其他托管引擎(AssemblyAI、Azure 及列表中的其余引擎)可透传但未经测试。

置信度:自研引擎返回 `null`,托管引擎返回 0-1 浮点数;流水线把 `null` 视为未知而不是低置信。
Expand Down
8 changes: 4 additions & 4 deletions internal/config/config.go
Original file line number Diff line number Diff line change
Expand Up @@ -425,10 +425,10 @@ type TelnyxConfig struct {
// "telnyx". The value is case-sensitive and sent to the endpoint
// verbatim. "Deepgram" (the default) streams partials, so barge-in and
// live captions work out of the box; the in-house "Telnyx" engine
// emits one final per utterance and no interims, so barge-in and live
// captions do not work with it and a startup log line says so. Other
// engines the endpoint fronts (AssemblyAI, Azure, ...) pass through
// untested.
// emits one final per utterance and no interims, so live captions show
// finals only and barge-in runs VAD-only after the backchannel window —
// a startup log line says so. Other engines the endpoint fronts
// (AssemblyAI, Azure, ...) pass through untested.
TranscriptionEngine string `toml:"transcription_engine"`
}

Expand Down
19 changes: 18 additions & 1 deletion internal/pipeline/inbound.go
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,15 @@ func (p *Pipeline) runInbound() {
}
defer sttClient.Close()

// Finals-only providers never confirm speech with partial text
// (issue #75). Asserted once here: a provider that does not implement
// stt.PartialsEmitter keeps emitsPartials true and the partials-driven
// path below exactly as it was.
emitsPartials := true
if ep, ok := sttClient.(stt.PartialsEmitter); ok {
emitsPartials = ep.EmitsPartials()
}

// Backchannel suppression state machine
var bargeInPending bool
var bargeInStart time.Time
Expand Down Expand Up @@ -172,6 +181,10 @@ func (p *Pipeline) runInbound() {
}
} else if !p.bargeInVAD.IsSpeaking() {
// Speech ended within the window — check for backchannel.
// A finals-only provider has no partial text here, so
// every burst that ends inside the window classifies as
// backchannel: with no text, there is no basis to cut
// the agent off mid-word.
partial, _ := latestPartial.Load("text")
partialStr, _ := partial.(string)
if !isMeaningfulBargeInTranscript(partialStr) {
Expand Down Expand Up @@ -200,8 +213,12 @@ func (p *Pipeline) runInbound() {
}
}
// else: still speaking within window, keep waiting
} else if p.bargeInVAD.IsSpeaking() && p.speaking.Load() && hasPartialText.Load() {
} else if p.bargeInVAD.IsSpeaking() && p.speaking.Load() && (!emitsPartials || hasPartialText.Load()) {
// Conditions met — start backchannel suppression window.
// A finals-only provider opens the window on VAD alone,
// since partial text never arrives to confirm the speech;
// the interrupt still cannot fire until the window has
// fully elapsed.
bargeInPending = true
bargeInStart = time.Now()
hasPartialText.Store(false)
Expand Down
Loading
Loading