Skip to content

Add Telnyx streaming STT provider - #78

Merged
jason-shen merged 2 commits into
streamcoreai:mainfrom
a692570:telnyx-stt-provider
Sep 12, 2026
Merged

jason-shen merged 2 commits into
streamcoreai:mainfrom
a692570:telnyx-stt-provider

Conversation

@a692570

@a692570 a692570 commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Telnyx STT provider, shaped per #75

Adds stt.provider = "telnyx": streaming transcription over the Telnyx speech-to-text WebSocket, reusing the [telnyx] section and API key from #74. Implements option 2 from #75; the pipeline capability flag stays out per the scoping there, and #75's option 3 remains open as a separate follow-up discussion.

One endpoint, two socket lifecycles

telnyx.transcription_engine picks the recognizer and is sent to the endpoint verbatim: case-sensitive, no trimming, no case normalization. Two values are claimed:

  • Deepgram (default) streams interims, so it gets one long-lived socket per session, the same shape as deepgram.go: partials flow into onResult with IsFinal: false, and barge-in and live captions work out of the box. A finals-only default would silently switch both off for anyone who just sets the provider and starts talking.
  • Telnyx (in-house) emits exactly one final per utterance, only after audio stops, and holds the socket open (verified live: no server close within 20s; interim_results ignored). The client therefore endpoints the speech itself with internal/vad.Detector and opens one socket per utterance: 200ms onset, 600ms silence to close (the silence timeout openai.go uses for the same job), the finished utterance is uploaded at the fastest rate verified against the live endpoint (200ms of audio per write, 30ms pause), the single final is awaited, and the client closes the socket. One startup log line says barge-in and live captions are off for the session, so the operator learns it from the log, not from a caller talking over the agent.

Design note on "open on speech onset": I took the buffer-then-dial shape from openai.go rather than dialing against the first syllable. This engine emits no partials, so live streaming buys nothing, and a ~600ms dial racing the onset audio is a failure mode the buffered version does not have. Onset frames are preserved by a 1s idle lookback that becomes the utterance prefix. SendAudio never errors across utterance boundaries (an error would stop the pipeline inbound loop); a spent socket is simply replaced by the next utterance's dial.

Docs and config

  • config.toml.example: transcription_engine under [telnyx] with the two verified values; telnyx added to the [stt] list. TestConfigExampleDocumentsEveryField holds it.
  • docs/providers.md / docs/providers.zh-CN.md: Telnyx STT section, the Telnyx note placed directly under the whisper note, and the requested whisper-note honesty fix: finals-only means barge-in and live captions do not work, not just "no streaming partials".
  • README provider notes (EN and zh).

Tests (hermetic, fake WebSocket server)

Using the grok_conn_test.go pattern (gorilla Upgrader + httptest, with telnyxSTTURL overridable the same way grokDialURL is):

  • frame routing for both engines, is_final handling, confidence: null mapped to 0 (unknown), errors-array detail surfacing, malformed-frame rejection
  • session path: partial, final, then another partial on the same socket (the final does not close it)
  • utterance path: two utterances produce two dials, every frame of each utterance reaches the server, exactly one final each, the client closes each socket after its final, SendAudio keeps accepting audio across the boundary
  • the engine string reaches the URL unmodified (MixedCaseEngine dials with casing intact); an unset engine defaults to Deepgram; silence alone never dials

Real-audio smoke (per CONTRIBUTING)

Input: 13.5s of 16kHz mono speech ("The quick brown fox jumps over the lazy dog. Testing one two three, this is the Telnyx STT capture for StreamCore."), fed as 20ms frames at 1.5x realtime.

  • Deepgram engine: 7 interims then 2 per-sentence finals, one socket for the whole clip. "The quick brown fox jumps over the lazy dog." (conf 0.999); "Testing one, two, three. This is the Telnyx STT capture for StreamCore." (conf 0.998).
  • Telnyx engine: startup log line fires, exactly one final for the whole clip, confidence null: "The quick brown fox jumps over the lazy dog, testing 1-2-3, this is the Telnix STT capture for Streamcore." (the engine's genuine output on the synthetic voice, quoted as-is). Final at T+12s after a 9.3s feed (dial ~0.6s + paced upload + server end-of-speech detection); the client closed the socket.

Gates

gofmt -l . clean, go build ./..., go vet ./..., go test -race ./... all green. go.mod/go.sum untouched (no new dependencies; internal/vad was already in the graph).

@jason-shen
jason-shen merged commit 731cbeb into streamcoreai:main Sep 12, 2026
1 check passed
@jason-shen

Copy link
Copy Markdown
Member

thank you @a692570

@a692570

a692570 commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Per the promise on #74: here's the Telnyx STT addition for llms-full.txt, matching the same style.

1. The STT provider list line gains | telnyx:

provider = "deepgram"   # aliyun | assemblyai | deepgram | openai | telnyx | vibevoice | volcengine

2. Replace the [telnyx] block with the combined version (it serves both directions now, one key covers both):

[telnyx]                # TTS and STT. Telnyx streaming text-to-speech and speech-to-text
api_key = ""            # or TELNYX_API_KEY env var
voice = ""              # TTS; defaults to Telnyx.Qwen3TTS.d9348e0d-988a-42cc-a64e-18093fe45c03 ("Delta");
                        # any id from GET /v2/text-to-speech/voices, availability varies by account,
                        # a voice your key is not provisioned for fails the dial with HTTP 403
voice_speed = 1.0       # TTS; clamped to 0.8-1.2
transcription_engine = "Deepgram"  # STT. "Deepgram" streams partials, so barge-in and live
                        # captions work; "Telnyx" is the in-house engine, finals-only, so both are
                        # off and a startup log line says so. Case-sensitive, sent verbatim; other
                        # hosted engines the endpoint fronts pass through untested

That closes out the llms-full.txt lag for everything merged so far. Whenever you get a chance.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants