Add Telnyx streaming STT provider - #78
Merged
Merged
Conversation
…, per-utterance lifecycle
Member
|
thank you @a692570 |
Contributor
Author
|
Per the promise on #74: here's the Telnyx STT addition for llms-full.txt, matching the same style. 1. The STT provider list line gains provider = "deepgram" # aliyun | assemblyai | deepgram | openai | telnyx | vibevoice | volcengine2. Replace the [telnyx] # TTS and STT. Telnyx streaming text-to-speech and speech-to-text
api_key = "" # or TELNYX_API_KEY env var
voice = "" # TTS; defaults to Telnyx.Qwen3TTS.d9348e0d-988a-42cc-a64e-18093fe45c03 ("Delta");
# any id from GET /v2/text-to-speech/voices, availability varies by account,
# a voice your key is not provisioned for fails the dial with HTTP 403
voice_speed = 1.0 # TTS; clamped to 0.8-1.2
transcription_engine = "Deepgram" # STT. "Deepgram" streams partials, so barge-in and live
# captions work; "Telnyx" is the in-house engine, finals-only, so both are
# off and a startup log line says so. Case-sensitive, sent verbatim; other
# hosted engines the endpoint fronts pass through untestedThat closes out the llms-full.txt lag for everything merged so far. Whenever you get a chance. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Telnyx STT provider, shaped per #75
Adds
stt.provider = "telnyx": streaming transcription over the Telnyx speech-to-text WebSocket, reusing the[telnyx]section and API key from #74. Implements option 2 from #75; the pipeline capability flag stays out per the scoping there, and #75's option 3 remains open as a separate follow-up discussion.One endpoint, two socket lifecycles
telnyx.transcription_enginepicks the recognizer and is sent to the endpoint verbatim: case-sensitive, no trimming, no case normalization. Two values are claimed:Deepgram(default) streams interims, so it gets one long-lived socket per session, the same shape asdeepgram.go: partials flow intoonResultwithIsFinal: false, and barge-in and live captions work out of the box. A finals-only default would silently switch both off for anyone who just sets the provider and starts talking.Telnyx(in-house) emits exactly one final per utterance, only after audio stops, and holds the socket open (verified live: no server close within 20s;interim_resultsignored). The client therefore endpoints the speech itself withinternal/vad.Detectorand opens one socket per utterance: 200ms onset, 600ms silence to close (the silence timeoutopenai.gouses for the same job), the finished utterance is uploaded at the fastest rate verified against the live endpoint (200ms of audio per write, 30ms pause), the single final is awaited, and the client closes the socket. One startup log line says barge-in and live captions are off for the session, so the operator learns it from the log, not from a caller talking over the agent.Design note on "open on speech onset": I took the buffer-then-dial shape from
openai.gorather than dialing against the first syllable. This engine emits no partials, so live streaming buys nothing, and a ~600ms dial racing the onset audio is a failure mode the buffered version does not have. Onset frames are preserved by a 1s idle lookback that becomes the utterance prefix.SendAudionever errors across utterance boundaries (an error would stop the pipeline inbound loop); a spent socket is simply replaced by the next utterance's dial.Docs and config
config.toml.example:transcription_engineunder[telnyx]with the two verified values;telnyxadded to the[stt]list.TestConfigExampleDocumentsEveryFieldholds it.docs/providers.md/docs/providers.zh-CN.md: Telnyx STT section, the Telnyx note placed directly under the whisper note, and the requested whisper-note honesty fix: finals-only means barge-in and live captions do not work, not just "no streaming partials".Tests (hermetic, fake WebSocket server)
Using the
grok_conn_test.gopattern (gorillaUpgrader+httptest, withtelnyxSTTURLoverridable the same waygrokDialURLis):is_finalhandling,confidence: nullmapped to 0 (unknown), errors-array detail surfacing, malformed-frame rejectionSendAudiokeeps accepting audio across the boundaryMixedCaseEnginedials with casing intact); an unset engine defaults to Deepgram; silence alone never dialsReal-audio smoke (per CONTRIBUTING)
Input: 13.5s of 16kHz mono speech ("The quick brown fox jumps over the lazy dog. Testing one two three, this is the Telnyx STT capture for StreamCore."), fed as 20ms frames at 1.5x realtime.
Gates
gofmt -l .clean,go build ./...,go vet ./...,go test -race ./...all green.go.mod/go.sumuntouched (no new dependencies;internal/vadwas already in the graph).