Skip to content

Add EmitsPartials capability flag for finals-only STT providers - #82

Merged
jason-shen merged 1 commit into
streamcoreai:mainfrom
a692570:stt-emits-partials
Sep 22, 2026
Merged

jason-shen merged 1 commit into
streamcoreai:mainfrom
a692570:stt-emits-partials

Conversation

@a692570

@a692570 a692570 commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

Summary

Implements the shape sketched in #75: an optional EmitsPartials() bool capability interface in internal/stt, type-asserted once in runInbound and defaulting to true, so finals-only providers get a degraded barge-in instead of none.

  • stt.PartialsEmitter sits next to Client; providers that stream interims return true, finals-only providers return false, and providers that don't implement it keep the partials-driven path exactly as before.
  • runInbound asserts once after stt.NewClient. When false, the backchannel candidate opens on VAD alone, and the interrupt still cannot fire until the full 600ms backchannelWindow has elapsed: a burst that ends inside the window classifies as backchannel, since there is no partial text to say otherwise. Live captions are unchanged (finals flow as before).
  • Wired where needed: openaiClient returns false (batch transcription); telnyxSessionClient returns true, telnyxUtteranceClient (the exact-case in-house Telnyx engine) returns false, the same split NewTelnyxClient already uses for socket lifecycle.

Why

openai has had dead barge-in since day one and nobody wrote it down; the in-house Telnyx engine shipped the same hole. Waiting for partial text that never arrives means hasPartialText never goes true and the barge-in trigger never fires. VAD-only after the full window is the honest degraded mode: sustained interruptions work, short unclassifiable bursts stay suppressed, and nothing about any other provider changes.

Existing-provider safety

For every provider without the interface (and telnyx hosted engines), emitsPartials is true and the trigger short-circuits to the original hasPartialText.Load() expression, same bytes, same frames, same state machine. No dispatch or config schema changes; the only existing-provider edits are the EmitsPartials methods, plus rewording the telnyx startup log line and one stale comment that claimed barge-in was off.

Tests

  • internal/stt: plain Client does not satisfy PartialsEmitter (the default-true contract); telnyx true/false by engine through NewTelnyxClient; openai false.
  • internal/pipeline: hermetic runInbound drives: no partial text means no barge-in (current path unchanged); partials plus sustained speech fires; VAD-only holds the interrupt until the window elapses and fires after; a short burst inside the window is suppressed.
  • CI green: gofmt -l ., go build ./..., go vet ./..., go test -race ./....

Docs

All "barge-in and live captions are off" claims (providers.md and zh-CN, README and zh-CN, configuration.md and zh-CN, the config.toml.example comment, the config.go doc comment, and the telnyx startup log) now describe the degraded mode instead.

Follows #78 (Telnyx STT) and option 3 from #75.

@jason-shen
jason-shen merged commit bba9c57 into streamcoreai:main Sep 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants