Skip to content

RVC VST: persistent latency after GPU overload; recovery patch available #2859

Description

@AutoGavy

Problem

When GPU contention temporarily makes the RVC worker slower than real time, the VST bridge can retain old audio in both its input and output queues. After the GPU recovers, the worker catches up and fills the output queue, but the host can only drain that queue in real time. In practice this can leave the plugin replaying speech with roughly nine seconds of persistent delay until Cantabile/the plugin is restarted, even though GPU utilization has already returned to normal.

The current inference-time display does not expose this queueing age, so inference can appear healthy while old audio is still being replayed.

Root cause

The delay is not only the duration of one slow inference call. Old samples remain in the C++ input/output queues. Processing the backlog does not eliminate its age: it moves stale audio into the output queue, where it is consumed at real-time speed. Without an expiry policy and a discontinuity reset, the plugin can remain permanently behind live input.

Proposed patch

A complete patch is available here:

The repository currently has Pull Requests disabled, so I am opening this issue with the full implementation details.

1. Bounded cumulative audio age and overload recovery

  • Timestamp samples at plugin entry with a monotonic clock and retain that lineage through the input queue, inference request, and output queue. Finishing inference does not renew a sample's age.
  • Add MAX ACCEPTABLE LATENCY, default 300 ms, range 100-2000 ms.
  • Treat the value as one cumulative plugin-internal age budget across input queueing, inference, and output queueing—not a separate allowance for each stage and not an end-to-end microphone-to-OBS guarantee.
  • Preserve continuity while audio remains within the budget. Once it expires, discard stale input/output and late inference responses instead of replaying old speech, then resume from fresh audio.
  • When input expires, discard the expired prefix and, when required, advance to the newest complete fresh block. Retain a partial fresh block until it can be processed. Queues longer than three blocks remain intact when they are still within budget.
  • Re-check queued output against the current value, including when the user lowers the budget while audio is queued.
  • Treat input overflow as a discontinuity, invalidate the retained backlog, and resume with newly captured samples.
  • Maintain SPSC ownership: only consumers advance read cursors. Output invalidation publishes a monotonic flush boundary so samples written after the boundary survive; active rings are not reset concurrently from two threads.
  • Treat response sequence mismatches as worker errors/restart conditions rather than reusing shared memory that may still belong to an in-flight request.
  • Display inference duration, oldest returned wet-sample age, and discarded block-equivalents (drop) separately. age is sampled at the output callback and excludes sound-card, host, OBS/broadcast, and other external buffering.

Under sustained overload this policy can produce gaps/silence, but it prevents expired speech from accumulating and being replayed indefinitely.

2. Reset inference history after a discontinuity

  • Bump the worker IPC protocol to v2 and add a reset-stream request flag at request-header offset 68, bit 0.
  • Reset input, resampled-input, RMS, SOLA, and pitch-cache buffers in place before the next request, preserving CUDA Graph buffer identities.
  • Invalidate cached pitch, formant, and index values after the reset.
  • IPC v1 and v2 are intentionally incompatible; the plugin binary and worker/rvc_worker.py must be distributed together.

3. Dedicated worker GPU scheduling priority

  • Add GPU PRIORITY, enabled by default.
  • Request Windows GPU scheduling class HIGH using D3DKMTSetProcessSchedulingPriorityClass for the exact dedicated RVC worker process handle, then verify it using D3DKMTGetProcessSchedulingPriorityClass.
  • Report accepted, unavailable, failed, and unconfirmed states in the UI. Disabling the switch restores the worker's original class.
  • Perform the scheduling request on the bridge thread after worker/CUDA initialization, never on the audio callback.
  • This does not change other python.exe processes, request REALTIME, elevate privileges, reserve GPU/VRAM, or guarantee preemption over games.

4. UI controls and clearer toggle state

  • Add the GPU PRIORITY switch, MAX ACCEPTABLE LATENCY slider, overload guidance, and status text; enlarge the editor to accommodate them.
  • Change the ON fill of both ENGINE and GPU PRIORITY to red (RGB 190, 42, 48) with light text.
  • Previously, the pressed fill and label were both light, causing an enabled toggle to appear almost solid white and making its state difficult to read. The OFF styling remains unchanged.
  • Bump the plugin version to 0.1.4.

5. Persist realtime tuning parameters

Before 0.1.4, %LOCALAPPDATA%\RVCRealtime\settings.ini persisted paths only. The patch stores UI edits for all 13 tuning controls in [Parameters]:

  • Pitch
  • Formant
  • Index Rate
  • RMS Mix
  • Gate/Threshold
  • Block
  • Crossfade
  • Context
  • F0 Method
  • Mix
  • Output Gain
  • GPU Priority
  • Max Acceptable Latency

Persistence behavior:

  • Debounce writes for about 200 ms and flush pending edits when the editor closes or the instance unloads.
  • Perform file I/O only from UI/idle/teardown callbacks, never from audio processing or the DSP parameter callback.
  • Write only explicitly edited keys to reduce cross-instance overwrites.
  • Use locale-independent round-trip formatting and ignore non-finite or out-of-range saved values.
  • Retry failed writes and report the failure in the UI.
  • Keep ENGINE host-controlled rather than making it a global default; fresh instances remain off.
  • Preserve host project/preset precedence. Automation, startup synchronization, and project/preset recall do not overwrite the saved global defaults.
  • Synchronize saved VST3 defaults to the controller before the host first queries them.

Compatibility

  • Existing parameter IDs 0-11 are unchanged.
  • GPU PRIORITY and MAX ACCEPTABLE LATENCY are appended as IDs 12 and 13.
  • Legacy host state containing 12 parameters remains accepted, with defaults appended for the two new controls.
  • Existing [Paths] settings remain supported.
  • The fake worker used by recovery tests is isolated to the test build and is never packaged with the plugin.

Tests and validation

  • Release x64 VST3 build: passed.
  • ring-recovery: passed.
  • python-reset: passed.
  • bridge-recovery: passed.
  • git diff --check: passed.

Automated coverage includes timestamp/ring wrap behavior, partial discards, cumulative age across stages, live budget changes, retaining more than three blocks inside budget, a simulated nine-second backlog, concurrent SPSC producer/consumer activity, stale IPC responses, restart recovery, and in-place Python reset ordering/buffer identity.

Live validation with an actual model, audio device/host, and sustained game GPU contention is still required.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions