Skip to content

索引检索被静默禁用导致音频卡顿/断续(nprobe 未设置) #2864

Description

@tayunxunyin

极其严重的bug,会导致模型产生严重卡顿和人机感,我不懂代码,遇到后叫dsh帮我查找修复并总结了一下,以下是dsh对bug的描述: # RVC 实时变声:索引检索被静默禁用导致音频卡顿/断续(nprobe 未设置)

环境

  • RVC 版本:官方 release 2.3.260718(整合包 RVC20260718Nvidia50x0.7z)
  • GPU:NVIDIA RTX 5070 Ti Laptop (sm_120, 12227 MiB)
  • torch:2.7.1+cu128,faiss 1.14.3
  • 启动方式:go-realtime_gui.bat(realtime_gui.py)
  • 模型:RVC v2 / 48kHz,index 为 added_IVF978_Flat_nprobe_1_*.index(38156 向量)

现象

实时变声中,命令行反复打印:

索引无效:必须使用added_xxxx.index,不能使用trained_xxxx.index

但使用的确实是 added_xxxx.index。实际听感:

  • 声音断续、卡顿,明显「人机味」
  • 音色在不同音频块之间跳动、不稳定
  • 关闭索引(index_rate=0)后反而变得明显更流畅

根本原因

1. 提示文本与实际判断条件无关

infer/rtrvc.py 第 250-270 行:

if hasattr(self, "index") and self.index_rate != 0:
    npy = feats[0][skip_head // 2 :].cpu().numpy().astype("float32")
    score, ix = self.index.search(npy, k=8)
    if (ix >= 0).all():                    # ← 真正的判断条件
        weight = np.square(1 / score)
        weight /= weight.sum(axis=1, keepdims=True)
        npy = np.sum(self.big_npy[ix] * np.expand_dims(weight, axis=2), axis=1)
        if self.config.is_half:
            npy = npy.astype("float16")
        feats[0][skip_head // 2 :] = (
            torch.from_numpy(npy).unsqueeze(0).to(self.device) * self.index_rate
            + (1 - self.index_rate) * feats[0][skip_head // 2 :]
        )
    else:
        printt(i18n("索引无效:必须使用added_xxxx.index,不能使用trained_xxxx.index"))

判断条件是 (ix >= 0).all(),完全没检查文件名是否包含 trained。提示文本具有误导性。

2. -1 来自 FAISS,因为 nprobe 从未被设置

infer/rtrvc.py 第 84-90 行加载索引:

if index_rate != 0:
    self.index = faiss.read_index(index_path)
    self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)

这里没有设置 self.index.nprobe,IndexIVFFlat 的默认值是 1。

索引构建时使用了很激进的聚类数(nlist ≈ ntotal / 39):

索引 ntotal nlist 向量/簇 nprobe=1 时 -1 出现率
miaomiao_e300 139795 3584 39.0 1.12%
sisi_e200 9768 250 39.1 1.40%
sagiri_e310 21738 557 39.0 0.10%
jyy_e266 5674 145 39.1 0.00%

因为平均每个聚类只有约 39 个向量,当一个查询落入稀疏聚类时,凑不满 k=8 个邻居,FAISS 对缺失位置返回 -1。

3. 后果:特征流被不定时打断

(ix >= 0).all() 只要发现任意一个 -1,就会整块放弃检索,直接用未检索的 HuBERT 特征继续推理。

实时变声按块处理后,同一个音频里有的块走「检索后的特征」,有的块走「原始特征」——模型每一块的输入特征分布都在跳变,于是产生卡顿和音色不稳定。

复现(500 次随机查询,k=8):

nprobe=1   -> 负值 45/4000 (1.12%),受影响行 15/500
nprobe=8   -> 负值 0/4000  (0.00%),受影响行 0/500

建议的修复

主修复:加载索引时设置 nprobe

infer/rtrvc.py,在第 85-86 行之间插入一行:

if index_rate != 0:
    self.index = faiss.read_index(index_path)
    self.index.nprobe = 8          # ← 新增;建议同时把 nprobe 写进索引文件名或配置
    self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)

infer/rtrvc.py 的 change_index_rate(第 141-146 行)也需要同样处理:

def change_index_rate(self, new_index_rate):
    if new_index_rate != 0 and self.index_rate == 0:
        self.index = faiss.read_index(self.index_path)
        self.index.nprobe = 8      # ← 新增
        self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)

次要修复:判断条件应逐行处理 -1,而不是全丢弃

(ix >= 0).all() 过于严格。建议改为只对存在 -1 的行回退到原始特征,其余行正常检索:

score, ix = self.index.search(npy, k=8)
valid = (ix >= 0).all(axis=1)
if valid.any():
    w = np.square(1 / score[valid])
    w /= w.sum(axis=1, keepdims=True)
    retrieved = np.sum(self.big_npy[ix[valid]] * np.expand_dims(w, axis=2), axis=1)
    if self.config.is_half:
        retrieved = retrieved.astype("float16")
    blended = (torch.from_numpy(retrieved).unsqueeze(0).to(self.device)
               * self.index_rate
               + (1 - self.index_rate) * feats[0][skip_head // 2 :][valid])
    feats[0][skip_head // 2 :][valid] = blended

另外:离线路径 infer/vc/pipeline.py 第 183-186 行没有这个检查

score, ix = index.search(npy, k=8)
weight = np.square(1 / score)
weight /= weight.sum(axis=1, keepdims=True)
npy = np.sum(index_vectors[ix] * np.expand_dims(weight, axis=2), axis=1)

-1 会被 NumPy 当作 index_vectors[-1](即最后一个向量)解释,静默取到错误数据。
实测影响很小,因为 weight = (1/score)² 让最近邻的权重压倒性大、远处邻居权重趋近 0(实测 -1 位置的权重和为 0.0000),但仍是隐患,建议一并加保护。

补充说明

  • 提示文本建议同步修正。真正的原因是 nprobe 过小导致 FAISS 返回 -1,与使用 trained_ 还是 added_ 索引无关。当前文本会把人引向错误方向(社区里已有人去排查索引文件名,见 Index search FAILED or disabled怎么解决? #1907、request help about installation or cant provide index after training #2585)。
  • 建议索引文件名的 nprobe 与运行时代码保持一致。生成索引时文件名里已经写了 nprobe_1(例如 added_IVF978_Flat_nprobe_1_*.index),说明这个参数是已知的,但读取端没有使用它。可以在 read_index 后用正则从文件名解析 nprobe 并设置。
  • 我这边用 monkey-patch 的方式验证过修复效果:设置 nprobe=8 后,-1 出现率从 1.12% 降到 0.00%,实时变声的卡顿与音色跳变问题消失,同一模型的主观听感显著改善。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions