极其严重的bug,会导致模型产生严重卡顿和人机感,我不懂代码,遇到后叫dsh帮我查找修复并总结了一下,以下是dsh对bug的描述: # RVC 实时变声:索引检索被静默禁用导致音频卡顿/断续(nprobe 未设置)
环境
- RVC 版本:官方 release
2.3.260718(整合包 RVC20260718Nvidia50x0.7z)
- GPU:NVIDIA RTX 5070 Ti Laptop (sm_120, 12227 MiB)
- torch:2.7.1+cu128,faiss 1.14.3
- 启动方式:
go-realtime_gui.bat(realtime_gui.py)
- 模型:RVC v2 / 48kHz,index 为
added_IVF978_Flat_nprobe_1_*.index(38156 向量)
现象
实时变声中,命令行反复打印:
索引无效:必须使用added_xxxx.index,不能使用trained_xxxx.index
但使用的确实是 added_xxxx.index。实际听感:
- 声音断续、卡顿,明显「人机味」
- 音色在不同音频块之间跳动、不稳定
- 关闭索引(
index_rate=0)后反而变得明显更流畅
根本原因
1. 提示文本与实际判断条件无关
infer/rtrvc.py 第 250-270 行:
if hasattr(self, "index") and self.index_rate != 0:
npy = feats[0][skip_head // 2 :].cpu().numpy().astype("float32")
score, ix = self.index.search(npy, k=8)
if (ix >= 0).all(): # ← 真正的判断条件
weight = np.square(1 / score)
weight /= weight.sum(axis=1, keepdims=True)
npy = np.sum(self.big_npy[ix] * np.expand_dims(weight, axis=2), axis=1)
if self.config.is_half:
npy = npy.astype("float16")
feats[0][skip_head // 2 :] = (
torch.from_numpy(npy).unsqueeze(0).to(self.device) * self.index_rate
+ (1 - self.index_rate) * feats[0][skip_head // 2 :]
)
else:
printt(i18n("索引无效:必须使用added_xxxx.index,不能使用trained_xxxx.index"))
判断条件是 (ix >= 0).all(),完全没检查文件名是否包含 trained。提示文本具有误导性。
2. -1 来自 FAISS,因为 nprobe 从未被设置
infer/rtrvc.py 第 84-90 行加载索引:
if index_rate != 0:
self.index = faiss.read_index(index_path)
self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)
这里没有设置 self.index.nprobe,IndexIVFFlat 的默认值是 1。
索引构建时使用了很激进的聚类数(nlist ≈ ntotal / 39):
| 索引 |
ntotal |
nlist |
向量/簇 |
nprobe=1 时 -1 出现率 |
| miaomiao_e300 |
139795 |
3584 |
39.0 |
1.12% |
| sisi_e200 |
9768 |
250 |
39.1 |
1.40% |
| sagiri_e310 |
21738 |
557 |
39.0 |
0.10% |
| jyy_e266 |
5674 |
145 |
39.1 |
0.00% |
因为平均每个聚类只有约 39 个向量,当一个查询落入稀疏聚类时,凑不满 k=8 个邻居,FAISS 对缺失位置返回 -1。
3. 后果:特征流被不定时打断
(ix >= 0).all() 只要发现任意一个 -1,就会整块放弃检索,直接用未检索的 HuBERT 特征继续推理。
实时变声按块处理后,同一个音频里有的块走「检索后的特征」,有的块走「原始特征」——模型每一块的输入特征分布都在跳变,于是产生卡顿和音色不稳定。
复现(500 次随机查询,k=8):
nprobe=1 -> 负值 45/4000 (1.12%),受影响行 15/500
nprobe=8 -> 负值 0/4000 (0.00%),受影响行 0/500
建议的修复
主修复:加载索引时设置 nprobe
infer/rtrvc.py,在第 85-86 行之间插入一行:
if index_rate != 0:
self.index = faiss.read_index(index_path)
self.index.nprobe = 8 # ← 新增;建议同时把 nprobe 写进索引文件名或配置
self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)
infer/rtrvc.py 的 change_index_rate(第 141-146 行)也需要同样处理:
def change_index_rate(self, new_index_rate):
if new_index_rate != 0 and self.index_rate == 0:
self.index = faiss.read_index(self.index_path)
self.index.nprobe = 8 # ← 新增
self.big_npy = self.index.reconstruct_n(0, self.index.ntotal)
次要修复:判断条件应逐行处理 -1,而不是全丢弃
(ix >= 0).all() 过于严格。建议改为只对存在 -1 的行回退到原始特征,其余行正常检索:
score, ix = self.index.search(npy, k=8)
valid = (ix >= 0).all(axis=1)
if valid.any():
w = np.square(1 / score[valid])
w /= w.sum(axis=1, keepdims=True)
retrieved = np.sum(self.big_npy[ix[valid]] * np.expand_dims(w, axis=2), axis=1)
if self.config.is_half:
retrieved = retrieved.astype("float16")
blended = (torch.from_numpy(retrieved).unsqueeze(0).to(self.device)
* self.index_rate
+ (1 - self.index_rate) * feats[0][skip_head // 2 :][valid])
feats[0][skip_head // 2 :][valid] = blended
另外:离线路径 infer/vc/pipeline.py 第 183-186 行没有这个检查
score, ix = index.search(npy, k=8)
weight = np.square(1 / score)
weight /= weight.sum(axis=1, keepdims=True)
npy = np.sum(index_vectors[ix] * np.expand_dims(weight, axis=2), axis=1)
-1 会被 NumPy 当作 index_vectors[-1](即最后一个向量)解释,静默取到错误数据。
实测影响很小,因为 weight = (1/score)² 让最近邻的权重压倒性大、远处邻居权重趋近 0(实测 -1 位置的权重和为 0.0000),但仍是隐患,建议一并加保护。
补充说明
极其严重的bug,会导致模型产生严重卡顿和人机感,我不懂代码,遇到后叫dsh帮我查找修复并总结了一下,以下是dsh对bug的描述: # RVC 实时变声:索引检索被静默禁用导致音频卡顿/断续(nprobe 未设置)
环境
2.3.260718(整合包RVC20260718Nvidia50x0.7z)go-realtime_gui.bat(realtime_gui.py)added_IVF978_Flat_nprobe_1_*.index(38156 向量)现象
实时变声中,命令行反复打印:
但使用的确实是
added_xxxx.index。实际听感:index_rate=0)后反而变得明显更流畅根本原因
1. 提示文本与实际判断条件无关
infer/rtrvc.py第 250-270 行:判断条件是
(ix >= 0).all(),完全没检查文件名是否包含trained。提示文本具有误导性。2.
-1来自 FAISS,因为nprobe从未被设置infer/rtrvc.py第 84-90 行加载索引:这里没有设置
self.index.nprobe,IndexIVFFlat的默认值是1。索引构建时使用了很激进的聚类数(
nlist ≈ ntotal / 39):-1出现率因为平均每个聚类只有约 39 个向量,当一个查询落入稀疏聚类时,凑不满
k=8个邻居,FAISS 对缺失位置返回-1。3. 后果:特征流被不定时打断
(ix >= 0).all()只要发现任意一个-1,就会整块放弃检索,直接用未检索的 HuBERT 特征继续推理。实时变声按块处理后,同一个音频里有的块走「检索后的特征」,有的块走「原始特征」——模型每一块的输入特征分布都在跳变,于是产生卡顿和音色不稳定。
复现(500 次随机查询,k=8):
建议的修复
主修复:加载索引时设置
nprobeinfer/rtrvc.py,在第 85-86 行之间插入一行:infer/rtrvc.py的change_index_rate(第 141-146 行)也需要同样处理:次要修复:判断条件应逐行处理
-1,而不是全丢弃(ix >= 0).all()过于严格。建议改为只对存在-1的行回退到原始特征,其余行正常检索:另外:离线路径
infer/vc/pipeline.py第 183-186 行没有这个检查-1会被 NumPy 当作index_vectors[-1](即最后一个向量)解释,静默取到错误数据。实测影响很小,因为
weight = (1/score)²让最近邻的权重压倒性大、远处邻居权重趋近 0(实测-1位置的权重和为0.0000),但仍是隐患,建议一并加保护。补充说明
nprobe过小导致 FAISS 返回-1,与使用trained_还是added_索引无关。当前文本会把人引向错误方向(社区里已有人去排查索引文件名,见 Index search FAILED or disabled怎么解决? #1907、request help about installation or cant provide index after training #2585)。nprobe与运行时代码保持一致。生成索引时文件名里已经写了nprobe_1(例如added_IVF978_Flat_nprobe_1_*.index),说明这个参数是已知的,但读取端没有使用它。可以在read_index后用正则从文件名解析nprobe并设置。nprobe=8后,-1出现率从 1.12% 降到 0.00%,实时变声的卡顿与音色跳变问题消失,同一模型的主观听感显著改善。