π Bug
FunASR has two cmvn readers in the installed package, and both walk the file line by line while indexing the result without any check:
funasr/frontends/wav_frontend.py:15 load_cmvn β used by WavFrontend / WavFrontendOnline, i.e. the frontend of the Paraformer and SenseVoice models most people load first;
funasr/frontends/default.py:390 MultiChannelFrontend._load_cmvn β used by MultiChannelFrontend (emotion2vec / campplus).
for i in range(len(lines)):
line_item = lines[i].split()
if line_item[0] == "<AddShift>": # blank line -> IndexError
line_item = lines[i + 1].split() # tag on the last line -> IndexError
if line_item[0] == "<LearnRateCoef>":
Two distinct symptoms come out of this, both measured on main @ 66d7a4c2:
cmvn_file content |
wav_frontend.load_cmvn |
MultiChannelFrontend._load_cmvn |
runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn (a real artifact this repository ships) |
OK, shape (2, 560) |
OK, shape (560,) |
| the same file plus one trailing blank line |
IndexError: list index out of range |
IndexError: list index out of range |
a blank (or whitespace-only) line between <AddShift> and its <LearnRateCoef> line |
IndexError: list index out of range |
IndexError: list index out of range |
<AddShift> 2 2 as the last line (truncated file) |
IndexError: list index out of range |
IndexError: list index out of range |
statistics present but the tags do not start a line on their own, e.g. the whole <Nnet> on one line |
returns a tensor of shape (2, 0) β no error |
returns two empty arrays β no error |
The first four crash while the frontend is being constructed, so AutoModel(...) or a training run dies with a bare IndexError that never mentions the cmvn file. The fifth is worse because it is silent: load_cmvn returns empty statistics, WavFrontend builds without complaint, and the failure only appears at the first inference.
To Reproduce
- Install with:
pip install -e . (source checkout of main @ 66d7a4c2)
- Copy the cmvn artifact the repository already ships and append a blank line:
cp runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn /tmp/am.mvn && printf '\n' >> /tmp/am.mvn
- Run the code sample below.
- See error:
IndexError: list index out of range
Code sample
from funasr.frontends.wav_frontend import WavFrontend
WavFrontend(cmvn_file="am.mvn", fs=16000, win_length=400, hop_length=160, n_mels=80)
# IndexError: list index out of range -- only when the file has an extra blank line;
# without the blank line the same call succeeds.
# silent variant: tags do not start a line on their own
from funasr.frontends.wav_frontend import load_cmvn
open("compact.mvn", "w").write(
"<Nnet> <AddShift> <LearnRateCoef> 1 [ -1.0 -2.0 ] </LearnRateCoef> "
"<Rescale> <LearnRateCoef> 1 [ 0.5 0.5 ] </LearnRateCoef> </Nnet>\n"
)
cmvn = load_cmvn("compact.mvn") # shape (2, 0), no error
Measured with n_mels=3 on the compact file, the frontend then builds fine and the first input_feats += self.mean / apply_cmvn call fails at inference time:
RuntimeError: The size of tensor a (3) must match the size of tensor b (0) at non-singleton dimension 1
MultiChannelFrontend behaves the same way (input_feats += self.mean against an empty registered buffer, funasr/frontends/default.py:355).
Expected behavior
A cmvn file that contains valid statistics should load regardless of blank or whitespace-only separator lines, and a file that yields no statistics should raise an error naming the file instead of returning empty arrays. Files that parse correctly today must keep parsing identically β the am.mvn above is the regression case.
Error logs
File ".../funasr/frontends/wav_frontend.py", line 27, in load_cmvn
if line_item[0] == "<AddShift>":
IndexError: list index out of range
Environment
- OS: Windows 11 (also reproduces on Linux β the loop is platform-independent)
- Python version: 3.13.7
- FunASR version: source,
main @ 66d7a4c2
- PyTorch / torchaudio version: 2.14.1+cpu / matching
- numpy: 2.5.3
- Install method: source (
pip install -e .)
- Device:
cpu
- No GPU, no model download needed β the reproduction uses the cmvn text file the repository already ships.
Audio details
Not applicable: no audio is involved, the defect is in the plain-text cmvn parser.
Suggested fix
For the two package readers:
- drop blank/whitespace-only lines before pairing a section tag with the line below it β that also makes
<AddShift> + blank line + <LearnRateCoef> parse, and matches how this repository already reads other plain-text inputs (funasr/bin/realtime_ws.py:1363 skips blank hotword lines, funasr/utils/compute_det_ctc.py:59 skips blank stat lines);
- bound the one-line lookahead so a truncated file cannot raise
IndexError;
- raise a
ValueError naming the file when neither <AddShift> nor <Rescale> produced statistics, instead of handing back empty arrays β the same "say which cmvn file is wrong" convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148 raises FileNotFoundError("cmvn file not exits")).
Five more copies of the same loop exist under runtime/ (runtime/python/libtorch/, runtime/python/onnxruntime/, runtime/triton_gpu/, export_lfr_cmvn_pe_onnx.py); they are standalone export/deployment helpers rather than the installed package, so I would scope any PR to the two package readers.
Filed with an AI coding agent, based on the measurements above that I ran locally on this machine.
π Bug
FunASR has two cmvn readers in the installed package, and both walk the file line by line while indexing the result without any check:
funasr/frontends/wav_frontend.py:15 load_cmvnβ used byWavFrontend/WavFrontendOnline, i.e. the frontend of the Paraformer and SenseVoice models most people load first;funasr/frontends/default.py:390 MultiChannelFrontend._load_cmvnβ used byMultiChannelFrontend(emotion2vec / campplus).Two distinct symptoms come out of this, both measured on
main@66d7a4c2:cmvn_filecontentwav_frontend.load_cmvnMultiChannelFrontend._load_cmvnruntime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn(a real artifact this repository ships)(2, 560)(560,)IndexError: list index out of rangeIndexError: list index out of range<AddShift>and its<LearnRateCoef>lineIndexError: list index out of rangeIndexError: list index out of range<AddShift> 2 2as the last line (truncated file)IndexError: list index out of rangeIndexError: list index out of range<Nnet>on one line(2, 0)β no errorThe first four crash while the frontend is being constructed, so
AutoModel(...)or a training run dies with a bareIndexErrorthat never mentions the cmvn file. The fifth is worse because it is silent:load_cmvnreturns empty statistics,WavFrontendbuilds without complaint, and the failure only appears at the first inference.To Reproduce
pip install -e .(source checkout ofmain@66d7a4c2)cp runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn /tmp/am.mvn && printf '\n' >> /tmp/am.mvnIndexError: list index out of rangeCode sample
Measured with
n_mels=3on the compact file, the frontend then builds fine and the firstinput_feats += self.mean/apply_cmvncall fails at inference time:MultiChannelFrontendbehaves the same way (input_feats += self.meanagainst an empty registered buffer,funasr/frontends/default.py:355).Expected behavior
A cmvn file that contains valid statistics should load regardless of blank or whitespace-only separator lines, and a file that yields no statistics should raise an error naming the file instead of returning empty arrays. Files that parse correctly today must keep parsing identically β the
am.mvnabove is the regression case.Error logs
Environment
main@66d7a4c2pip install -e .)cpuAudio details
Not applicable: no audio is involved, the defect is in the plain-text cmvn parser.
Suggested fix
For the two package readers:
<AddShift>+ blank line +<LearnRateCoef>parse, and matches how this repository already reads other plain-text inputs (funasr/bin/realtime_ws.py:1363skips blank hotword lines,funasr/utils/compute_det_ctc.py:59skips blank stat lines);IndexError;ValueErrornaming the file when neither<AddShift>nor<Rescale>produced statistics, instead of handing back empty arrays β the same "say which cmvn file is wrong" convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148raisesFileNotFoundError("cmvn file not exits")).Five more copies of the same loop exist under
runtime/(runtime/python/libtorch/,runtime/python/onnxruntime/,runtime/triton_gpu/,export_lfr_cmvn_pe_onnx.py); they are standalone export/deployment helpers rather than the installed package, so I would scope any PR to the two package readers.Filed with an AI coding agent, based on the measurements above that I ran locally on this machine.