Skip to content

Bug: cmvn readers raise IndexError on blank lines and return empty statistics without errorΒ #3759

Description

@Lesereingrape

πŸ› Bug

FunASR has two cmvn readers in the installed package, and both walk the file line by line while indexing the result without any check:

  • funasr/frontends/wav_frontend.py:15 load_cmvn β€” used by WavFrontend / WavFrontendOnline, i.e. the frontend of the Paraformer and SenseVoice models most people load first;
  • funasr/frontends/default.py:390 MultiChannelFrontend._load_cmvn β€” used by MultiChannelFrontend (emotion2vec / campplus).
    for i in range(len(lines)):
        line_item = lines[i].split()
        if line_item[0] == "<AddShift>":          # blank line -> IndexError
            line_item = lines[i + 1].split()      # tag on the last line -> IndexError
            if line_item[0] == "<LearnRateCoef>":

Two distinct symptoms come out of this, both measured on main @ 66d7a4c2:

cmvn_file content wav_frontend.load_cmvn MultiChannelFrontend._load_cmvn
runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn (a real artifact this repository ships) OK, shape (2, 560) OK, shape (560,)
the same file plus one trailing blank line IndexError: list index out of range IndexError: list index out of range
a blank (or whitespace-only) line between <AddShift> and its <LearnRateCoef> line IndexError: list index out of range IndexError: list index out of range
<AddShift> 2 2 as the last line (truncated file) IndexError: list index out of range IndexError: list index out of range
statistics present but the tags do not start a line on their own, e.g. the whole <Nnet> on one line returns a tensor of shape (2, 0) β€” no error returns two empty arrays β€” no error

The first four crash while the frontend is being constructed, so AutoModel(...) or a training run dies with a bare IndexError that never mentions the cmvn file. The fifth is worse because it is silent: load_cmvn returns empty statistics, WavFrontend builds without complaint, and the failure only appears at the first inference.

To Reproduce

  1. Install with: pip install -e . (source checkout of main @ 66d7a4c2)
  2. Copy the cmvn artifact the repository already ships and append a blank line:
    cp runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/am.mvn /tmp/am.mvn && printf '\n' >> /tmp/am.mvn
  3. Run the code sample below.
  4. See error: IndexError: list index out of range

Code sample

from funasr.frontends.wav_frontend import WavFrontend

WavFrontend(cmvn_file="am.mvn", fs=16000, win_length=400, hop_length=160, n_mels=80)
# IndexError: list index out of range   -- only when the file has an extra blank line;
# without the blank line the same call succeeds.

# silent variant: tags do not start a line on their own
from funasr.frontends.wav_frontend import load_cmvn
open("compact.mvn", "w").write(
    "<Nnet> <AddShift> <LearnRateCoef> 1 [ -1.0 -2.0 ] </LearnRateCoef> "
    "<Rescale> <LearnRateCoef> 1 [ 0.5 0.5 ] </LearnRateCoef> </Nnet>\n"
)
cmvn = load_cmvn("compact.mvn")   # shape (2, 0), no error

Measured with n_mels=3 on the compact file, the frontend then builds fine and the first input_feats += self.mean / apply_cmvn call fails at inference time:

RuntimeError: The size of tensor a (3) must match the size of tensor b (0) at non-singleton dimension 1

MultiChannelFrontend behaves the same way (input_feats += self.mean against an empty registered buffer, funasr/frontends/default.py:355).

Expected behavior

A cmvn file that contains valid statistics should load regardless of blank or whitespace-only separator lines, and a file that yields no statistics should raise an error naming the file instead of returning empty arrays. Files that parse correctly today must keep parsing identically β€” the am.mvn above is the regression case.

Error logs

  File ".../funasr/frontends/wav_frontend.py", line 27, in load_cmvn
    if line_item[0] == "<AddShift>":
IndexError: list index out of range

Environment

  • OS: Windows 11 (also reproduces on Linux β€” the loop is platform-independent)
  • Python version: 3.13.7
  • FunASR version: source, main @ 66d7a4c2
  • PyTorch / torchaudio version: 2.14.1+cpu / matching
  • numpy: 2.5.3
  • Install method: source (pip install -e .)
  • Device: cpu
  • No GPU, no model download needed β€” the reproduction uses the cmvn text file the repository already ships.

Audio details

Not applicable: no audio is involved, the defect is in the plain-text cmvn parser.

Suggested fix

For the two package readers:

  1. drop blank/whitespace-only lines before pairing a section tag with the line below it β€” that also makes <AddShift> + blank line + <LearnRateCoef> parse, and matches how this repository already reads other plain-text inputs (funasr/bin/realtime_ws.py:1363 skips blank hotword lines, funasr/utils/compute_det_ctc.py:59 skips blank stat lines);
  2. bound the one-line lookahead so a truncated file cannot raise IndexError;
  3. raise a ValueError naming the file when neither <AddShift> nor <Rescale> produced statistics, instead of handing back empty arrays β€” the same "say which cmvn file is wrong" convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148 raises FileNotFoundError("cmvn file not exits")).

Five more copies of the same loop exist under runtime/ (runtime/python/libtorch/, runtime/python/onnxruntime/, runtime/triton_gpu/, export_lfr_cmvn_pe_onnx.py); they are standalone export/deployment helpers rather than the installed package, so I would scope any PR to the two package readers.

Filed with an AI coding agent, based on the measurements above that I ran locally on this machine.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions