Skip to content

fix: CJK 转写分段拼接修复 - #94

Merged
IchenDEV merged 1 commit into
mainfrom
agent/qa/20adf000e3c1
Sep 18, 2026
Merged

IchenDEV merged 1 commit into
mainfrom
agent/qa/20adf000e3c1

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

背景

WhisperKit 在 VAD 分段下每个 chunk 返回一条文本,原实现统一用空格拼接,导致中文/日文/韩文段间插入空格,污染原始转写与后续 LLM 输入。

改动

  • 新增 Sources/Speech/TranscriptSegmentJoiner.swift:zh/yue/ja/ko 段间无分隔符,其他语言用单空格;语言未知(auto)时按相邻边界字符判断(任一侧为 CJK 则无空格)。逐段 trim、丢弃空段。
  • Sources/Speech/WhisperEngine.swift:文件转写改用该拼接器,并把 language 透传给流式会话。
  • Sources/Speech/WhisperStreamingSession.swift:新增 language 属性(init 默认 nil),partial 拼接同样走拼接器。
  • Tests/OpenTypeTests/TranscriptSegmentJoinerTests.swift(新增):zh/ja/ko 无空格、en 有空格、auto 内容回退、trim/空段。
  • Tests/OpenTypeTests/SpeechRecognitionQualityTests.swift:新增 auto 语言下热词 prompt token 非空的回归测试(复核确认 Add Volcengine Doubao ASR speech engine support #1 为假阳性,仅锁定现有行为)。

测试

  • swift build → Build complete! (164.77 secs)
  • swift test --filter 'TranscriptSegmentJoinerTests|SpeechRecognitionQualityTests' → 13 passed / 0 failures
  • 扩展 StreamingSpeechSupportTests|TranscriptionSanitizerTests|RealtimeAudioConversionTests → 合计 38 passed / 0 failures

备注

#3/#4(录音能量闸门与弱音频启发式阈值)本轮不改:缺带实测 averageRMS/maxRMS 的真实语料,无法标定,后续单独立项。

WhisperKit returns one text per VAD chunk; joining with a single space inserted spaces between Chinese/Japanese/Korean segments. Add TranscriptSegmentJoiner that selects the separator from the language (auto falls back to boundary characters) and use it for both file and streaming transcription.

Also add a regression test locking hotword prompt tokens when the language is auto.

Co-authored-by: multica-agent <github@multica.ai>
@IchenDEV
IchenDEV merged commit 2df1257 into main Sep 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant