Skip to content

feat: ASR 热词注入(Volc / Qwen3-ASR) - #98

Merged
IchenDEV merged 1 commit into
mainfrom
agent/qa/asr-hotword-injection
Sep 18, 2026
Merged

IchenDEV merged 1 commit into
mainfrom
agent/qa/asr-hotword-injection

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

摘要

按各引擎的公开 API 能力,把现有词库(SpeechRecognitionContext.phrases)接到明确支持热词/上下文偏置的引擎上;不支持的不改代码,只给结论。应用到的两个引擎均补了请求构造单测。

引擎 → 是否支持 → 本轮做了什么 → 验证方式

引擎 是否支持 依据 本轮做了什么 验证方式
Volc(豆包大模型流式 ASR) ✅ 支持 官方《大模型流式语音识别 API》request.corpus.context:双向流式支持热词直传,格式 {"hotwords":[{"word":"..."}]}(另有 boosting_table_id 词表方案) configureRecognition 存词库 → hotwordContext(for:) 生成 JSON 字符串 → fullClientRequestPayload 写入 request.corpus.context;流式 partial 与最终文件转写共用同一请求构造 Tests/OpenTypeTests/VolcSpeechEnginePayloadTests.swift(4 条)
QwenNative(Qwen3-ASR) ✅ 支持 mlx-audio-swift Qwen3ASRModel.generate(context:),buildPromptText 把 context 注入 system prompt(其自带单测 qwen3ASRPromptTextIncludesContextAndAssistantPrefix) configureRecognition 存词库 → contextualPrompt() → 传入 generate(context:) QwenNativeASREngineTests.testRecognitionContextPromptReachesTheModelCall + SpeechRecognitionQualityTests.testContextualPromptListsTermsAndRespectsBudget
MLXSTT(FireRed / Mega-ASR 通用封装) ❌ 不支持(当前封装) 通用协议 STTGenerationModel.generate(audio:generationParameters:) 与 STTGenerateParameters 均无 context/prompt 字段;MLXSTTEngine 持有 any STTGenerationModel,没有注入通道 不改代码 代码路径审查
FireRed(FireRedASR2-AED) ❌ 不支持(当前封装) FireRedASR2Model.generate 的两个重载都只有 beam/maxLen/language 参数,无 hotword/context 不改代码 代码路径审查

实现要点

  • Volc:hotwords 列表上限 100 条、总字符上限 300,保守低于官方文档中双向流式「100 tokens」的直传预算;不引入占位逻辑,超出预算的词条整条跳过。corpus 仅在词条非空时出现。
  • Qwen3-ASR:contextualPrompt(maximumCharacters: 600) 以 Terms: a, b, c 形式按词库排序注入,超预算的词条整条跳过(保留后续短词条)。
  • 两个引擎的 configureRecognition 均在既有调用点被调用(VoicePipeline.start、InputSessionCoordinator),无需改调用方。

测试

  • swift build → Build complete! (3.93 secs)
  • 聚焦:VolcSpeechEnginePayloadTests|SpeechRecognitionQualityTests|QwenNativeASREngineTests → 19 passed / 2 skipped / 0 failures(Volc 4 条含「词条确实进入 corpus.context」与预算丢弃;Qwen 断言 configureRecognition 后的上下文串即传给模型的参数)
  • 扩展 Speech|Volc|Qwen|Recognition|Processing|Whisper|ASR|Sanitizer|Integration|Prompt → 190 passed / 7 skipped / 0 failures

范围

6 个文件(3 源 + 3 测试),不改阈值、护栏、UI;FireRed/Mega-ASR 明确只出结论。后续若 mlx-audio-swift 的通用协议开放 context 字段,再统一接入。

Volc: build request.corpus.context as the documented {hotwords:[{word:...}]} JSON string from the recognition phrases, bounded for the 双向流式 direct-pass budget, and include it in the full client request. The streaming partials and the final file transcription share the same request path.

Qwen3-ASR: mlx-audio-swift exposes Qwen3ASRModel.generate(context:) which injects the text into the system prompt. Add configureRecognition storage and pass a bounded terms prompt; the generic STTGenerationModel protocol has no context parameter, so no other engine is touched.
Co-authored-by: multica-agent <github@multica.ai>
@IchenDEV
IchenDEV merged commit 14367cf into main Sep 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant