feat: ASR 热词注入(Volc / Qwen3-ASR) - #98
Merged
Merged
Conversation
Volc: build request.corpus.context as the documented {hotwords:[{word:...}]} JSON string from the recognition phrases, bounded for the 双向流式 direct-pass budget, and include it in the full client request. The streaming partials and the final file transcription share the same request path.
Qwen3-ASR: mlx-audio-swift exposes Qwen3ASRModel.generate(context:) which injects the text into the system prompt. Add configureRecognition storage and pass a bounded terms prompt; the generic STTGenerationModel protocol has no context parameter, so no other engine is touched.
Co-authored-by: multica-agent <github@multica.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
摘要
按各引擎的公开 API 能力,把现有词库(
SpeechRecognitionContext.phrases)接到明确支持热词/上下文偏置的引擎上;不支持的不改代码,只给结论。应用到的两个引擎均补了请求构造单测。引擎 → 是否支持 → 本轮做了什么 → 验证方式
request.corpus.context:双向流式支持热词直传,格式{"hotwords":[{"word":"..."}]}(另有boosting_table_id词表方案)configureRecognition存词库 →hotwordContext(for:)生成 JSON 字符串 →fullClientRequestPayload写入request.corpus.context;流式 partial 与最终文件转写共用同一请求构造Tests/OpenTypeTests/VolcSpeechEnginePayloadTests.swift(4 条)Qwen3ASRModel.generate(context:),buildPromptText把 context 注入 system prompt(其自带单测qwen3ASRPromptTextIncludesContextAndAssistantPrefix)configureRecognition存词库 →contextualPrompt()→ 传入generate(context:)QwenNativeASREngineTests.testRecognitionContextPromptReachesTheModelCall+SpeechRecognitionQualityTests.testContextualPromptListsTermsAndRespectsBudgetSTTGenerationModel.generate(audio:generationParameters:)与STTGenerateParameters均无 context/prompt 字段;MLXSTTEngine持有any STTGenerationModel,没有注入通道FireRedASR2Model.generate的两个重载都只有 beam/maxLen/language 参数,无 hotword/context实现要点
hotwords列表上限 100 条、总字符上限 300,保守低于官方文档中双向流式「100 tokens」的直传预算;不引入占位逻辑,超出预算的词条整条跳过。corpus仅在词条非空时出现。contextualPrompt(maximumCharacters: 600)以Terms: a, b, c形式按词库排序注入,超预算的词条整条跳过(保留后续短词条)。configureRecognition均在既有调用点被调用(VoicePipeline.start、InputSessionCoordinator),无需改调用方。测试
swift build→Build complete! (3.93 secs)VolcSpeechEnginePayloadTests|SpeechRecognitionQualityTests|QwenNativeASREngineTests→ 19 passed / 2 skipped / 0 failures(Volc 4 条含「词条确实进入corpus.context」与预算丢弃;Qwen 断言configureRecognition后的上下文串即传给模型的参数)Speech|Volc|Qwen|Recognition|Processing|Whisper|ASR|Sanitizer|Integration|Prompt→ 190 passed / 7 skipped / 0 failures范围
6 个文件(3 源 + 3 测试),不改阈值、护栏、UI;FireRed/Mega-ASR 明确只出结论。后续若 mlx-audio-swift 的通用协议开放 context 字段,再统一接入。