refactor: 音频活动阈值解耦 + RMS 诊断(不改默认数值,release 不输出) - #95
Merged
Merged
Conversation
Split the shared audio activity constants into AudioActivityThresholds with independent gate and weakSpeechEvidence groups, so relaxing the recording gate no longer shifts TranscriptionSanitizer's collapse/hallucination behavior. Default values are unchanged. Add AudioCaptureDiagnostic, a numeric-only record (averageRMS/maxRMS/frameCount/gateRejected) with a stable single-line format. It is logged per recording from the menu bar and integration paths only in debug builds when UTTER_AUDIO_DIAGNOSTICS=1 is set, so release builds never emit it. Co-authored-by: multica-agent <github@multica.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
背景
#3/#4 标定前的准备:把「录制能量闸门」与「弱音频启发式」从共用常量解耦,并新增一条仅数值的 RMS 诊断,供后续用真实语料标定阈值。不改变任何默认阈值数值,也不改变 release 行为。
改动
Sources/Audio/AudioActivityThresholds.swift:gate与weakSpeechEvidence两组独立配置(单一来源)。默认值与原常量逐项一致:gate0.0015 / 0.005,weak0.004 / 0.012。Sources/Audio/AudioCaptureManager.swift:AudioCaptureActivity持有thresholds,hasMeaningfulAudio/hasWeakSpeechEvidence各读自己那组,判定语义不变。Sources/Audio/AudioCaptureDiagnostics.swift:AudioCaptureDiagnostic记录averageRMS/maxRMS/frameCount/gateRejected,仅数值,固定单行格式:audio-activity averageRMS=0.002000 maxRMS=0.002000 frames=16000 gateRejected=falseSources/App/VoicePipeline+Processing.swift(菜单栏听写)与Sources/Integration/InputSessionCoordinator.swift(集成 API),每次录音记一条,包含被闸门拦下的录音。诊断采集方式
诊断日志仅在 debug 构建 且设置环境变量
UTTER_AUDIO_DIAGNOSTICS=1时输出;release 构建恒不输出。抓取日志中
audio-activity前缀行,得到每条录音的averageRMS/maxRMS/frames/gateRejected(不含音频与转写内容)。把这些判定与scripts/evaluate-voice-quality.py的 CER / 幻听结果对照,即可标定两组阈值。测试
swift build→Build complete! (4.10 secs)swift test --filter 'AudioCaptureActivityTests|ConfigurationTests|TranscriptionSanitizerTests|VoicePipelinePolicyTests|QualityProbeTests'→ 87 passed / 0 failuresTests/OpenTypeTests/AudioCaptureActivityTests.swift(5 条):默认数值锁定、两组阈值互不影响、诊断格式稳定、诊断开关仅 debug + opt-in 生效。未包含
#7(
VocabularyReplacementEngine.swift词典替换的 CJK 边界)留待后续。