Skip to content

refactor: 音频活动阈值解耦 + RMS 诊断(不改默认数值,release 不输出) - #95

Merged
IchenDEV merged 1 commit into
mainfrom
agent/qa/audio-activity-threshold-prep
Sep 18, 2026
Merged

IchenDEV merged 1 commit into
mainfrom
agent/qa/audio-activity-threshold-prep

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

背景

#3/#4 标定前的准备:把「录制能量闸门」与「弱音频启发式」从共用常量解耦,并新增一条仅数值的 RMS 诊断,供后续用真实语料标定阈值。不改变任何默认阈值数值,也不改变 release 行为。

改动

  • 新增 Sources/Audio/AudioActivityThresholds.swift:gate 与 weakSpeechEvidence 两组独立配置(单一来源)。默认值与原常量逐项一致:gate 0.0015 / 0.005,weak 0.004 / 0.012。
  • Sources/Audio/AudioCaptureManager.swift:AudioCaptureActivity 持有 thresholds,hasMeaningfulAudio / hasWeakSpeechEvidence 各读自己那组,判定语义不变。
  • 新增 Sources/Audio/AudioCaptureDiagnostics.swift:AudioCaptureDiagnostic 记录 averageRMS / maxRMS / frameCount / gateRejected,仅数值,固定单行格式:
    audio-activity averageRMS=0.002000 maxRMS=0.002000 frames=16000 gateRejected=false
  • 调用点:Sources/App/VoicePipeline+Processing.swift(菜单栏听写)与 Sources/Integration/InputSessionCoordinator.swift(集成 API),每次录音记一条,包含被闸门拦下的录音。

诊断采集方式

诊断日志仅在 debug 构建 且设置环境变量 UTTER_AUDIO_DIAGNOSTICS=1 时输出;release 构建恒不输出。

# 开发模式
UTTER_AUDIO_DIAGNOSTICS=1 swift run OpenType

# 或构建并启动 dev .app
UTTER_AUDIO_DIAGNOSTICS=1 bash scripts/build-and-run.sh

抓取日志中 audio-activity 前缀行,得到每条录音的 averageRMS / maxRMS / frames / gateRejected(不含音频与转写内容)。把这些判定与 scripts/evaluate-voice-quality.py 的 CER / 幻听结果对照,即可标定两组阈值。

测试

  • swift build → Build complete! (4.10 secs)
  • swift test --filter 'AudioCaptureActivityTests|ConfigurationTests|TranscriptionSanitizerTests|VoicePipelinePolicyTests|QualityProbeTests' → 87 passed / 0 failures
  • 新增 Tests/OpenTypeTests/AudioCaptureActivityTests.swift(5 条):默认数值锁定、两组阈值互不影响、诊断格式稳定、诊断开关仅 debug + opt-in 生效。

未包含

#7(VocabularyReplacementEngine.swift 词典替换的 CJK 边界)留待后续。

Split the shared audio activity constants into AudioActivityThresholds with independent gate and weakSpeechEvidence groups, so relaxing the recording gate no longer shifts TranscriptionSanitizer's collapse/hallucination behavior. Default values are unchanged.

Add AudioCaptureDiagnostic, a numeric-only record (averageRMS/maxRMS/frameCount/gateRejected) with a stable single-line format. It is logged per recording from the menu bar and integration paths only in debug builds when UTTER_AUDIO_DIAGNOSTICS=1 is set, so release builds never emit it.

Co-authored-by: multica-agent <github@multica.ai>
@IchenDEV
IchenDEV merged commit 8ec8aab into main Sep 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant