Skip to content

feat: 音频活动阈值开放为用户可设的灵敏度档位 - #99

Merged
IchenDEV merged 1 commit into
mainfrom
agent/qa/audio-sensitivity-setting
Sep 18, 2026
Merged

IchenDEV merged 1 commit into
mainfrom
agent/qa/audio-sensitivity-setting

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

背景

接 #95:AudioActivityThresholds 已把闸门与弱语音启发式解耦,本轮把它开放给用户,用「灵敏度」档位而非原始 RMS。

改动

设置面(General → 音频区)新增两个独立档位,各 3 档(保守 / 标准 / 灵敏):

档位 gate(语音检测)avg / peak weakSpeechEvidence(低音量清理)avg / peak
保守 0.0030 / 0.0100 0.0020 / 0.0060
标准 0.0015 / 0.0050 0.0040 / 0.0120
灵敏 0.00075 / 0.0025 0.0080 / 0.0240
  • 标准档严格等于原默认常数;老用户升级后行为不变。
  • 两组档位相互独立(gate 调整不影响 weakSpeechEvidence,反之亦然),保持 refactor: 音频活动阈值解耦 + RMS 诊断(不改默认数值,release 不输出) #95 的解耦语义。
  • 语义:gate 档位越高越「只接受响亮音频」(灵敏 = 更容易接受轻声/远场);weak 档位越高越容易把安静录音判为弱语音(触发重复折叠/幻听尾巴清理)。

持久化、校验与生效

  • AppSettings 新增 audioGateSensitivity / audioWeakSpeechSensitivity,沿用 @Published + UserDefaults 持久化约定。
  • 非法/未知持久化值回退 standard;解析出的阈值经 AudioActivityThresholds.clamped 做上下限截断([0.0001, 0.05],非有限值回退默认)。
  • VoicePipeline 与 InputSessionCoordinator 在每次录音开始前读取 settings.audioActivityThresholds 写入 AudioCaptureManager.thresholds,修改后即时生效、无需重启。
  • 未触碰诊断开关(UTTER_AUDIO_DIAGNOSTICS)与 release 行为。

测试

新增 Tests/OpenTypeTests/AudioSensitivityThresholdTests.swift(6 条):

  • 标准档 == 出厂默认(gate + weak + 合成阈值)。
  • 调 gate 档位后闸门判定随之改变(0.002 被保守拒绝、标准接受;0.001 被标准拒绝、灵敏接受)。
  • gate 档位不影响 hasWeakSpeechEvidence(refactor: 音频活动阈值解耦 + RMS 诊断(不改默认数值,release 不输出) #95 解耦语义)。
  • weak 档位只改弱语音判定、不改闸门。
  • 越界值截断到边界、NaN 回退默认。
  • 设置默认值/持久化/非法值回退默认。

结果:

  • swift build → Build complete! (4.27 secs)
  • 聚焦 6 个套件 → 80 passed / 0 failures
  • 扩展 Processing|Speech|Sanitizer|Audio|Configuration|Voice|Text|Prompt|Lexicon|Integration|Recognition|Transcript|Whisper|ASR|Guard|Fidelity|Replacement|Output|Sensitivity → 434 passed / 7 skipped / 0 failures
  • plutil -lint 两份 Localizable.strings 均 OK。

范围

10 个文件(7 源 + 2 本地化 + 1 测试),未改护栏、ASR 引擎与诊断逻辑。

Add two independent sensitivity presets (conservative/standard/sensitive) in the Audio settings section: one for the recording gate (hasMeaningfulAudio) and one for the weak-speech heuristic. Standard resolves exactly to the shipped defaults, so existing users keep their behavior.

Settings persist via UserDefaults, unknown raw values fall back to standard, and resolved thresholds are clamped to safe bounds. VoicePipeline and InputSessionCoordinator read the resolved thresholds before each recording, so changes apply without a restart. The diagnostics switch and release behavior are untouched.

Co-authored-by: multica-agent <github@multica.ai>
@IchenDEV
IchenDEV merged commit e736450 into main Sep 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant