Filter Whisper "Do anything" hallucination on weak audio - #92
Merged
Merged
Conversation
When no speech is present but background noise passes the energy threshold, Whisper may hallucinate "Do anything". Add a whole-transcript hallucination phrase set checked during weak-audio filtering so the phrase is rejected instead of being forwarded to the LLM. Strong audio keeps the phrase intact for genuine dictation. Co-authored-by: Cursor <cursoragent@cursor.com>
IchenDEV
marked this pull request as ready for review
September 11, 2026 08:07
- Add generic MLXSTTEngine that loads any mlx-audio-swift STT model via STTGenerationModel protocol (FireRedASR2, Mega-ASR, etc.) - Add .firered and .megaASR speech engine types, replacing .mimo - Register FireRedASR2-AED-mlx (~4.6 GB) and Mega-ASR-6bit (~2.0 GB) in ModelCatalogASR with required files and download estimates - Integrate MLXSTTEngine into SpeechEngineProvider and VoicePipeline - Update engine picker UI and model management sections - Add en/zh-Hans localization strings for new engines Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
Fix: silence (no speech) no longer produces "Do anything" output. When Whisper hallucinates this phrase on weak audio, the transcription sanitizer now rejects it before it reaches the LLM.
SDLC bundle and risk
Verification
swift test— all 9TranscriptionSanitizerTestspass, including 2 new casesbash scripts/ci-basic-checks.shResidual risk and rollback
None. The filter only activates when
hasWeakSpeechEvidenceis true, so genuine "Do anything" dictation with clear audio is unaffected.Reviewer focus
weakAudioWholeTranscriptHallucinationsset inTranscriptionSanitizer.swift— verify the phrase list is appropriate and the guard-only-on-weak-audio policy is correct.testDropsDoAnythingHallucinationOnWeakAudioandtestKeepsDoAnythingWhenAudioIsStrongcover both sides.Made with Cursor