Skip to content

Filter Whisper "Do anything" hallucination on weak audio - #92

Merged
IchenDEV merged 3 commits into
mainfrom
cursor/fix-do-anything-hallucination
Sep 17, 2026
Merged

IchenDEV merged 3 commits into
mainfrom
cursor/fix-do-anything-hallucination

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

Outcome

Fix: silence (no speech) no longer produces "Do anything" output. When Whisper hallucinates this phrase on weak audio, the transcription sanitizer now rejects it before it reaches the LLM.

SDLC bundle and risk

  • Bundle: N/A (trivial single-phrase filter addition)
  • Risk: trivial
  • Human decisions still required: none

Verification

  • swift test — all 9 TranscriptionSanitizerTests pass, including 2 new cases
  • bash scripts/ci-basic-checks.sh
  • Real-device QA: confirm silence no longer outputs "Do anything"

Residual risk and rollback

None. The filter only activates when hasWeakSpeechEvidence is true, so genuine "Do anything" dictation with clear audio is unaffected.

Reviewer focus

  • weakAudioWholeTranscriptHallucinations set in TranscriptionSanitizer.swift — verify the phrase list is appropriate and the guard-only-on-weak-audio policy is correct.
  • New tests testDropsDoAnythingHallucinationOnWeakAudio and testKeepsDoAnythingWhenAudioIsStrong cover both sides.

Made with Cursor

When no speech is present but background noise passes the energy
threshold, Whisper may hallucinate "Do anything". Add a
whole-transcript hallucination phrase set checked during weak-audio
filtering so the phrase is rejected instead of being forwarded to the
LLM. Strong audio keeps the phrase intact for genuine dictation.

Co-authored-by: Cursor <cursoragent@cursor.com>
@IchenDEV
IchenDEV marked this pull request as ready for review September 11, 2026 08:07
IchenDEV and others added 2 commits September 17, 2026 15:11
- Add generic MLXSTTEngine that loads any mlx-audio-swift STT model
  via STTGenerationModel protocol (FireRedASR2, Mega-ASR, etc.)
- Add .firered and .megaASR speech engine types, replacing .mimo
- Register FireRedASR2-AED-mlx (~4.6 GB) and Mega-ASR-6bit (~2.0 GB)
  in ModelCatalogASR with required files and download estimates
- Integrate MLXSTTEngine into SpeechEngineProvider and VoicePipeline
- Update engine picker UI and model management sections
- Add en/zh-Hans localization strings for new engines

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@IchenDEV
IchenDEV merged commit d1454cd into main Sep 17, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant