fix: reject near-total text replacements in correction learning - #93
Merged
Merged
Conversation
When no speech is present but background noise passes the energy threshold, Whisper may hallucinate "Do anything". Add a whole-transcript hallucination phrase set checked during weak-audio filtering so the phrase is rejected instead of being forwarded to the LLM. Strong audio keeps the phrase intact for genuine dictation. Co-authored-by: Cursor <cursoragent@cursor.com>
CorrectionObservationPolicy now requires at least 25% overlap (common prefix + suffix) for texts longer than 4 characters. This prevents accidental learning when the target app replaces the entire text field content (e.g. UI navigation changing the field to 'Follow up'), which was being misclassified as a user correction. Co-authored-by: Cursor <cursoragent@cursor.com>
IchenDEV
marked this pull request as ready for review
September 17, 2026 05:13
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
Fixes accidental correction learning when the target app replaces the entire text field content (e.g. UI navigation changing the field to "Follow up"). Previously,
CorrectionObservationPolicyonly rejected appends and prepends but allowed near-total replacements to pass through as valid corrections.SDLC bundle and risk
Verification
swift test— all 8CorrectionCandidateClassifierTestspass, including 2 new testsbash scripts/ci-basic-checks.shpython3 scripts/sdlc.py validate --worktreeResidual risk and rollback
Minimal. The new guard only applies to texts >4 characters and requires ≥25% overlap (common prefix + suffix ≥
max(2, length/4)). Legitimate corrections that change a small part of the text are unaffected. Revert is a single commit.Reviewer focus
The overlap threshold (
insertedCount / 4) — verify it balances false-positive rejection against genuine short-text corrections.Made with Cursor