test(sse): cover the split boundary the existing test claims to cover - #274
Merged
Merged
Conversation
`sse_helpers_split_across_chunks` declares that a UTF-8 character straddles the chunk boundary, but its boundary falls after an ASCII `"`: the multi-byte character sits wholly inside the second chunk. A naive implementation (per-chunk `String::from_utf8_lossy`, no remainder carried) therefore produces byte-identical output on that input, so the test cannot fail for the bug it claims to guard. Keep the test and its input, correct the comment to describe what it actually exercises, and add two tests that split a character inside its UTF-8 sequence: - `sse_helpers_split_inside_a_multibyte_char`: 3-byte and 4-byte characters split in every position, asserting the character survives, no U+FFFD appears and the buffer drains. - `sse_helpers_keep_an_incomplete_char_in_the_remainder`: the incomplete tail is withheld until completed; a genuinely invalid byte yields exactly one U+FFFD and parsing resumes. Both feed the character as raw bytes so the test does not share the decoder assumption it is meant to check. Test-only: no production code changes.
Owner
Author
|
@/Users/argszero/scm/github.com/argszero/aitokenpool/.emrg/sessions/emrg-evolution-aitokenpool-opensource-task/tmp/c2168-review.txt |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
sse_helpers_split_across_chunks(src/sse.rs:1628) is documented as covering "the event is split acrosstwo chunks and a UTF-8 character straddles the boundary". Its second chunk boundary, however, sits after
an ASCII
"— the multi-byte character sits wholly inside the second chunk. Measured: replacingappend_utf8_safewith a naive implementation (per-chunkString::from_utf8_lossy, no remainder carried)produces byte-identical output on that input, so the test cannot fail for the bug it claims to guard.
This corrects the comment to describe what the input actually exercises, and adds two tests that put a
character inside a UTF-8 sequence so the guard can fail.
This is not a bug fix. Production code is correct on every input measured (12 legs, listed below); the
defect is in the test's claim, not in the code it guards.
Related Issue
Changes
src/sse.rs(test module only):sse_helpers_split_across_chunks— kept, same input, comment corrected to say what it actuallycovers (a chunk concatenation, not a straddling character).
sse_helpers_split_inside_a_multibyte_char(new, 6 legs) — 3-byte你split2+1/1+2/1+1+1and 4-byte
😀split1+3/2+2/3+1. Each leg asserts the character survives, that noU+FFFDappears, and that the buffer is drained.sse_helpers_keep_an_incomplete_char_in_the_remainder(new, 2 legs) — half a character is withheld inremainderand emitted once completed; a genuinely invalid byte yields exactly oneU+FFFDandparsing continues (
A\u{FFFD}B). White-box on purpose: this is the property a naive implementationloses.
[0xE4, 0xBD, 0xA0]/[0xF0, 0x9F, 0x98, 0x80]as bytes, notas string literals — a literal would make the test share the decoder assumption it is supposed to check.
Why the new legs can fail (measured, not asserted)
from_utf8_lossy)…{\"a\":\"|你\"}\n\n)event: x␊data: {"a":"你"}[E4 BD]|[A0]你��(twoU+FFFD)Compiled A/B against a verbatim copy of the live helpers (
rustc --test, four arms, all as declared):The third row is the point of this PR: the existing test passes against a broken implementation.
Re-measured in-tree on this branch (not on the standalone sketch): the whole body of
append_utf8_safewassubstituted in place with the naive implementation,
cargo test sse::was run, and the file was then restoredand checked byte-identical (
md5 536da10ed48033ce77600cb779b60567before and after;git diff --statunchanged).Same verdict: retained test
ok, both new testsFAILED(panics atsrc/sse.rs:1659and:1677).Honest boundary
boundary, invalid byte, CRLF/LF mixed framing,
strip_sse_fieldin all four shapes) are correct on thecurrent code.
append_utf8_safe'sremainder.len() > 3branch is unreachable — a valid incomplete UTF-8 prefix is atmost 3 bytes, so
error_len() == Nonecan never leave more. It is a defensive branch, left untouched.append_utf8_safe/take_sse_blockhave no consumer outsidesrc/sse.rs.Tests
cargo testpasses — 310 passed / 0 failed measured on this branch (mainis 308; the delta isexactly the two tests added here, and
src/sse.rs's#[test]count goes 19 → 21)cargo fmt --checkpassescargo clippy --all-targets -- -D warningspassespasses on that same naive implementation (A/B above, 4 arms, each arm's expectation declared up front)
fn sse_helpers_split_across_chunks1 → 1 (retained,not replaced);
fn sse_helpers_split_inside_a_multibyte_char0 → 1;fn sse_helpers_keep_an_incomplete_char_in_the_remainder0 → 1;#[test]insrc/sse.rs19 → 21;the byte anchor
0xE4, 0xBD, 0xA00 → 1;FFFDoccurrences 1 → 3Checklist
fix/…)test(sse): …)main)