Skip to content

fix(server): decode streamed SSE text from the cumulative id list - #105

Open
ahrazzle wants to merge 2 commits into
Edge0-AI:mainfrom
ahrazzle:fix/sse-incremental-decode
Open

ahrazzle wants to merge 2 commits into
Edge0-AI:mainfrom
ahrazzle:fix/sse-incremental-decode

Conversation

@ahrazzle

@ahrazzle ahrazzle commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

What changed

Streamed SSE responses now decode the full id list cumulatively and emit only the newly completed suffix, holding back partial multi-byte characters until they resolve.

Root cause

The streaming path decoded each token id in isolation (python/src/edge0/server/app.py:82). A byte-level BPE token carrying part of a multi-byte character streamed as U+FFFD while the non-streaming path returned correct text.

Changes

  • Added _incremental_suffix helper to extract the new suffix from a cumulative decode
  • Modified _chat_stream to maintain a running list of seen ids
  • Flush held-back partials at end of generation so streamed equals non-streaming text

Rebase needed

The multi-platform layout refactor (cad29a1) moved the touched files to python/src/edge0/server/app.py and python/tests/test_server.py. This branch predates it and no longer applies (mergeable: conflicting). Rebase onto main first. The diff carries over unchanged: a local rebase onto feafe31 applied with rename detection and zero content conflicts.

Validation

Live edge0-8b server run (stdlib transport), prompt asking the model to echo 🌊🏖️🦀🍣:

path text
base, streaming 9 U+FFFD, no emoji
head, streaming 🌊🏖️🦀🍣 (0 U+FFFD)

The non-streaming path decodes the whole id list and was already correct.

Test: python/tests/test_server.py::test_chat_stream_incremental_decode_of_split_characters

  • Red on base (base tree + head test file)
  • Green on head (66 passed, 1 skipped)
  • Re-verified 2026-10-03 on the rebased tree: test passes (34 ids, 14 content deltas, streamed text == non-streamed text, no U+FFFD). The same test fails on unpatched base

Limits

Common CJK text in the shipped vocabulary is not affected. Corruption needs characters that split across byte tokens (emoji, rare CJK, some Korean).

O(n²) cumulative re-decode per request. This is the minimal fix. Throughput was not measured.

Automated posting by agentic team with human oversight.

Rebased onto main: upstream moved the tree under python/; the fix
applies unchanged to python/src/edge0/server/app.py.
@ahrazzle
ahrazzle force-pushed the fix/sse-incremental-decode branch from 72586c0 to 690724a Compare October 3, 2026 18:26
Resolve the conflict in _chat_stream: keep main's tool-call
buffering path (buffer the full response, parse once, emit a
single delta with tool_calls and finish_reason=tool_calls) and
layer the incremental SSE decode on top of the non-buffering
path -- on_token appends to seen_ids, decodes the cumulative id
list, and emits only the new suffix via _incremental_suffix,
with a final flush so concatenated deltas equal the non-streaming
text exactly.

ToolCallTok in the tests now decodes prefix-stably (one fixed
piece per id, like a real byte-level BPE tokenizer) so the
no-tools regression guard still observes one immediate delta per
token under cumulative decoding.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant