Skip to content

fix(protocol): give every announced responses item a unique output_index and its own order - #207

Merged
argszero merged 1 commit into
mainfrom
fix/responses-output-item-index
Sep 13, 2026
Merged

argszero merged 1 commit into
mainfrom
fix/responses-output-item-index

Conversation

@argszero

Copy link
Copy Markdown
Owner

Summary

The streamed openai_chat → responses translator gave two different items the same output_index, and its terminal object listed the items in an order contrary to its own incremental events.

ResponsesStreamState announced the message item at a hardcoded output_index: 0 while tool_output_index also started at 0, so a response carrying text + a tool call announced two items at index 0. A client that tracks items by output_index — the key the Responses protocol gives for exactly this purpose — has one item overwrite the other, usually the function_call an agent client then waits forever for. The terminal response.output was composed by the whole-body helper (function_call first, message last), which contradicts the order in which the stream had announced them.

Measured on the streamed path (text + two tool calls, after the fix):

item announcements (id, output_index): [("msg_c1", 0), ("fc_call_1", 1), ("fc_call_2", 2)]

and with the colliding indexing reinstated (positive control below):

assertion `left == right` failed: 索引必须互不相同:[0, 0, 1]

This survived until now because the shape-parity test added in #206 only compared the streamed terminal object with the whole-body one — never the stream's own events with its terminal.

Related Issue

Changes

  • 功能/修复说明 — 每个被宣布的条目从同一个单调计数器拿到 output_index(message 条目先,随后各 function_call),因此索引唯一、单调,且就是该条目在 output 数组里的位置;tool_output_index 那套 +1 / saturating_sub(1) 算术删掉
  • 功能/修复说明 — message 条目改为第一次拿到文本时宣布(此前在初始化块里无条件宣布):于是「流宣布过的条目」与「终局列出的条目」是同一个集合,没有文本就没有这个条目(与整包路径一致)
  • 功能/修复说明 — 终局 output 按宣布顺序拼装(流自己的权威);整包路径的顺序不变(function_call 在前),形状仍同源
  • 功能/修复说明 — 每个 function_call 的 response.function_call_arguments.delta 用它自己的索引(旧算术对「同一 chunk 里第二个工具调用」和「同一 call id 的续块」都会给出错的索引)
  • 功能/修复说明 — 条目形状拆成按条目的唯一真源 protocol::openai_chat_message_item / protocol::openai_chat_tool_call_item,两条路径共用;openai_chat_message_to_responses_output 改为调用它们(整包行为不变)
  • 涉及配置/数据结构的改动已同步示例文件 — 不涉及(无 config / schema / ui/ / i18n 改动)

Tests

  • cargo test 全部通过 — 199 passed(原 196,+3)
  • cargo fmt --check 通过
  • 新增/更新了单元测试(如适用)
    • responses_item_indices_are_unique_and_match_the_announced_order — 文本 + 两个工具调用 ⇒ 三个条目的索引多重集为 {0,1,2} 且按宣布顺序单调递增;终局 output 的 id 序列 == 宣布顺序;每个参数增量的 output_index == 它自己条目的索引(含同一 chunk 内两个工具调用、同一 call id 的续块)
    • responses_items_follow_the_announcement_order_when_tool_calls_come_first — 顺序规则是宣布顺序,不是「message 永远第一」
    • responses_stream_without_content_announces_no_items — 不变式的边界:无内容 ⇒ 不宣布条目、终局 output 为空
    • responses_terminal_shape_agrees_between_stream_and_whole_body — 改为按身份(message / call_id)比对条目形状而不是按下标,并把两条路径各自的顺序规则写成显式断言(流式=宣布顺序,整包=工具调用在前)
    • responses_terminal_event_is_emitted_once_with_the_full_response_object — 终局顺序期望随规则更新为 [message, function_call]

A/B(红前绿后,每条变异自带期望红集,还原在 finally,mtime 已刷新)

变异 期望红集 实测
索引撞车(message 条目写死 0) 索引唯一性 + 工具调用先到 197 passed / 2 red(= 期望)
终局按整包顺序拼装 唯一性 + 终局事件测试 + parity 196 passed / 3 red(= 期望)
参数增量用旧 next_output_index-1 索引唯一性 198 passed / 1 red(= 期望)

还原后 199 passed 全绿,且两个源文件 md5 与注入前逐字节相同。

Checklist

  • 分支命名符合约定(fix/responses-output-item-index)
  • Commit message 使用 Conventional Commits 格式(fix(protocol): ...)
  • 单一职责,改动最小化 — 只动 src/sse.rs + src/protocol.rs 的 responses 流式条目身份/顺序;status/截断语义(face ④)与初始化门(F1,可达性未证)刻意不动

…dex and its own order

The streamed openai_chat -> responses translator announced the message item at a
hardcoded `output_index: 0` while `tool_output_index` also started at 0, so a
response carrying text + a tool call announced two different items at the same
index; a client keying items by `output_index` (the key the protocol gives for
exactly this) has one item overwrite the other - typically the `function_call`
that an agent client then waits for. The terminal `response.output` also followed
the whole-body composition order (`function_call` first), contradicting the
stream's own announcement order.

- one monotonic counter allocates every announced item's index (message item,
  then each function_call), so an item's index is also its position in `output`
- the message item is announced when the first text arrives, so "the items the
  stream announced" and "the items the terminal lists" are the same set
- each function_call's argument deltas carry that item's own index (the old
  `saturating_sub(1)` arithmetic was wrong for a second tool call in the same
  chunk and for a repeated call id)
- the item shapes become per-item single sources
  (`protocol::openai_chat_message_item` / `openai_chat_tool_call_item`) called by
  both paths; the whole-body path keeps its own order (`function_call` first)
- tests compare the two paths by item identity instead of by position, and state
  each path's ordering rule explicitly

199 passed (was 196), fmt/clippy clean.
@argszero
argszero merged commit 254eacf into main Sep 13, 2026
1 check passed
@argszero
argszero deleted the fix/responses-output-item-index branch September 13, 2026 12:05
argszero added a commit that referenced this pull request Sep 14, 2026
…pseek-flash (#233)

The DeepSeek docs (https://api-docs.deepseek.com/zh-cn/, checked 2026-09-14)
now name the model `deepseek-flash`; the old `deepseek-v4-flash` and
`deepseek-v4-flash-vision-exp` are retired (requests to those names still
answer, but DeepSeek routes them to DeepSeek-V4.1-Flash and bills Flash
prices). `deepseek-v4-pro` is unchanged.

Why this mattered: `POST /api/sharings` validates that a share is priceable
with `SELECT 1 FROM models WHERE provider = ?1 AND model = ?2`, so listing a
share of the new model was rejected because the name was in neither
`config/config.example.toml` nor the seeded `models` table.

Rename only - deliberately no alias handling. The retired names are dead and
upstream-routed, and "a deleted model bills 0" is the documented behaviour
(`admin.models.sub`), so dropping them introduces no new defect class.

- config/config.example.toml: the flash row becomes `deepseek-flash` with the
  official prices (idle cache-hit 0.02 / uncached 1.0 / output 4.0 CNY per 1M;
  peak 0.04 / 2.0 / 8.0), `vision = true` (V4.1-Flash is natively multimodal)
  and context 1M / max output 384K; the separate `deepseek-v4-flash-vision-exp`
  entry is folded into it, so the catalogue is 13 models.
- ui/js/data.js: the MODELS / MARKET mirrors follow (13 models, 7 on sale).
- src/catalog_gate.rs: MODEL_COUNT 14 -> 13, KNOWN_MODEL / KNOWN_INPUT /
  KNOWN_OUTPUT updated to the new name and prices.
- src/config.rs, src/db.rs, src/routes/sharing.rs: fixtures and catalog
  assertions updated; `config.rs` now also asserts that the retired names are
  gone and that `deepseek-flash` is a vision model.
- docs/plan-api-matrix.md: stop hardcoding the catalogue size (point at the
  single source of truth instead).
- ui/index.html: data.js cache-bust.

Deployment note: `seed_models` is a full sync - `config`'s `[[models]]` is
authoritative and rows absent from it are deleted at startup - so a deployment
must apply the same rename to its own (gitignored) config.toml and restart.

Tests: `cargo test` 239 passed / 0 failed (the count is unchanged - no test was
added or removed, the catalogue assertions were updated in place);
`cargo fmt --check` exit 0. Clippy is clean on CI's stable toolchain (main's
last run: "Run clippy" success). The only complete toolchain installable in
this sandbox is rustc/clippy 1.95.0, whose clippy additionally flags one
pre-existing `collapsible_match` in `src/protocol.rs:662` - a file this change
does not touch (it landed in #207 and its CI has been green since), so it is
left alone rather than fixed opportunistically here.
argszero added a commit that referenced this pull request Sep 14, 2026
…ggregates (#234)

Host report (rant 2026-09-14T16:51:14): monthly aggregate queries take 4~10s on
the dev deployment.

The root cause is not NAS throughput (measured: a 4KB hot read on the NAS is
1.9µs, on par with a local disk) but **how many pages a query touches**:
`strftime('%Y-%m', time) = strftime('%Y-%m', 'now')` wraps the indexed column in
a function, so the index is unusable and SQLite falls back to a full scan. Same
db, same query: strftime 3845ms -> range predicate 2321ms -> covering index
119ms (18-32x).

- 12 production predicates rewritten to **closed** ranges (equivalent to the old
  form and index-usable): wallet.rs (month consume / month earn / dashboard
  month-by-type + the 7-day series JOIN), ops.rs (month_calls / month_in /
  month_out + today-by-hour), admin.rs (per-member, per-model, per-dept),
  org.rs (per-dept month_cost). The closed upper bound matters: the lower bound
  alone would admit future-month rows (asserted).
- src/db.rs: SCHEMA_VERSION 13 -> 14 with four new indexes -
  `transactions(user_id, time, type, pts)` and `transactions(time, type, pts)`
  for the with/without-user_id month aggregates, plus `usage_records(time)` and
  `usage_records(user_id, time)` (that table previously had zero indexes).
- New src/perf_gate.rs (test-only module): fails the suite if a date function
  wraps a time column in production again, with a positive control on the
  rewrites and a self-test proving the detector actually fires.
- The wallet test that still uses the old form is kept as the independent spec
  oracle - the rewrite must keep satisfying the old semantics.

Tests: `cargo test` 239 -> 245 passed / 0 failed (4 perf_gate + 2 db).
`cargo fmt --check` exit 0. Clippy is clean on CI's stable toolchain; the
sandbox's only complete toolchain (1.95.0) additionally reports one pre-existing
`collapsible_match` in `src/protocol.rs:662` (out of this diff, landed in #207),
left untouched.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant