Skip to content
This repository was archived by the owner on Sep 8, 2026. It is now read-only.
This repository was archived by the owner on Sep 8, 2026. It is now read-only.

agent memory: structured truncated tool_result on the wire (not Tool: crumbs) #549

Description

@btipling

Parent

#548 (agent session architecture)

Problem

formatPromptWithHistory drops every tool_run row (lib/sessionStore.ts).
Those cards are display-only (plan #345). The model already saw tool results
this turn; on the next turn it gets only user/assistant/system/error
prose, flattened into one { prompt } string.

Operators see: the agent re-reads the same files, forgets what it grepped,
never builds session memory, and talks like a new hire each send.

We persist up to 8 MiB of those cards in Blob/localStorage and then
omit them from inference.

Goal (locked direction)

Put truncated tool_result on the wire as structured messages — not
Tool: … lines stuffed into formatPromptWithHistory.

Peer harnesses (OpenCode / Pi / Oh My Pi) send a real message array:

user / assistant + tool_use / tool_result (thinking only if the
provider wants it). That is how the next turn still knows what was read.

Invincible today: single-shot { prompt } + display-only tool_run.

Success: after a turn that read_files / greps / change_dirs, the next
user message can be “use what you already found” and the model does not
rediscover the tree from zero.

How (directional — lock in create-plan)

  • Host/server keep a model-facing trace (call id, name, args, truncated
    result, ok) distinct from the Wasm paint card (tool_run stays
    display-primary).
  • Default send: bounded result (paths + status + short excerpt). OpenCode
    ~2k lines / 50 KB; Pi ~2k chars at compact time. Do not re-send
    TOOL_RUN_PREVIEW_MAX_CHARS (100k) per item.
  • Prefer AI SDK / Gateway native tool-result messages over a flattened
    transcript. Flattening is a fallback if a provider cannot take the array —
    still truncated, still paired (never a call without its result).
  • Abort/cancel: persist committed tools the same way we persist change_dir
    (honest, not a hallucinated summary of a cancelled mid-tool).

Constraints (lock)

Peer harness notes

Idea Take
Native tool_use / tool_result in the messages array This issue
Truncate tool output; do not omit This issue
Never split a call from its result at a cut #552
Files touched listed in the compact summary #552
Skills catalog vs full bodies #557

Non-goals

Suggested next

create-plan against #548 A1 before coding.

Orrery (2026-08-15)

  • Hygiene: few summary lines in the model-facing result; verbose exec/test output stays on the sandbox disk (file path in the result). Same Anthropic “parallel Claudes” lesson as their exec tool. Caps for that excerpt land in the A1 create-plan table — do not default to TOOL_RUN_PREVIEW_MAX_CHARS.
  • Assembly order: model → tools → system → history is locked before the request is built (inference: cache-stable two-block system prompt (don’t bust the KV prefix) #558). Structured messages are the history part, not a second system dump.

Related tools (not this board)

Tool-side shaping (so there is something honest to fold):

  • #562 first-class search
  • #563 windowed read_file
  • #564 stale str_replace window
  • #565 exec summary + disk log

Those issues are sandbox/tools, not agent-session. This issue remains structured tool_result on the wire.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions