Skip to content

Prompt-surface judgment rewrite: acceptance = blog-principles conformance per file (size is reporting-only) #1280

Description

@waleedkadous

Problem

The always-on prompt surface a builder consumes is ~21,900 served words (measured: 1252-word-baseline.md), dominated by process recipes: porch phase tasks (11,430w across a 10-iteration project), protocol.md inlined into every spawn (3,703w), CLAUDE.md how-to prose, and consult-type preambles. Spec 1252's measurement proved deduplication alone yields only −7%: the surface is not duplicated, it is over-instructed.

The new rules of context engineering for Claude-5-generation models reports >80% of Claude Code's system prompt was deleted with no measurable performance loss, by replacing rules with judgment, deleting worst-case padding, designing interfaces instead of examples, and progressive disclosure. That operation — content deletion on judgment-trust grounds — was explicitly a Non-goal of Spec 1252 and has never been attempted here.

Goal (SUPERSEDED — see Amendment below)

Rewrite the always-on instruction surface for frontier-model consumers: >50% reduction in served always-on words, measured with 1252's committed measurement script (counts served/expanded words; phantom-savings-proof).

Attack order by size: phase-task prompts → protocol.md spawn inlining (keep the state machine, gates, artifact contracts; delete narrative process) → CLAUDE.md residual how-tos (→ progressive-disclosure skills) → consult-type prompts (rubric + verdict contract; delete process prose).

Baked Decisions

  • All prompt consumers are frontier models (Claude 5, GPT 5.6, Gemini 3.6 class). No weak-model tier, no fallback scaffolding variant, no tiering mechanism. One form.
  • Scar rules are exempt and verbatim — the eight compressed canonicals developed in Spec 1252 Phase 5 (six repo rules + shellper verified-orphan + Tower-restart permission) ship with the rewrite; the registry/enforcement concept from 1252 is rebuilt fit-for-purpose around the post-shrink surface, not before it.
  • Validation is A/B, not observational: same issues executed by builders on old vs new prompts, compared on outcomes (gate friction, review rounds, correctness). Spec 1252's M12 established that observational baselines (n=17) can only detect large regressions — insufficient at deletion scale. The A/B design is a first-class spec section.
  • Spec must define a rollback story (prompt surfaces are files; reverting is cheap — say so concretely).

Prior art (required reading for the spec phase)

Protocol

SPIR — spec phase must produce: the per-surface cut plan with word targets, the A/B eval design, and the scar-rule carriage plan, before any rewriting.

AMENDMENT — 2026-08-01 (owner redirect, supersedes the Goal above)

The owner has redirected the acceptance model; reviewers should judge against THIS, not the original Goal:

  1. The goal is NOT a particular size. Acceptance = conformance to the blog post's principles (P1–P7, quoted verbatim in the spec, each restated as a per-file question answerable from a diff). A file that is principle-conformant at MORE words passes.
  2. Size becomes reporting-only. The measurement instrument (corrected per the spec's M0 — the original committed script measured a dead prompt tree) still runs before/after and reports honestly, including deletion-vs-relocation separation, but no criterion passes or fails on a word count. The original '>50% reduction measured with 1252's committed script' target is superseded on both counts.
  3. Architect personal inspection is a mandatory gate mechanic: the architect inspects the old-vs-new diff of every changed file (~66 distinct decisions), per-phase, against a manifest that fails the phase if it omits a changed file.
  4. Unchanged from the charter: frontier-fleet Baked Decisions, scar-rules exemption (resolved against the blog's P7 in the spec: the blog deletes guardrails against bad output; scar rules guard irreversible acts), and mandatory A/B behavioral validation.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/cross-cuttingTouches multiple areas — needs coordinated handling

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions