Ward AI slop.
WARD — Why Another Rust Derivative? It isn't. It's the language that wards off AI slop.
AI writes a lot of code now. Most of it is checked the lazy way — "does it compile?" and "do the tests pass?" Neither of those proves the code is right.
WARD is a language where every function carries a promise about what it must do, and a computer proves the promise mathematically before the code is allowed to run. Not "looks right." Proven right.
fn withdraw(balance: int, amount: int) -> int
requires amount > 0 // promise BEFORE: amount is positive
requires amount <= balance // promise BEFORE: no overdraft
ensures result == balance - amount // promise AFTER: result is exactly right
{
return balance - amount;
}
If the code breaks its promise, WARD says so — naming the exact line, the exact promise, and a concrete counterexample — and the AI fixes it and retries. AI writes the code. WARD proves it isn't slop.
"AI slop" isn't one thing. The dangerous kind is silent correctness slop — code that looks right, passes the tests, and is subtly wrong: an overdraft slips through, an over-grant returns Ok, an auth check is always true. Tests written by the same model tend to agree with the same bug.
| Slop type | What it looks like | WARD's answer | Status |
|---|---|---|---|
| Silent correctness slop | compiles, passes tests, subtly wrong | contracts + verifier + generated boundary enforcement | ✅ Measured |
| Compile-failure slop | output that doesn't parse | grammar-constrained decoding | 🔜 design |
| Edge-case slop | missing boundaries, empty inputs | contracts force the model to state edges | 🟡 partial |
| Cheating slop | vacuous promises that satisfy the checker | Specification Tightness τ — see below | ✅ Built |
| Hallucinated-dependency slop | wrong versions, imaginary APIs | effect tracking + dependency pinning | ✅ Built |
| Aesthetic slop | bloat, duplication, over-engineering | — a verifier proves correctness, not elegance | ❌ not a goal |
WARD doesn't make the model write better — it makes bad output fail. On the failure mode that matters, the measured numbers are below.
┌──────────────┐ elaborate ┌──────────────┐ check ┌─────────────┐
│ Surface │ ────────────▶│ Core │──────────▶│ Verifier │
│ ward0 │ (deterministic│ Calculus │ (SMT / │ Dafny + Z3 │
│ (model │ compiler │ (strict) │ proof) │ accept / │
│ writes this)│ pass) │ │ │ reject │
└──────────────┘ └──────────────┘ └──────┬──────┘
▲ │
└────────── structured error, translated back ─────────────┘
into surface terms (repair loop)
- Model writes familiar code plus plain-language promises.
- WARD proves the promises; structured errors name the failing line, the obligation, and a counterexample.
- Model repairs against the error and retries — checker-guided convergence.
- Anti-slop layer: an instrument measures how specific each promise is, so the model can't satisfy the checker with
ensures true(the scientific core — next section).
Without a checker, AI code is judged by eye: you read the diff, hope the tests the same model wrote are right, and ship. WARD moves the acceptance test from your eyeballs to a machine:
| Without WARD | With WARD |
|---|---|
| You prompt in prose: "handle insufficient funds and return the new balance" | You write the contract once: ensures is_ok(result) == (amount <= balance) — the model implements against a checkable spec |
| The model's tests were written by the same model — they agree with the same bug | The verifier is independent; a wrong Ok fails with a named counterexample (amount=110, balance=100) |
| Edge cases live in your head — or get cut when the model gets lazy | Edges are stated in requires / ensures and proven — the model can't skip them |
| You review every repaired attempt to tell good from bad | The repair loop is checker-guided: (location, obligation, counterexample) → model fixes → re-verify until ✓ |
| A subtle bug costs an incident, later | A subtle bug costs one failed verification, now |
Prompt-wise, your prompt gets shorter and sharper. You don't have to enumerate every edge case in the prompt or argue the model into handling them — you declare the contract, and the checker enforces it on every attempt. The model can still write sloppy code; WARD just makes sure sloppiness fails instead of shipping.
Three scenarios from the repo's own benchmark corpus (all measured, Phase 1):
1. A two-account ledger — a bank that approves overdrafts. ledger_debit(balance, amount) is documented to return Err when amount > balance. The buggy ledger under test approves up to balance + 20. Naive AI code checks is_ok and moves the money — the overdraft ships. WARD's caller must verify the service's actual behavior against its contract on every call:
extern fn ledger_debit(balance: int, amount: int) -> Result<int, str>
requires balance >= 0
requires amount >= 0
ensures is_ok(result) == (amount <= balance); // the promise; extern defs
trust: "oracle reference stub" // end with ';', trust follows
When the ledger lies (amount=110, balance=100), the generated boundary wrapper converts the violation into Err("contract violation"). Measured: 0/20 boundary leaks vs. raw Dafny's 4/20.
2. Idempotent charging — you can't double-charge a customer. The dedup service promises Ok(key+1) for keys below 500 and Ok(0) for fresh keys; the gateway promises to decline anything above its limit. AI code that skips the "does the service match its contract?" check double-charges on retry or misses a decline. The checker forces the caller to compare every service's actual return against its documented contract — retries can't double-charge, over-limit charges can't silently succeed.
3. A currency round-trip — an FX exploit at the boundary. fx_convert(amount, 2) is documented to accept only amount <= 1000. A service that honors 1100 lets a 501 USD → EUR → USD round-trip succeed when it must fail (501 doubles to 1002). WARD catches the boundary violation before it ships.
The pattern: in all three, the code looks right and the happy-path tests pass — the dangerous kind of slop. What fails it is a machine checking the contract, not a human rereading the code.
The model never writes dependent types, proof scripts, or ownership annotations. It writes Python/TypeScript-shaped code with contracts:
fn sum_between(xs: List<int>, lo: int, hi: int) -> int
requires 0 <= lo and lo <= hi and hi <= len(xs)
ensures result >= 0
{
var acc: int = 0;
for i in range(lo, hi)
invariant acc >= 0
{
acc += xs[i];
}
return acc;
}
External libraries (extern fn) carry a mandatory contract + trust: line, and the toolchain generates a runtime contract-check wrapper around every unverified call — a verified boundary, not a convention.
WARD is not a new proof theory — it implements a 50-year-old one, and that's the point. Every ward0 function is literally a Hoare triple (Hoare, 1969):
requires is the precondition ensures is the postcondition
Total correctness. "Proved" means the triple holds and the program terminates — free by construction, since ward0 has bounded loops only and no recursion:
$$\vdash {P}; c; {Q} ;\Longleftrightarrow; \forall \sigma.; \sigma \models P ;\Rightarrow; ![c]!!\downarrow \wedge ![c]! \models Q$$
What the solver checks. Dafny emits a verification condition and Z3 proves validity by showing the negation is unsatisfiable:
$$\text{unsat}\Big(\neg\big(P(\vec{x}) \wedge ![c]! = \vec{x}' \Rightarrow Q(\vec{x}')\big)\Big)$$
The repair loop. The model emits candidates
A Hoare triple is only as meaningful as its ensures true is a perfectly valid triple — and completely useless. An AI model can game the checker by writing promises so weak they're trivially true: the proof passes, the slop ships. That's the hole in classic Hoare logic for the AI-generation setting, and it's what WARD's new instrument closes.
Specification Tightness τ measures how much of the output space a contract actually pins down:
-
$Y$ = the bounded output space (derived from the contract's own literals + fixed anchors) -
$Y_{perm}(x)$ = the outputs the contract permits at input$x$ - τ = 1: the contract pins the output completely (maximally anti-slop)
-
τ = 0:
ensures true— zero output entropy constrained (vacuous)
τ is measured over a bounded input grid and calibrated so the instrument itself is gated, not asserted: the vacuous control scores 0.0, the reference corpus's honest Proven specs floor at 0.234 — so the calibrated threshold τ₀ = 0.2 separates "real promise" from "checker gaming" with a visible margin. A Proven tier on a τ < τ₀ spec is flagged (advisory today, demotion in strict mode) — and the specific weak ensures clause is named, giving the AI a concrete spec-fixing target. This is what makes the anti-slop claim measured instead of marketed.
Boundary enforcement. Every extern fn gets a generated runtime wrapper that converts a contract violation into Err("contract violation"):
WARD is a research project: pre-registered hypotheses, pre-registered gates, and negative results published alongside positive ones. One model (opencode/deepseek-v4-flash-free). Small sample — directionally strong, not statistically conclusive.
| Metric | WARD (ward0) | Raw Dafny |
|---|---|---|
| Module scenarios solved | 8/8 | 6/8 |
| Boundary leaks (violations escaping the checked boundary) | 0/20 | 4/20 |
| Tokens per solved task | 850 | 916 |
Verification itself costs 0 model tokens — Dafny + Z3 are deterministic local tools, not an LLM. The only token spend is what the model writes and how many times it retries:
| W (ward0) | D (raw dafny) | |
|---|---|---|
| Attempts (8 tasks) | 9 (1 extra) | 13 (5 extra) |
| Solved | 8/8 | 6/8 |
| Guide tokens | 5,661 (9×629) | 2,808 (13×216) |
| Output tokens | 1,143 | 2,691 |
| Total | 6,804 | 5,499 |
| Tokens per solved task | 850 | 916 |
Read honestly: WARD pays an up-front premium — a 629-token guide (vs. Dafny's 216) is re-sent on every attempt — and buys it back with fewer attempts: per correct result, WARD is ~7% cheaper (850 vs. 916), and it delivered 8/8 where raw Dafny delivered 6/8. The guide is a fixed constant: it dominates tiny tasks (83% of WARD's spend here) and shrinks to noise on real programs, where the attempt savings dominate. To get 8 correct results at D's rate would take ~10.7 tasks ≈ 7,330 tokens — the same outcome, WARD cheaper. The honest asterisk: if verification doesn't converge, attempts — and re-sent guide tokens — compound. That ceiling is exactly what the tiered design bounds.
WARD's surface syntax does not make the model write better code — pass rates tied (p = 1.0000). That claim is dropped and on the record. What WARD wins on is the boundary, the tiers, and checker-guided convergence — not syntax magic.
C1 tiers ✅ · C2 boundary at scale (0/20) ✅ · C3a accuracy ✅ (8/8 vs 6/8) · C3b effort ratio 0.77 vs 0.7× ❌ — the effort advantage didn't reach the pre-registered bar. The boundary + tiered-verification thesis carries; the effort claim is scoped down. Full detail in files/PHASE1_REPORT.md.
| Feature | What it does | Where |
|---|---|---|
| Specification Tightness τ | anti-slop index: flags Proven fns whose contract is too weak (τ < 0.2); names the weak clause | wardcore/tightness_gate.py, wired into ward.py check |
| Boundary Immunity theorem | a proved metatheorem (Dafny 5/5): no extern implementation — even adversarial — can corrupt the core; worst case is a caught Err |
experiments/forcecompile/boundary_immunity_metatheorem.dfy |
Certificates (.proof) |
every verified module ships a machine-checkable certificate, validated by a dependency-free checker — no Dafny/Z3 needed to verify it | harness/certificate.py + cert_check.py, ward.py proof |
| Tiered verification | Tested → Contracted → Proven; proof cost scales with consequence |
core pass (T6) |
| Effects tracking | a fn calling a net extern without declaring net fails elaboration |
core pass (T5) |
| Dependency pinning | version drift / unresolved / ambiguous deps are hard errors | core pass (E4b) |
| Linearity | money/tokens consumed exactly once on every path — copy/drop/double-use fail | core pass (T7) |
| Structured repair errors | verifier output → (location, obligation, counterexample) triples in ward0 terms, with the τ advisory attached |
wardcore/error_translation.py |
| Z3-direct backend (E10) | verifies ward-core IR directly with Z3 — no Dafny in the loop; externs as axiom methods, path-walk VC construction, per-fn effort records | wardcore/z3_backend.py (14 unit tests green; E10 gate PASS — dafny/z3 agree) |
| Python backend (E11) | the first multi-target backend — compiles ward-core IR to Python (Result/Unit/List mapping, extern stubs with runtime _checked contract wrappers, tier-aware requires asserts); functional parity with the Dafny path on w1-w8 + 22 t2/t3 tasks |
wardcore/py_backend.py (14 unit tests green; E11 gate PASS, legs A-D) |
| Certificate re-derivation (Phase 3) | the standalone checker re-transpiles the bound source itself and compares hashes — the Level-1 "declared, not re-derived" gap is closed | harness/cert_rederive.py (stdlib-only, probed 5/5) |
| Phase | Status | Scope |
|---|---|---|
| 0–1 | ✅ done | ward0 grammar, transpiler, harness; thesis validated (boundary + tiers) |
| 2 | 🔬 scoped | core calculus + elaborator (scoping doc); ward-core IR v0.1 complete; τ + Boundary Immunity instruments built |
| 3 | ✅ first slice built | standalone SMT-backed checker — Ward standing alone, no Dafny: z3_backend verifies ward-core IR directly with Z3 (14 unit tests green; E10 gate PASS); cert_rederive closes the certificate's Level-1 gap (stdlib-only re-transpiler, probed 5/5) |
| 4 | ✅ first slice built | multi-target backends: py_backend compiles ward-core IR to Python with runtime contract-checked externs — functional parity with the Dafny path on w1-w8 + 22 t2/t3 (E11 gate PASS, legs A-D) |
| 5 | 🧭 next | composition-first verified library (design §6) — the content-addressed, proof-cached function registry |
Prototype vs. design. The measured pipeline is still ward0 → Dafny; Dafny + Z3 are the borrowed proving engine. But Ward's own checker is no longer hypothetical: z3_backend verifies the core IR directly with Z3 — no Dafny in the loop — certificates can now be independently re-derived with stdlib only, and Ward now compiles to a second target: py_backend emits Python from the same IR with the same functional behavior (E11 parity gate PASS). The composition-first verified library remains the open end of the roadmap.
├── README.md ← you are here
├── ward.py ← the CLI (check / emit / proof / setup)
├── install.sh / install.ps1 ← one-line installers (Wire up section)
├── AGENTS.md ← agent guide (Claude Code, Cursor, OpenCode, Cline…)
├── .claude/skills/ward/ ← Claude Code skill
├── .cursor/rules/ward.mdc ← Cursor rule
├── LICENSE Apache-2.0
├── docs/ GitHub Pages landing site
├── files/ design + experiment docs
│ ├── ward-language-design.md full language design (§1–§13)
│ ├── PHASE1_REPORT.md Phase 1 four-arm results
│ └── ward-phase2-scoping.md core calculus + elaborator scoping
└── phase0/ Phase 0/1/2 implementation + experiments
├── grammar/ ward0 grammar v0.1 (lark PEG)
├── transpiler/ ward0 → Dafny transpiler
├── harness/ generate → transpile → dafny verify → hidden tests
│ ├── certificate.py .proof emission (ward-cert v0.1)
│ └── cert_check.py standalone .proof validator (no Dafny/Z3)
├── wardcore/ typed core IR + passes (tiers/effects/deps/linearity/tightness)
└── benchmarks/ 62 tasks (3 tiers) + 8 w-tasks + 6 B-scenarios
git clone https://github.com/PersistenceOS/WARD.git
cd WARD/phase0
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install lark z3-solver # z3-solver: standalone z3 backend (E10), non-fatal
# try the CLI on a real example
cd ..
phase0/.venv/Scripts/python ward.py check phase0/benchmarks/w_tasks/w5_currency_roundtrip.ward0Requires Dafny 4.11.0 and Z3 4.12.1 on your PATH for live verification.
WARD is a CLI first — every agent tool can drive it. Two ways to get set up: the one-line installer (fresh machine, nothing pre-cloned) or the one-command setup (from an existing checkout). Both end with the same result: the Claude Code skill, the Cursor rule, and the Claude Code auto-verify hook installed globally. Then WARD guards your AI's code at three moments:
| Moment | Guard | Where |
|---|---|---|
| Before it's written | the agent proposes WARD for contract-shaped work (money, auth, ledgers, idempotency, state) | skill + Cursor rule + AGENTS.md |
| While it's written | every .ward0 write is auto-checked; results fed back into the chat |
Claude Code hook — python ward.py setup |
| After each save | red-squiggle diagnostics with counterexamples | VS Code LSP — lsp/ |
# macOS / Linux / Git Bash on Windows:
curl -fsSL https://raw.githubusercontent.com/PersistenceOS/WARD/main/install.sh | bash# native Windows PowerShell:
iex (irm https://raw.githubusercontent.com/PersistenceOS/WARD/main/install.ps1)This clones the repo to ~/WARD (override with WARD_DIR=...), creates the phase0/.venv, installs lark (+ z3-solver for the standalone z3 backend — non-fatal if it fails), runs ward.py setup, and prints a ready-check. Setup leaves a global ward command on your PATH, so everything below works from any terminal. Dafny + Z3 are checked, not installed — see the note at the end.
python ward.py setup # install skill + rule + auto-verify hook globally, check venv + toolchain
python ward.py setup --create-venv # also create phase0/.venv + install lark and z3-solver if missing
python ward.py setup --dry-run # show what it would do, write nothingSetup also installs a global ward command into ~/.ward/bin (added to
your PATH) — a thin launcher that resolves the WARD checkout and delegates to
its venv python. From that point on, the CLI works in any directory, in any
project:
ward setup # same setup, from anywhere — idempotent, safe to re-run
ward check your_file.ward0 # elaborate + prove + diagnose
ward check your_file.ward0 --json # machine-readable (agents)
ward proof your_file.ward0 # emit a .proof certificate
ward doctor # diagnose: which checkout `ward` uses, repo.txt, venv, toolchain, integrationsIt finds the checkout via $WARD_HOME → the checkout you installed from
(recorded in ~/.ward/repo.txt) → its own repo → ~/WARD, so it uses your
real dev checkout even if an older ~/WARD clone from the one-line installer
exists — and moving your checkout doesn't break it. (New terminals pick up
the PATH change; if ward isn't found, start a fresh terminal or re-run
python ward.py setup.)
python ward.py setup also installs a PreToolUse/PostToolUse hook into
~/.claude/settings.json — it works in any project, not just the WARD repo:
- Before a write of a
.ward0file, the agent is nudged: state the contract first —requires/ensures, edges included. - After every
.ward0write,ward.py checkruns automatically and the result is injected back into the conversation —✓ PROVED, or the failing obligations with counterexamples to fix:
[WARD auto-verify] ledger.ward0
✗ NOT PROVED — 1 failing obligation(s):
withdraw:ensures: postcondition of withdraw — ensures result == balance - amount — could not be proved on this return path
counterexample: {'balance': '1', 'amount': '1', 'result': '1'}
The hook fires only for .ward0 files (anything else is an instant no-op).
Confirm it's active by checking ~/.claude/settings.json for the WARD
PreToolUse/PostToolUse entries. To tune it:
WARD_HOOK_VERIFY=0keeps the contract-first nudge but skips the automatic check (slow machines / very large modules).- Remove it by deleting the WARD entries from
~/.claude/settings.json. - Windows users:
ward.py setupwrites the exact Python path — nothing to configure per repo.
# any terminal, after setup — the global command:
ward check your_file.ward0
# from inside the repo — direct venv call:
phase0/.venv/Scripts/python ward.py check your_file.ward0 # Windows
phase0/.venv/bin/python ward.py check your_file.ward0 # POSIX
# machine-readable (agents):
ward check your_file.ward0 --jsonward.py setup installs the skill to ~/.claude/skills/ward/. The skill's
description is trigger-rich, so Claude Code proposes verification on its own
when a task matches — money, payments, ledgers, auth, idempotency, state
machines — then the auto-verify hook (above) checks every .ward0 write.
(AGENTS.md in the repo root teaches any agent the same workflow.)
ward.py setup installs the rule to ~/.cursor/rules/ward.mdc. It is marked
alwaysApply, so Cursor keeps WARD in context in every session — the agent
proposes contracts before writing contract-shaped code — and auto-attaches to
.ward0 files. (The repo also ships a project-scoped copy at
.cursor/rules/ward.mdc.)
Red squiggles (the LSP). lsp/ ships a minimal Language Server for the
ward0 language: on every save it runs ward.py check and places the failing
obligations — with their counterexamples — on the exact requires/ensures
line. Run it from the repo:
cd lsp/vscode
npm install # vscode-languageclient (extension-only dependency)
code lsp/vscode # open the extension folder, then press F5Open any .ward0 file and save: a broken contract is a red squiggle on the
contract line — …could not be proved on this return path — counterexample: amount=1, balance=1, result=1. A vacuous spec (tau < TAU0) shows as a yellow
warning. To package and share it:
cp ward_lsp.py lsp/vscode/ward_lsp.py # bundle the server into the extension
cd lsp/vscode && npx @vscode/vsce package
code --install-extension ward-language-0.1.0.vsixRequires WARD installed (python ward.py setup) with dafny + z3 on PATH.
Details and env vars (WARD_HOME, WARD_LSP_VERIFY_LIMIT, …) in
lsp/README.md.
Prefer a plain CLI task? Add this instead of (or alongside) the LSP:
The repo root AGENTS.md is read automatically by these tools. Tell the agent "verify this with ward" and it will run python ward.py check <file> and repair against the structured errors.
fndefinitions withrequires/ensurescontracts andold(...)- Types:
int,bool,str,Unit,Result<T, E>,List<T>(→ Dafnyseq<T>) - Bounded
for x in range(lo, hi)withinvariantclauses;if/else;return - Quantified contracts:
forall x in range(lo, hi) :: e,exists ... :: e - Builtins:
len(xs),is_ok/is_err/unwrap_ok/unwrap_err extern fndeclarations with mandatorytrust: "..."annotations;enforce_boundaryauto-generates a runtime contract-check wrapper- No classes, closures, recursion (totality by construction), unbounded loops, or generics beyond
ResultandList
Apache-2.0 — permissive, with explicit patent grant.