Web Skill Factory: evolving reusable, verified, code-native skills for web agents#54
Web Skill Factory: evolving reusable, verified, code-native skills for web agents#54DEM1TASSE wants to merge 165 commits into
Conversation
…tool
A built-in submodule turning solved tasks into reusable, executable code skills:
- skills/{library,retrieve,decide,gate,update,llm}: store / retrieve (relevance) /
decide (use·adapt·skip utility) / admission gate (gold|self_verify|none) /
evolve (incremental growth on existing library) — backend-agnostic via configure_llm
over webwright's own Model abstraction (no hardcoded gateway/key/path)
- tools/skill_use.py: solve-time tool (agent invokes like self_reflection/image_qa) ->
retrieve+decide -> JSON recommendation (use/adapt/skip + source path)
- python -m webwright.skills.update --manifest batch.json --library ./lib : batch growth
- tests/skills: 5 unit tests pass against the migrated module (logic == original)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…skill_use CLI - skills/prompt.with_skill_hint: prepend skill-library usage hint to task prompt (non-invasive; webwright merges system_template by replacement, so prompt-level is the clean way) - config/skill_mode.yaml: optional overlay doc + step budget for skill-reuse runs - llm._model(): bare CLI (python -m webwright.tools.skill_use) builds model from SKILL_MODEL_NAME/ENDPOINT (or OPENAI_*) env -> same backend as agent, no hardcoded gateway Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
- README: what the module is, the two plug points (skill_use tool + update CLI), components table, gate semantics, backend config, results summary - llm._model(): bare CLI builds model from SKILL_MODEL_NAME/ENDPOINT (or OPENAI_*) env Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
- README: Skill Library section (what it is, reuse via skill_use tool, grow via update CLI, end-to-end validation summary) - tests/skills: 5 unit tests for library/gate/update/evolve/retrieve+decide Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
Remove _grow / update() / _UPDATERS dispatch — evolve() is the single entry now; drop the test_update test that exercised the removed grow path. Keep retrieve/llm fallbacks (useful). Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
…val) Three bugs hit when update.refine emits a large skill on a slow gateway: - llm() ignored max_tokens -> model default ~4000 truncated the refined skill mid-code - llm() had no timeout override -> model default 120s ReadTimeout'd on the ~16k-token refine (now request_timeout_seconds defaults 600, env SKILL_MODEL_TIMEOUT) - _extract_code returned raw text (with ```python fence) when the closing fence was missing (truncated) -> skill failed to compile; now strips the opening fence anyway Co-Authored-By: Demi Wang <86202027+DEM1TASSE@users.noreply.github.com>
…lve-time reuse, direct skill run) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@microsoft-github-policy-service agree company="Microsoft" |
…te+manifest -> update -> reuse); fix output_schema examples to gate's {type} form
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…bArena numbers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- traces_from_manifest: 'admit' is now REQUIRED per run — a missing gate verdict raises instead of silently defaulting to admitted (was the main pollution risk) - _slug: templates longer than 48 chars get a short content-hash suffix so two templates sharing a long prefix can no longer overwrite each other's skill - skill_use.recommend: the decision's skill_id must be one of the RETRIEVED candidates; anything else (LLM hallucination, even an existing library id) downgrades to skip - tests for all three Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… is truthy) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tive paths, missing answer file) Independent cleanroom reproduction (fresh clone + venv, README-only, public GitHub tasks) surfaced usability failures; mechanism itself reproduced end-to-end in 25 min. - with_skill_hint resolves the library path to ABSOLUTE (F2): the hint's command runs in the agent's workspace, where a relative ./library silently resolved to a nonexistent dir -> empty library -> every lookup skipped, no error, answer still right - skill_use.recommend: a missing/empty library now answers skip with an explicit 'warning: library empty at <abspath>' BEFORE Library() can mkdir the bogus path (F3) - README: step 1 now tells the agent to write agent_response.json (stock webwright does not produce it; the gate/manifest flow assumed it) with a copyable ANSWER_SPEC (F1); absolute-path + --library-beats-env notes (F2/F4); custom endpoint tip (F5) - skills/__init__ no longer eagerly imports update -> no more runpy RuntimeWarning on 'python -m webwright.skills.update' (F6); import evolve/Trace from the submodule - tests: hint abspath, empty/missing-library warning (incl. no-mkdir side effect) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s, filled-in inputs - example_library/: the commit-counting skill verbatim as evolve wrote it (runnable standalone via taskspec, no LLM in the loop; functionally verified against a local repo) - README: what a skill looks like (catalog card + the distilled git-log algorithm), measured held-out numbers (33->10 steps; wrong->correct rescues; honest note that reuse costs more than it saves on cheap tasks), three try-it paths - honest coverage-boundary demo: an unseen period shape raises cleanly; on the real held-out run the agent read the source and adapted around it - tasks/batch/taskspec example JSONs matching the how-to-use steps - links from the module README Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…, safe growth, measured results) + data-flow/interfaces diagram
…no manifest) python -m webwright.skills learn <runs_dir> --library ./library - auto: reads task.json/agent_response.json per run, gates (gold via --golds, else self_verify), groups tasks into templates + extracts params with one LLM call per chunk (default 25; existing templates passed in so chunks refine instead of duplicating), infers output_schema from the answer shape, site from start_url - idempotent: processed runs remembered in library/.learned.json; --dry-run plan mode - README: Quickstart (two commands) + use cases up top; old walkthrough demoted to 'Manual mode'; examples/solve_with_library.sh wrapper (hint + answer instruction) - unit tests for the LLM-free plumbing
…ks, leaks) External-user test of the friendly path surfaced that a trivially-easy config mistake (gateway key + unset OPENAI_ENDPOINT) silently disabled ALL reuse. Fixes: - skill_use: a hard error still degrades to skip (never block solving) but now says LOUDLY it is a LOOKUP FAILURE, not a no-match — error field in the JSON, hint about OPENAI_ENDPOINT/SKILL_MODEL_ENDPOINT, and a stderr line (F1+F2) - README Quickstart: gateway users must export OPENAI_ENDPOINT/OPENAI_MODEL for BOTH steps, stated where step-1 users actually look (F2) - learn: grouping-LLM failure now exits with a one-line actionable message instead of a 40-line traceback (F3); skipped-for-missing-answer runs get a visible summary with the correct pointer (the old message named a command that does not exist) (F4) - learn: strips the answer-output instruction from task text so it cannot leak into templates/skill_ids (F7) - solve_with_library.sh: usage check instead of passing empty args into the CLI (F6)
…ers get them too)
…ssion tests - README: "Only verified solves get in" -> "Validation-gated, exactly as strong as the gate you give it" — states plainly that the default self_verify checks shape only and that the WebArena numbers used the gold gate; learn prints the same warning at run time when no --golds is given - examples/learned_library/: a skill produced by "skills learn" from 3 real GitHub solves — n_solves=3, owner/repo lifted to parameters, two extraction strategies as fallbacks; verified standalone on an unseen repo (numpy/numpy -> v2.5.1, no model); test_learned_example.py locks n_solves>=3 + lifted params + no leak - regression tests for the interface-test findings: F1 (skill_use surfaces hard errors as ERROR, not quiet skip), F3 (learn exits with an actionable message)
There was a problem hiding this comment.
Pull request overview
This PR introduces a new webwright.skills subsystem that turns previously solved tasks into reusable, executable “skills”, enabling solve-time reuse (via a CLI tool) and offline library growth (via learn/update pipelines) while keeping the main agent loop unchanged.
Changes:
- Adds a disk-backed skill library (
Skill/Library) plus retrieve/decide/gate/evolve/learn modules to store, select, admit, and incrementally refine skills. - Adds
webwright.tools.skill_useas a solve-time CLI that recommendsuse|adapt|skipand provides the source path for reuse. - Adds docs/config/examples and new tests to validate deterministic plumbing and example artifacts.
Reviewed changes
Copilot reviewed 29 out of 31 changed files in this pull request and generated 13 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/skills/test_retrieve_decide.py | Adds deterministic tests for retrieve/decide + skill_use/prompt behavior (currently not pytest-discoverable). |
| tests/skills/test_library.py | Adds tests for disk persistence of Library (currently not pytest-discoverable). |
| tests/skills/test_learned_example.py | Adds a check that the checked-in learned example is aggregated/parameterized (currently not pytest-discoverable). |
| tests/skills/test_learn.py | Adds tests for learn plumbing + regression handling (currently not pytest-discoverable). |
| tests/skills/test_gate.py | Adds tests for gate admission logic (currently not pytest-discoverable). |
| tests/skills/test_evolve.py | Adds tests for evolve behavior and slug collision avoidance (currently not pytest-discoverable). |
| src/webwright/tools/skill_use.py | Introduces solve-time library recommendation tool with guardrails against missing/empty libraries and hallucinated skill IDs. |
| src/webwright/skills/update.py | Implements incremental library evolution and refinement prompt construction + manifest ingestion. |
| src/webwright/skills/retrieve.py | Implements LLM-based retrieval plus a simple deterministic keyword-overlap fallback. |
| src/webwright/skills/decide.py | Implements LLM-based use/adapt/skip decision over retrieved candidates. |
| src/webwright/skills/gate.py | Implements admission gate (gold/self_verify/none/auto) to prevent wrong solves from entering the library. |
| src/webwright/skills/learn.py | Adds “friendly” pipeline to learn skills from run folders with gating, chunked grouping, and an idempotent ledger. |
| src/webwright/skills/library.py | Adds on-disk skill storage (<id>/skill.py + meta.json) and simple list/get/add APIs. |
| src/webwright/skills/llm.py | Adds backend-agnostic LLM helper using Webwright’s Model abstraction. |
| src/webwright/skills/prompt.py | Adds with_skill_hint() helper that prepends a bash command hint to consult the skill library. |
| src/webwright/skills/init.py | Exposes the public webwright.skills API surface for consumers. |
| src/webwright/skills/main.py | Adds `python -m webwright.skills <learn |
| src/webwright/skills/README.md | Adds comprehensive module documentation, usage patterns, and rationale. |
| src/webwright/skills/pipeline_diagram.svg | Adds diagram documenting data flow and interfaces for the skills pipeline. |
| src/webwright/config/skill_mode.yaml | Adds optional config overlay to increase step budget for skill reuse runs. |
| src/webwright/skills/examples/README.md | Adds examples overview and how-to for running skills/tools and batch pipeline. |
| src/webwright/skills/examples/solve_with_library.sh | Adds helper script to prepend hint + answer spec and run Webwright with a library. |
| src/webwright/skills/examples/taskspec.example.json | Adds example taskspec input for running a skill standalone. |
| src/webwright/skills/examples/tasks.example.json | Adds example batch task list input with params/golds. |
| src/webwright/skills/examples/batch.example.json | Adds example manifest for update (admit/params/schema/etc). |
| src/webwright/skills/examples/example_library/how_many_commits_did_user_make_period_in_the_cur/skill.py | Adds a runnable example skill produced by the pipeline. |
| src/webwright/skills/examples/example_library/how_many_commits_did_user_make_period_in_the_cur/meta.json | Adds metadata for the example skill. |
| src/webwright/skills/examples/learned_library/what_is_the_latest_release_version_of_ow_c29dab8/skill.py | Adds a checked-in “learned” skill example aggregated from multiple solves. |
| src/webwright/skills/examples/learned_library/what_is_the_latest_release_version_of_ow_c29dab8/meta.json | Adds metadata for the learned skill example. |
| README.md | Adds top-level README section linking to the new skill-library capability and docs. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| f"answer={json.dumps(tr.answer, ensure_ascii=False)[:120]})\n```python\n{tr.code}\n```" | ||
| ) | ||
| sys_prompt = _REFINE_SYS + (_REFINE_INCREMENTAL if existing else "") | ||
| code = _extract_code(llm(sys_prompt, "\n\n".join(blocks), max_tokens=16000, timeout=400)) |
There was a problem hiding this comment.
Fixed in 6513d87 — removed the kwarg (it was indeed swallowed by llm()'s **_). Timeout is governed by the model's request_timeout_seconds (600s default in _model()).
| if answer is None and (d / "agent_response.json").exists(): | ||
| try: | ||
| answer = json.load(open(d / "agent_response.json")).get("retrieved_data") | ||
| except Exception: | ||
| pass |
There was a problem hiding this comment.
Fixed in 6513d87 — switched to Path.read_text(encoding="utf-8").
| manifest = json.load(open(a.manifest, encoding="utf-8")) | ||
| traces = traces_from_manifest(manifest) |
There was a problem hiding this comment.
Fixed in 6513d87 — switched to Path.read_text(encoding="utf-8").
| def llm_json(system: str, user: str, **kw: Any) -> dict: | ||
| """Call + parse the first {...} JSON object out of the reply.""" | ||
| txt = llm(system, user, **kw) | ||
| match = re.search(r"\{.*\}", txt, re.S) | ||
| if not match: | ||
| return {} |
There was a problem hiding this comment.
Fixed in 6513d87 — llm_json now walks the reply with json.JSONDecoder().raw_decode and returns the first valid JSON object, skipping non-JSON brace snippets; matches the docstring.
| meta = json.loads((d / "meta.json").read_text()) | ||
| code = (d / "skill.py").read_text() if (d / "skill.py").exists() else "" | ||
| return Skill(skill_id=skill_id, code=code, meta=meta) | ||
|
|
||
| def add(self, skill: Skill) -> None: | ||
| d = self._dir(skill.skill_id) | ||
| d.mkdir(parents=True, exist_ok=True) | ||
| (d / "skill.py").write_text(skill.code) | ||
| (d / "meta.json").write_text(json.dumps(skill.meta, ensure_ascii=False, indent=2)) |
There was a problem hiding this comment.
Fixed in 6513d87 — explicit UTF-8 on all skill/meta reads and writes in library.py (and the same for learn.py's ledger/run files).
| if __name__ == "__main__": | ||
| run() |
There was a problem hiding this comment.
Fixed in 6513d87 — added a pytest entry point (test_all).
| if __name__ == "__main__": | ||
| run() |
There was a problem hiding this comment.
Fixed in 6513d87 — added a pytest entry point (test_all).
| if __name__ == "__main__": | ||
| run() |
There was a problem hiding this comment.
Fixed in 6513d87 — added a pytest entry point (test_all).
| if __name__ == "__main__": | ||
| run() |
There was a problem hiding this comment.
Fixed in 6513d87 — added a pytest entry point (test_all).
| if __name__ == "__main__": | ||
| run() | ||
| run_regressions() |
There was a problem hiding this comment.
Fixed in 6513d87 — added a pytest entry point (test_all) that runs both run() and run_regressions().
The README table (held-out 70% vs 55%, 14.7 vs 17.1 steps; train 26/30 vs
23/30) previously existed only as prose. Now:
- evals/webarena/results/: sanitized per-task records of the exact run behind
the table — task id, answer, gold score, steps, skill verdict; one command
("reproduce.py table --results results") re-derives the table, no setup
- evals/webarena/reproduce.py: self-contained driver that re-runs the whole
experiment (train -> gold-gated update -> held-out with/base -> table)
against your own WebArena deployment via microsoft/webarena-verified;
resumable, parallelizable per template
- run_all.sh + model.eval.yaml (agent-model overrides; gateway pointer)
- tests/skills/test_eval_snapshot.py: CI-locks the records to the published
numbers and enforces snapshot sanitization (no local paths/hosts/keys)
- skills README + examples README now link the records instead of asking for
trust; CI also triggers on evals/webarena/**
- prompt.py: shell-quote task and library in the skill_use hint (shlex.quote) —
$VAR / $(...) / backticks expanded in bash even inside the old double quotes;
regression test added
- llm.py: llm_json now scans for the FIRST valid JSON object (raw_decode) as
documented, instead of a greedy first-{ to last-} regex that could span
unrelated braces
- update.py: drop the misleading llm(..., timeout=400) kwarg (silently
swallowed; the model's request_timeout_seconds already governs); read JSON
files via read_text instead of unclosed open()
- library.py / learn.py: explicit UTF-8 on every skill/meta/ledger read+write
(locale-independent on Windows)
- tests: pytest-discoverable test_all() entry points in all 7 files (CI keeps
running them as scripts too)
There was a problem hiding this comment.
It's the guided tour of the examples/ directory (the module README links here for every "see examples"): two real, checked-in skill libraries — learned_library/ is the Quickstart loop's actual output (3 GitHub solves -> learn -> owner/repo lifted to parameters, runs standalone on unseen repos with no model), example_library/ is verbatim update.evolve output from the WebArena eval — plus the solve wrapper and filled-in copies of every input file the manual pipeline asks you to write. 8ea59be makes this explicit in the file's opening paragraph and adds the learned_library provenance section.
There was a problem hiding this comment.
this file looks redundant, can we remove?
There was a problem hiding this comment.
Agreed — removed in 8ea59be. Nothing referenced it: the skill hint is prompt-level (with_skill_hint), and no documented path needed the step_limit bump.
| # 2. turn everything you've solved into skills — no manifest, no fields to learn | ||
| python -m webwright.skills learn outputs/ --library ./library | ||
| ``` | ||
|
|
There was a problem hiding this comment.
@DEM1TASSE It is better to add the complete example in the quick start session.
It needs to additionally include how to use the skill library.
There was a problem hiding this comment.
Done in 8ea59be — the Quickstart is now the complete loop on a copy-pasteable example (public GitHub): solve 3 instances -> learn -> an unseen instance reuses the skill (with the expected skill_decision.json shown), plus how to use the library without the agent (querying skill_use directly, and running the learned skill standalone with no model — verified pandas-dev/pandas -> v3.0.4). The same loop's output is checked in at examples/learned_library/.
…larify examples/ - README Quickstart is now the complete loop on a runnable example (public GitHub): solve 3 instances -> learn -> an UNSEEN instance reuses the skill, plus using the library without the agent (skill_use query, running the learned skill standalone, verified: pandas-dev/pandas -> v3.0.4 with a params-only taskspec) - remove config/skill_mode.yaml: nothing referenced it (the skill hint is prompt-level via with_skill_hint; the step_limit bump was never needed by any documented path) - examples/README: state the directory's purpose up front, and document learned_library's provenance (the Quickstart loop's checked-in result) with the unseen-repo runs and the CI test that locks it
The gate section described two gates and left a hole between them that cost a real build every draw it had: a solve can hold a right answer and contain no method. self_verify admits it — the answer is right, and the answer is all it looks at — and then distillation, forbidden from copying instance values, has to invent an extractor the trace never had and fails replay on that instance forever. So the section now shows the script that caused it, says what happens to it, and names dropped_lookup next to dropped_wrong: one counts answers the gate rejected, the other counts answers it accepted from scripts that never earned them. It also records why the test is "every field verbatim" and not the obvious "the answer appears in the code" — the obvious one flags all three shipped trajectories, because an airline name is a vocabulary entry and a time lands in an assertion. That's the kind of thing a reader will re-derive and get wrong unless the measurement is written down: 0 of the 5 solves that distilled, 1 of the 1 that couldn't.
The measurement belongs in the commit that made the decision, not in a reference page. Gone: the 0-of-5 / 1-of-1 counts, "the narrowness is the point", and the sentence about times landing in assertions. What's left is the script, what happens to it, and the one clause a maintainer needs before loosening the test — "the answer appears in the code" flags working solves too. 24 lines to 16.
Updated README to enhance clarity and detail about the Web Skill Factory, including integration points and setup instructions.
The README now says "build and quickstart.sh translate the environment variables into the appropriate agent configuration". That was true of build and not of this script, which read MODEL_CFG or nothing — so exporting a gateway sent `ask` there and `solve` to api.openai.com, and the fix was a yaml you had to write yourself. Same trick build uses: the CLI takes inline `model.key=value` specs, so the script passes OPENAI_ENDPOINT/OPENAI_MODEL through as overrides on top of the defaults. No yaml, no MODEL_CFG, nothing to remember. MODEL_CFG still wins when set, which is now what it's for: putting the agent on a different model than the distiller. Checked all three paths: env set -> base.yaml + model_openai.yaml + the two overrides; MODEL_CFG set -> base.yaml + your yaml; neither -> the old defaults. warn_gateway_agent goes with it. It existed to tell you the script was about to ignore your gateway; it can't. That's twice today a warning turned out to be a fix I hadn't written yet. The header comment was also wrong from the day it was written — "MODEL_CFG ... for the agent in solve/build" — build has never read MODEL_CFG, only -c.
Updated instructions for setting up the module and clarified the importance of keeping the virtual environment activated.
…the way out Verifying from a fresh clone on a Mac, step 1 dies at the airport field: choose_airport presses Control+A to clear the box, which is select-all only on Linux/Windows — on macOS it's move-to-line-start, so "SEA" gets appended to the prefilled origin and no clean option matches. The failure is in the code and the platform, not the user's setup: a strict-replayed skill reproduces its *training* run (Linux, Google Flights as it was), which is not generalization. That's the research point, so the note frames it that way rather than as a caveat. And it gives the way out instead of a dead end: when standalone won't run, `ask` and `solve` carry the skill as a prior the agent reads and adapts around the difference, and keeping that solve lets `learn` fold it back in — so the next standalone run covers the new platform too. reference-as-prior and refine, which the module already has, now have a concrete reason to exist in the docs. Placed under step 1, where the reader just ran the thing that breaks.
… 16000 to 4000 Reported from a Mac: `solve` reusing the checked-in skill looped forever, re-emitting a 700-line final_script.py and never running it. It's the agent's output being truncated mid-script — the write never completes, the JSON is invalid, the harness retries, repeat. Confirmed against a Linux run's config snapshot: it used max_output_tokens: 16000 and worked. The regression is mine, from making a gateway reachable by env vars alone. That env path stands in for a hand-written model yaml — and examples/ model_gateway.example.yaml sets max_output_tokens: 16000 precisely because base.yaml's 4000 is too small for the agent to emit a reused skill. I forwarded openai_endpoint and model_name from the env and stopped there, so the synthesized config inherited base.yaml's 4000. The convenience quietly quartered the budget, and a big skill then can't be written in one response. So both env-synthesis paths — build's _agent_cfg and quickstart.sh's solve config — now also carry model.max_output_tokens, defaulting to 16000 to match the yaml they replace, overridable via SKILL_AGENT_MAX_TOKENS. An explicit -c or MODEL_CFG is untouched: your yaml still owns the budget. Inline ints parse as ints (the loader yaml.safe_loads the value), verified end to end through the config merge. Tests pin it and the mutation was checked: drop the append and the "env path carries a usable output budget" test fails. Does not fix the Mac run itself — the reused skill still clears the airport box with Control+A, which isn't select-all on macOS (documented in the README's step 1). This only restores the budget my change took away.
…al output_schema - collect_runs: recover the answer from the trajectory exit message when a run has no agent_response.json (plain solves), map exit_status->status, skip failed/showcase runs loudly - canonicalize_answers: define ONE canonical output_schema per template at aggregation and reshape each member's answer to it (mechanical when already structured; else one reshape-only LLM call). schema now belongs to the template, not each solve - _GROUP_SYS: split by site / core action / optimization objective / output schema; merge paraphrases + filters (params). objective is intent, not a parameter
…ews entry - point at src/webwright/skill_factory/ (the module was renamed from skills/) - lead with what a skill is here (a program, not a document) and the two gates - name the current entry points: learn / build / update, and skill_use at solve time - replace the old GitHub walkthrough with the shipped results: standalone ~40 s / zero tokens, and WebArena held-out 55% -> 70% (+15 pp)
…n the reference The old figure packed data flow, function signatures and JSON shapes into three columns; it answered 'what are the interfaces' rather than 'what does this do'. The new one keeps three things only: the components, the loop, and what a skill is on disk, with the batch and the widening made explicit. The old figure is still the best picture of how the modules talk to each other, so it moves to docs/skill_factory/reference.md next to the file-by-file table, as skill_factory_interfaces.
The section described the solve-time half in full and left the growth half to one arrow in the diagram. Adds the other side: how solves are grouped into a template, how the group is aligned into one parameterized program with primitives, what the two gates admit, and why an existing skill is widened in place instead of rebuilt.
The second caption line ran past the card's right edge. Drops 'in place', which the widened row below already says.
Updated README.md for clarity and consistency in language.
…ives are for - de- duplication -> de-duplication (a line-wrap artifact) - two em dashes had no spaces where every other one does; rewritten without them - How it works described the parameter half of distillation only. Adds the other half: the site-driving core is factored into named functions, which is what lets a later task on the same site reuse it and change only its final step (the 'adapt' verdict).
Revised the explanation of library growth and task alignment in the README.
…the next paragraph
… the source
Verified the OpenCLI column against jackwener/OpenCLI. Five of seven cells hold
as written; two overstated and are now accurate against the repo:
- "agent can adapt it" said OpenCLI edits source "only to repair the shared
adapter when it breaks." Too narrow: `opencli adapter eject` + `opencli browser
init` are general local-override / private-authoring paths, not repair-only
(README "Modify an official adapter locally"). The real distinction isn't
repair-vs-not — it's that OpenCLI's edit is a persisted, maintained adapter
file, where ours is a per-task reshape of source-in-hand that leaves the library
untouched. Reworded to that.
- "grows from your runs: nothing accumulates" was too absolute. opencli-sitemap-
author captures a discovered workflow, and site knowledge persists to
~/.opencli/sites/<site>/notes.md. Knowledge does accumulate — as adapters and
sitemaps authored from a session. What OpenCLI lacks is a mechanism that widens
an existing adapter from repeated runs, or replays accumulated past answers
(grep for widen/accumulate/self-evolve/regression in README+skills: none). Now
says "only at authoring time — ... repeated runs never widen an existing adapter
and past answers aren't replayed."
Left correct as-is: a skill is one CLI command per capability (clis/<site>/<cmd>);
author-declared args (hackernews/top.js: args:[{name:'limit',type:'int'...}]);
verify at authoring + tests (browser recon verify + verify/<cmd>.json + *.test.js).
Confirmed the first-row framing too: OpenCLI's adapter is a parameterized pipeline
that runs with no model (top.js is fetch->map->filter, browser:false, zero LLM
calls), so grouping it with programs-that-run rather than docs is right. Its own
"skill" (SKILL.md) is the agent-facing doc; the executable unit is the adapter —
noted under the table so the terminology collision doesn't mislead.
Added a source-attribution line under the table pointing at the exact files.
…ne they lack Stress-tested each Ours cell against jackwener/OpenCLI as an adversarial author. Most hold; two rested on claims OpenCLI can rebut, so they're retired in favor of the differentiators that actually survive. - domain: dropped "private, cross-site, multi-step workflows no shared catalogue has." OpenCLI does all three — private local adapters via `opencli browser init` (~/.opencli/clis/), multi-step pipelines (top.js chains fetch→map→filter→fetch), even multi-domain in one adapter (semanticscholar hits 4 hosts). Claiming those as ours is a false differentiator. The un-rebuttable axis is specificity: ours is the one task *you* repeat, not a capability a shared catalogue would carry. OpenCLI's cell is now "a shared catalogue of a site's capabilities," which is what clis/ is. - verified: "reproduce its own answers, no model" was rebuttable — OpenCLI's verify fixtures are also a model-free output check. The real gap is what the check compares against: OpenCLI's fixtures are author-written expected values (verify/<cmd>.json: patterns/notEmpty/mustNotContain rules); ours replays the answers our own solves produced. Sharpened to "reproduce the answers from its own solves," and OpenCLI's cell now says "author-written fixtures checked at authoring time, plus unit tests." The input gate wording is now "a wrong solve never becomes material," which is genuinely unique because only ours builds a skill from solves at all. Left standing because they survive the test: distilled-from-solves (row 3) and params-from-observed-diffs (row 4) — none of the three do either; per-task reshape with the library untouched (row 6); widens-in-place-from-runs (row 7). Those four are the core, and each is something the other columns genuinely don't have.
…ces, drop the correctness overclaim Investigated HKUDS/OpenSpace the same way as OpenCLI — cloned, read the source. It's the closest neighbor, so it goes in the table; and it forced two honest corrections. Correctness overclaim, removed. The verified cell said "a wrong solve never becomes material." That's only true with --golds. The default self_verify gate checks shape + non-empty + the agent's own success report, so a wrong-but- plausible answer passes. And the output gate proves the skill *reproduces* its own recorded answers with no model — a consistency/determinism property, not a proof the answer is right. The cell now says exactly that. OpenSpace's gate is also not correctness: a judge (deterministic|llm|gdpval) rules the candidate behaves no worse than baseline (outcome can be needs-human-review). Neither system establishes ground truth without external labels; the table now says so. OpenSpace shares "grows from runs" and "agent can adapt" with us, so those rows can't be yes/no differentiators. Kept them anyway, reframed to the difference in kind, each cited to source: - grows: OpenSpace *proliferates* skills — CAPTURED per trace, DERIVED as a specialization of a parent (engine.py: DERIVED requires a parent). It never lifts parameters from several instances into one widened program. Ours widens *one* program in place, params from the diffs across your solves — a correct- but-narrow solve generalizes. - adapt: OpenSpace adapts by evolving the shared library (new revisions through gates); ours reshapes source per task and leaves the library untouched. - produced by: folded in online+offline — ours distills a folder of trajectories you already have (offline) or as you build (online); OpenSpace captures from a single trace. OpenSpace verification cited to behavior_eval.py (replay is the only committing gate; sandbox runner is host-supplied, require_replay_runner=True) and the deterministic|llm|gdpval judge policy. Skills are SKILL.md docs (skills/*/SKILL.md, pure prose), so it sits with the document columns on "a skill is." Attribution line now covers both repos with file-level pointers.
…w, parallel cells Both rows buried their distinction in prose and the cells answered different questions. Reframed each row as a single question every cell answers the same way. "agent can adapt it" -> "adapting it for a task means". Now the axis is what an edit entails, and the cells line up: nothing (fixed doc) / a persistent adapter you maintain (OpenCLI) / a gated library revision (OpenSpace) / a throwaway copy reworked for the one task, library untouched (ours). The ours differentiator — per-task and free, vs everyone else's persistent artifact — is now the first thing you read. "grows from your runs" -> "what grows from your runs". The axis is what grows: nothing / a document / nothing automatic / the library's count, a new skill per trace (OpenSpace) / one skill's breadth, generalized from the diffs across your solves (ours). This makes the OpenSpace contrast the point — they proliferate skills, each as narrow as its trace; we widen one skill. Same evidence as before (engine.py CAPTURED/DERIVED), just said plainly.
…ain (2026-07-23)
A source re-audit found both columns overstated. Re-cloned and checked each claim
against current main (OpenSpace HEAD unchanged from my earlier read — I had read
it too narrowly, not stale). Corrections, all cited:
OpenCLI was written as only its adapter layer. It has three surfaces (README):
site adapters, generic `browser` primitives, and a CLI/plugin hub (gh, docker,
custom). Fixed "a skill is." "what grows: nothing automatic" was too absolute —
site knowledge persists to ~/.opencli/sites/<site>/ (endpoints, field maps,
fixtures, notes); adapters and that memory accumulate through authoring, they
just aren't auto-generalized from repeated solves. "autofix the shared one" was
wrong: autofix patches a local override and reports upstream only with approval
(opencli-autofix). "verified" now notes the live-page check the authoring
workflow requires and "repo tests where present" rather than implying universal
unit tests.
OpenSpace's "verified" was my biggest error: I described behavior_eval.py's
internal replay/judge machinery ("no worse than baseline, needs-human-review")
as the mechanism. The project's actual model is a trust lifecycle — new skills
are provisional until independent cross-task success promotes them to trusted,
attributable failures demote them, a version is validated before it replaces the
old (types.py PROVISIONAL/TRUSTED; README "Provisional first" / "Independent
trust"). Also: a skill is a directory centered on SKILL.md (+ helper files), not
just a doc; evolution is FIX (revision) / DERIVED (coexisting specialization) /
CAPTURED (trace-grounded, execution + validation), not "one skill per trace";
"what grows" accumulates evidence/trust/lineage, and FIX grows history without
adding count. Benchmarks (Terminal-Bench/GDPval) are eval settings, not the
domain — domain reworded to general reusable agent workflows.
SkillOpt tightened too: it optimizes a natural-language doc by bounded
add/delete/replace edits (not monotonic growth), and a candidate replaces the
current doc only if it strictly improves on a validation split (the held-out
test is the final eval, not the gate).
Ours cells unchanged — the differentiators still hold once the neighbors are
described accurately: a parameterized program distilled from your solves, params
from observed diffs, model-free exact reproduction as the output gate. The <sub>
now says "verified against current main (2026-07-23)" with file/section pointers.
The table had grown to seven dense rows. Dropped the two most secondary: "domain / whose need" (positioning, not mechanism; overlapped "a skill is" and "produced by") and "adapting it for a task means" (a real but non-central difference). The remaining five are the whole argument, tight: what a skill is (a program, no model), produced by (distilling several of your solves, offline or online), parameters (the diffs observed across them), verified (model-free exact reproduction), grows (one skill generalized in place). Every remaining row is a core differentiator; the corrected OpenCLI/OpenSpace cells are unchanged.
The cells had turned into jargon — "deterministic CLI command," "validation split," "provisional/trusted," "author-tightened fixtures," "regression-replayed." Accurate, unreadable. Rewrote every cell in everyday words and made the row labels plain questions: what a skill is, where it comes from, how it handles different inputs, is it checked to work, does it get better as you use it. Ours in one line per row: a small program that runs by itself; built from a few times you did the task; figures out the varying inputs on its own; re-runs to reproduce its own answers exactly with no AI (consistent, not proof-of-correct); one skill gets broader with use. The precise, source-cited version stays in the footnote for anyone who wants the exact mechanisms.
…op insider terms
Last pass over-corrected into baby talk ("a note the AI reads", "it doesn't,
really", "eyeballs the live page"). Raised it back to plain professional English —
document, model, agent, verified, parameters, consistent — while keeping the
jargon out of the cells: no FIX/DERIVED/provisional/fixtures/validation-split/
args-schema in the table. Those exact, source-cited terms remain in the footnote
for readers who want the precise mechanisms. Row labels are clean questions: what
a skill is, how it's created, how it handles variation, how it's verified, does it
improve with use.
…ed wording Replaced the comparison with the reviewed version verbatim — it reads cleanly and every cell checks out against source. One fix only: the "how it's created / Ours" cell had a "[by llm/pipeline?]" placeholder; resolved it to what the code does — an LLM distills the runs into one program (update.py _refine). Moved it out of its own top-level section and into a collapsed <details> under How it works, right before "The Quick Start below demonstrates the complete workflow," so the section stays scannable and the comparison is there for anyone who opens it. Dropped the source-citation footnote in favor of a plain disclaimer: this is a friendly comparison based on our understanding; if we've mischaracterized a project, open an issue or PR and we'll fix it.
The "what a skill is" row said "the model reads" for SKILL.md and SkillOpt but "an agent reads" for OpenSpace — same meaning, inconsistent wording, no reason for it. Unified to "the model reads" (kept OpenSpace's helper-files nuance), which is also the axis this row is about: all three are documents that need the model to run, versus ours, a program that runs with no model.
"an LLM distills them into one program" was too narrow. Creating a skill is a pipeline: an input gate filters wrong solves, an LLM groups the runs and writes the program, then a deterministic model-free replay-verify keeps it only if it reproduces the recorded answers (learn.py -> update._refine -> _replay, with a draws x rounds retry). The LLM writes the code, but the pipeline — and especially the verify step — is the point. Reworded to say so.
…iewed wording Consolidated to four rows and applied the reviewed fixes: - "how it's created" and "what grows from your runs" merged into one honest row, "how the library evolves from runs" — ours also grows the library (add a skill for a new template, leave it unchanged on a plain reuse, refine it on an adapted run), not just "widens one skill." Same for the others, stated plainly: OpenSpace fixes/derives/captures documents; OpenCLI authors maintain commands. - "how it handles variation" -> "how one skill covers different inputs," asking the concrete question: does the skill take real inputs, and who set them. Ours: aligned verified runs, differences become explicit parameters. Others: prose the model improvises from, or hand-declared arguments. - "no model needed" -> "no model required (an agent can still use it)" — the skill runs standalone, but can also be used by an agent; the distinction from the document rows is that those always require a model. - Dropped the jargon "fixture" for "a saved expected result." - "the author" -> "a person or agent" for OpenCLI authoring/maintenance — confirmed in source: opencli-adapter-author is explicitly an agent workflow and autofix auto-repairs adapters. Ours cells unchanged in substance; the differentiators still stand once the neighbors are described accurately.
What
Adds Web Skill Factory (
webwright.skill_factory) — a self-evolving skill factory (MVP): turn solved tasks into reusable,executable code skills, retrieve + judge them at solve time, gate what enters the library, and grow
the library incrementally. A self-evolving loop on top of Webwright's code-as-action solves:
This is the reuse + accumulation layer on top of Webwright's code-as-action solves: it consumes
the
final_script.pyevery solve already produces (plain or crafted mode — both work), accumulatesskills across tasks, judges when a prior skill applies, and improves skills as more solves arrive —
with a gate so wrong solves don't pollute the library. It complements
crafted_cli: wherecrafted_cliparameterizes a single task's script by anticipating what might vary,update.refineparameterizes across multiple verified solves — the differences actually observed between
instances become the parameters.
Modular composition (~810 lines of core code)
Ten small, single-responsibility modules — each with a stable interface and a swappable
implementation:
skill_factory/library.pySkill+Library, skills on disk (skill.py+meta.json)skill_factory/retrieve.pyskill_factory/decide.pyskill_factory/gate.pyskill_factory/update.pyrefineparameterizes + decomposes into primitivesskill_factory/llm.pyModel(no endpoint/key hardcoded)skill_factory/prompt.pywith_skill_hint)skill_factory/learn.pylearn <runs_dir>: auto-group runs into templates, gate, evolve; no manifest to writeskill_factory/__main__.pypython -m webwright.skill_factory <learn|update>dispatchertools/skill_use.pyHow it plugs in (no change to the agent loop or default config)
skill_usetool, invoked from bash likeself_reflection/image_qa:{verdict: use|adapt|skip, skill_id, source_path, how_to_reuse}.updateCLI distills a batch of gate-passed solves into aparameterized, primitive-decomposed skill:
learngroups a folder of finished runs intotemplates (one LLM call per chunk), gates them, and evolves the library — idempotent,
--dry-run:examples/learned_library/checks in the skill this produced from 3 real Google Flightssolves — five parameters lifted (origin/destination city+code, date), verified on an unseen
route three independent ways (from scratch / reuse / standalone, same answer) — with a CI
test locking it.
Validation
WebArena: 10 templates × 3 domains — reuse lifts accuracy +15pp and saves steps on held-out tasks
10 retrieve-type task templates across shopping_admin / gitlab / map. Per template: 3 train
tasks build the library (solved from scratch; only gold-verified solves are admitted), 2 held-out
tasks (unseen instances of the template — different parameter values) measure reuse. Every task is
solved both WITH the library and from scratch (BASE) — 100 solves total.
Per-task records and a reproduction driver are kept in the companion research repo and can be
shipped here on request.
Highlights:
library; net reuse-wins 7 vs 1 regression across the 20 held-out tasks.
33 steps (scratch) to 10 (reuse); a map routing task from 29 to 16.
here. (The gate is exactly as strong as its verifier — the default
self_verifyis a shapecheck only; see the README's "validation-gated" section.)
update.refinelifts per-instance differences into parametersand bakes the aggregation logic (top-n ranking, commit counting, route-time extraction) into
primitives, so unseen instances of the template solve by a direct
useof the skill.skill from the shared library (grown to 10 skills over the run), including telling apart two
near-duplicate gitlab commit-counting skills (by-date vs by-period).
evolvebatches produce 4 independent skills — new templates get added, existing skills arerefined in place (working functions kept), skills with no new traces stay byte-identical, and
zero cross-contamination between skills; held-out reuse against the mixed-built library matches
the per-template-built one.
Real website (public GitHub, read-only): the full loop end-to-end
Solve two repos from scratch ->
updatebuilds a parameterized skill -> a held-out repo is solved byreusing it (the agent calls
skill_use, verdictuse, answer correct). Reuse pays off most onmulti-step tasks where saved exploration outweighs the lookup overhead (see the WebArena numbers);
on short single-page lookups it is roughly break-even.
7 unit-test files under
tests/skill_factory/(library / gate / evolve / retrieve+decide / learn /learned-example lock / eval-snapshot lock) run in CI on every push touching the module
(
.github/workflows/skills-tests.yml).Status: a deliberately simplistic MVP
Most steps are a single LLM call (retrieve = one catalog prompt, decide = one prompt, refine =
one batched prompt) — chosen for clarity, not yet for scale/accuracy. The point is the modular
shape: each stage has a stable interface, so swapping in something stronger (embedding retrieval,
a learned ranker, WebJudge / cross-source consistency for the real-website gate) is a localized
change that does not touch the others or the agent loop.
Scope
Purely additive (zero deletions), confined to
src/webwright/skill_factory/(module + examples,including a checked-in learned skill),
src/webwright/tools/skill_use.py,tests/skill_factory/and one CI workflow. The actual implementation is ~670 lines of logic(non-blank, non-comment, across the skills module + the
skill_usetool); the rest is tests,examples, eval records, and docs. No edits to the agent loop, models, or existing configs.
Module README:
src/webwright/skill_factory/README.md.