Skip to content

Repository files navigation

Plan2Agent

npm version npm downloads CI Node.js 22+ License: MIT

English | 한국어

Use Plan2Agent as a development assistant: describe the result you want, confirm the few decisions that change that result, and let it guide planning, implementation, recovery, and verification.

Install in 30 seconds

Plan2Agent requires Node.js 22.20.0 or newer.

npm install -g plan2agent
cd <project-dir>
p2a init --tools all --codex-profile quality
p2a next

p2a next reads the local project state and returns one concrete next action. Run it again whenever you finish a planning, approval, or development step.

Your first plan in 5 minutes

After initialization, either pass a one-sentence idea with p2a next --idea "<what to build>" or use a concise Markdown/text document with p2a next --entry <path>. The CLI stores the request locally, so chat is never the only copy of the requirement.

Agent Example
Codex Run p2a next --idea "Add release status", then use $p2a-harness.
Claude Code p2a next --idea "Add release status", then /p2a-harness
Gemini CLI p2a next --entry docs/idea.md, then /p2a:harness

Plan2Agent first explains what it understood, asks only about choices that would materially change the product, and confirms the scope and implementation plan before coding.

Natural-language request
  -> short understanding summary and only necessary questions
  -> scope confirmation
  -> implementation-plan confirmation
  -> implementation, automatic recovery, and risk-based verification
  -> recommended close, with optional code review or retrospective

Internally, scope and implementation-plan confirmation remain separate safety boundaries. Their compatibility names are Gate A and Gate B. An existing constitution is reused; the optional Gate ② appears only for a hard prohibition or consequential, difficult-to-reverse architecture/stack choice. Ordinary repository conventions remain advisory. Gate names, state ids, hashes, and artifact paths stay out of normal assistant messages and remain available through p2a next --details or JSON.

The harness writes canonical files under .plan2agent/artifacts/<project_id>/. Complete the suggested action and run:

p2a next

When Gate A-C readiness passes, next guides the transition into supervised execution and subsequent iterations.

Why Plan2Agent

AI coding tools are effective at implementation, but chat history is a fragile place to keep requirements, approvals, dependencies, and verification evidence. Plan2Agent adds a durable control layer around those tools.

Need Plan2Agent approach
Clear decisions before code Gate A scope and Gate B spec approval preserve product decisions; Gate ② is added only for material project-shape commitments.
Traceable implementation work Specs map to a Direct run, Planned checkpoints, or dependency-aware Orchestrated tasks.
Reviewable agent execution Foreground-supervised runs preserve mode, rationale, changed files, and verification evidence.
Portable project state Local JSON artifacts remain canonical across Codex, Claude Code, and Gemini CLI.
Controlled improvement Evaluation and proposal flows recommend maintenance without silently applying self-modifying patches.

Plan2Agent coordinates the workflow; it does not replace your coding agent, source control, or project management system.

What gets written

Planning and execution state stays local to the project:

.plan2agent/
  project.config.json
  constitution.json                 # conditional project-shape contract
  artifacts/<project_id>/
    decisions.jsonl
    gate-a-intake/
      intake.json
    gate-b-spec/
      spec.json
      experience-spec.json       # conditional
      visual-design/              # conditional offline HTML prototypes
    gate-c-task-graph/
      task-graph.json
    current-spec.json
    iterations/
    runs/
    eval/
    proposals/

When external Agent Skills are installed, the team-tracked p2a-skills.lock.json stays at the project root while validated copies live under .agents/skills/ and, when selected, .claude/skills/. The ignored .plan2agent/ manifest records only the local installation state.

decisions.jsonl is the source of truth for recorded approvals and revocations; the existing JSON approval audits remain compatible copies. All artifacts are validated against schemas shipped with the package. Generated Markdown is a human-readable view. Run evidence is temporary execution state by default: the current development bundle remains reviewable and portable, while opening the next iteration removes archived iteration runs. When a retry replaces a run, run-index.json keeps only bounded retrospective counters—never commands, output tails, notes, or run IDs—and drops those counters when the next iteration opens. Git and optional BuildLore projection are the durable history; archived P2A Gate documents are not runtime dependencies for current development. Unmined failed or blocked runs remain available to the proposal flow. Use p2a runs gc --dry-run to review indexed and orphan evidence before cleanup; persistent projects require an explicit --force. Each Git-backed run also records its current HEAD, branch, and dirty state. Set runTracking.persistence to persistent only when long-lived local run evidence is required. Use runTracking.generatedPaths to omit untracked generated files from automatic change collection. Successful standard checks keep shorter output tails; failures and verification bindings remain intact (recording policy). Optional runTracking.retrospectiveSignals thresholds let p2a next --json --contract v2 surface bounded current-iteration performance and process candidates before close. Safe process signals are detected by default; projects can set enabled: false to disable them, while performance budgets remain opt-in. At closeout, product review, P2A retrospective, and close are separate choices; when evidence is current and no signal exists, close is recommended. A clean review asks once to close instead of repeating the menu. Proposal writes remain separately approved and skipping retrospective never blocks close. Continuing the retrospective writes one short docs/retrospective/<project>-<iteration>.md report only after approval. The final maintenance task prints the same review/retrospective/finish choice without adding a persistent maintenance close state. That four-section report can be rendered with p2a proposals issue-preview and, only after an explicit --yes confirmation, published to the public P2A GitHub repository with publish-issue. Previewing never calls GitHub, and publication does not approve implementation. New iterations materialize current-development-contract.json from the approved current state. It contains only the objective, scope, durable project rules, current-iteration architecture/interface/ dependency constraints, preservation constraints, acceptance, verification, authority, and current task bindings required for implementation. p2a next and the normal run lifecycle validate that contract, the current task graph, any constitution, and active run without traversing archived iteration documents. Existing iterative projects can run p2a iteration migrate-current-contract --artifacts <artifact-root> once.

New runs store the current-contract execution envelope once by content hash and reference it from each run record, avoiding repeated copies while preserving hash-verified fail-closed validation. Existing inline records remain readable and p2a runs migrate-schema converts them in place.

Core workflow

1. Plan with approval gates

The planning harness turns an idea into structured intake, product and implementation specs, and validated execution readiness. Gate A presents a compact understanding summary and requires explicit confirmation. The same session reuses any constitution and opens Gate ② only for a material durable project-shape decision before continuing to Gate B. It records uncertainty rather than inventing it.

2. Execute the approved objective

After Gate B approval, use p2a next to start the next action authorized by the current development contract. New projects default to adaptive, while existing configs without an execution mode continue to resolve as orchestrated; explicit adaptive, direct, planned, and orchestrated policies remain supported without another mode approval. Planned mode records 2–5 ordered, command-verified resume checkpoints. New runs bind an execution envelope containing objective, current-contract hash, scope, durable project rules, current-iteration architecture/interface/dependency constraints, preservation conditions, non-goals, acceptance, verification, and authority boundaries.

For direct control of a prepared work item:

p2a execute start \
  --artifacts .plan2agent/artifacts/<project_id> \
  --task <task-id>

During implementation, projects may opt into fast changed-file checks with structured relatedVerification commands and p2a runs verify --related. With no flags, finish selects a related check for docs/metadata work and every configured test, lint, and typecheck for code work. Failed attempts stay in the run, while a corrected retry of the same check at the current revision controls completion.

After every task is done, one shared risk profile selects close evidence. Documentation and metadata need a current related check; configured project commands take precedence, with a packaged file-integrity check as the default. A canonical isolated-code implementation can reuse its current product-revision full verification. Multi-task/worktree integration, high-risk paths, or product code changed after verification require p2a execute verify-final. If documentation changes after a valid product pass, P2A keeps that pass and runs only p2a execute verify-final --scope relevant; a later product-file change requires full verification again.

If a material code-review finding appears after a task is done but before the iteration closes, use p2a execute remediate --artifacts <root> --task <task-id> --finding <text>. P2A keeps the reviewed run immutable, starts a linked run in the same iteration, blocks close while remediation is active, and returns the task to done only after verification passes. Work discovered after iteration close belongs in maintenance or a new iteration.

See the Execution Reference for start, remediation, resume, finish, retry, and bounded batch procedures.

3. Iterate without losing the baseline

Iterations preserve the approved spec, derive change tasks, track maintenance work, and archive closed history. A later Gate A reuses relevant confirmed answers from the baseline and asks again only where the new idea changes or conflicts with them. p2a next guides close/open transitions; p2a iteration exposes the lower-level controls.

4. Evaluate and improve

The eval flow grades run evidence, compares results, and groups recurring failures. The proposal flow can turn supported findings into human-reviewed maintenance tasks. It never applies a patch merely because a proposal exists. With the default active_only retention, run eval before opening the next iteration or starting the next maintenance task; use persistent for deliberate long-term local comparisons.

5. Keep optional long-term knowledge with BuildLore

BuildLore is a local-first, Git-backed knowledge tool. The current Plan2Agent contract remains the local execution source of truth. BuildLore owns optional long-term knowledge projected from completed development. When configured and relevant, the execution owner can consult it from the first attempt, explain progress, and offer evidence-based direction advice without adding an approval gate. Missing or unavailable knowledge does not stop development.

After attaching a BuildLore knowledge/ repository and registering the same project ID, enable the adapter and preview projection from .plan2agent/artifacts/<project-id>/:

p2a enhance buildlore
p2a buildlore status
p2a buildlore sync --dry-run
p2a buildlore sync

BuildLore selects supported approved planning and execution evidence, sanitizes it, and writes reviewable knowledge sources. A connected source repository or knowledge workspace can read approved, project-scoped memory without synchronizing or generating a Wiki:

p2a buildlore search --query "authentication decision" --mode lexical
p2a buildlore memory --task "Prepare the next implementation plan" --progressive --json
p2a buildlore lookup --kind evidence --id <canonical-evidence-id> --expect-generation <generation-digest> --json

Use canonical evidence IDs from the memory registry, not its short aliases, and keep lookup bound to the returned generation. Reads default to a 15-second timeout and 256 KiB output limit; --timeout-ms can set a 1–60,000 ms budget. Connected status uses the connection API; legacy knowledge status and projection commands remain available for existing installations.

Synchronization does not commit or push knowledge. BuildLore publication remains a separate, reviewable Git workflow.

To preserve a completed iteration independently of Wiki approval, explicitly capture it before opening the next iteration, then import the returned local bundlePath into BuildLore:

p2a knowledge capture --artifacts .plan2agent/artifacts/<project-id> --iteration <iteration-id> --json
p2a buildlore handoff import --file <capture-file> --commit --json
p2a buildlore handoff list --work-id <iteration-id> --json
p2a buildlore handoff read --id <handoff-id> --json
p2a buildlore handoff verify --id <handoff-id> --json

Use --task <maintenance-task-id> instead of --iteration for completed maintenance. Capture creates an unsanitized local retry bundle; do not commit or share it. BuildLore sanitizes before preservation and rejects suspected credentials. --commit explicitly commits only the preserved object locally; omitting it stores the object without creating a commit. Neither mode pushes, approves a Wiki, or changes the next development baseline. Reads work from the preserved object, but automatic capture, archived-source compilation, baseline restoration, and document cleanup are not yet connected: keep the original development artifacts (cleanupEligible: false). See the CLI reference for limits.

6. Manage project-scoped external Agent Skills

Plan2Agent uses the pinned skills@1.7.0 package as an isolated source adapter. Preview a Git or project-local skill before P2A copies validated regular files into the selected provider paths:

p2a skills source vercel-labs/agent-skills --list
p2a skills add vercel-labs/agent-skills \
  --skill web-design-guidelines \
  --tools codex,claude,gemini \
  --dry-run
p2a skills add vercel-labs/agent-skills \
  --skill web-design-guidelines \
  --tools codex,claude,gemini \
  --apply --expect-plan <dry-run-plan-sha256>
p2a skills list
p2a doctor --dev --strict

The tracked root p2a-skills.lock.json pins Git revisions and content/file hashes. The local manifest records separate external-skill:<name> ownership, so core P2A assets remain owned by init, update, and upgrade. Codex and Gemini share one .agents/skills copy; Claude receives an additional .claude/skills copy only when selected. Update and removal stop on local drift, and p2a skills sync restores missing copies while preserving modified or extra files. Every apply requires the reviewed --expect-plan digest. Review external skill instructions before applying them because they influence an agent with that agent's permissions. See the External Agent Skills guide for update, recovery, and trust boundaries.

CLI at a glance

Plan2Agent installs one p2a entrypoint:

Command Purpose
p2a init Initialize project state and provider assets.
p2a next Explain the current situation and recommend one next action; use --details for the internal command and state.
p2a decide Record Gate ① approvals, revocations, and scope changes in the decision ledger.
p2a decisions List decision history and trace a file to governing decisions with --why.
p2a shape Inspect, migrate, approve, and revoke the persistent project constitution.
p2a info Show project, artifact, task, and run status.
p2a doctor Diagnose configuration, assets, and local drift.
p2a update Apply project-managed assets pinned to the manifest package version.
p2a upgrade Preview or apply an npm-global package upgrade, then update the current project.
p2a enhance Enable optional capabilities such as BuildLore and proposals.
p2a validate Validate planning, task, run, eval, and proposal artifacts.
p2a iteration Manage iteration initialization, close/open cycles, diffs, and maintenance.
p2a tasks Inspect and transition task state.
p2a runs Record, verify, finish, and inspect run evidence.
p2a execute Supervise implementation and canonical final verification, visual, and acceptance runs through verified finish.
p2a eval Grade, compare, analyze, generate, and summarize execution evidence.
p2a proposals Mine and review proposals, or preview and explicitly publish a retrospective GitHub issue.
p2a buildlore Project, check, search, and retrieve optional BuildLore knowledge.
p2a knowledge capture Freeze completed work locally for a separate sanitized knowledge import.
p2a skills Preview, install, pin, update, restore, and remove external Agent Skills.

Run p2a --help for the top-level command surface and use the CLI Reference for detailed options and examples.

Safety model

Plan2Agent is contract-gated, confined, and local-first. Product meaning requires explicit approval; implementation choices and verification retries inside that approved envelope do not.

It is a good fit when you want:

  • explicit product decisions before implementation;
  • reviewable specs and task graphs only when orchestration benefits from them;
  • confined Codex or Claude execution, with Gemini kept read-only;
  • verification evidence and regression history;
  • human-approved maintenance and improvement loops.

It is not designed for:

  • unattended background coding;
  • unofficial provider API automation;
  • automatic dependency installation, merging, pushing, or PR creation without approval;
  • treating a remote service as the canonical project state;
  • replacing Git or an issue tracker.

Provider support

Canonical skills and subagent definitions live under .agents/. Plan2Agent generates and validates provider-specific surfaces for:

  • Codex
  • Claude Code
  • Gemini CLI

Parity checks keep provider mirrors aligned with the canonical definitions. The agent itself stays in the foreground tool session; Plan2Agent does not call provider APIs directly.

Companion projects

The core planning, validation, iteration, execution, eval, and proposal flows work without companion services.

Project Purpose
BuildLore Optional local-first, Git-backed projection and retrieval of Plan2Agent knowledge.
plan2agent-feature-radar Optional research workflow that exports evidence for planning without selecting requirements automatically.

Documentation

Developing Plan2Agent

Clone the repository and use Node.js 22.20.0 or newer. During development, run the core suite and the provider parity check:

npm test
node scripts/check_cli_parity.mjs

Before a PR or release, run npm run test:all. It runs the core suite, the long-running fixture gate, and the package/upgrade smoke tests once each. npm run test:full owns fixture and lifecycle coverage; node scripts/run_fixtures.mjs is its direct debugging equivalent, so do not run both in the same verification pass. Run node scripts/sync_cli_assets.mjs only when canonical provider assets changed, then confirm them with the parity check.

The runtime is Node.js ESM and uses the Node.js standard library. Repository structure:

.agents/       canonical skills and CLI-neutral subagents
.claude/       generated Claude Code mirrors
.codex/        generated Codex mirrors
.gemini/       generated Gemini CLI commands and agents
docs/          user guides and implementation references
fixtures/      golden and negative fixtures
schemas/       JSON schemas for Plan2Agent artifacts
scripts/       toolkit, validation, runtime, eval, proposal, and BuildLore adapter CLIs

Project status

Plan2Agent is under active development. Version 0.3.0 adds adaptive Direct, Planned, and Orchestrated execution, with Gate-derived execution envelopes and compatibility-preserving legacy orchestration. Detailed task graphs are now created only for Orchestrated work that benefits from dependency or ownership boundaries. The local-first planning, supervised execution, evaluation, proposal, and optional BuildLore workflows remain available. Autonomous provider execution and unapproved remote side effects remain outside the default safety model.

Plan2Agent is available under the MIT License.

About

Spec-driven development harness for AI coding agents with human-in-the-loop approval gates. Turns one-sentence ideas into approved specs, dependency-aware task graphs, and verified run logs for Claude Code, Codex, and Gemini CLI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages