Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
136 changes: 85 additions & 51 deletions ops/NEEDS_HUMAN.md
Original file line number Diff line number Diff line change
@@ -1,83 +1,117 @@
# NEEDS_HUMAN — Conflicting Work Package Context
# NEEDS_HUMAN — TARGET.md scopes already-completed work

**Situation:** This run has conflicting scope context that requires human clarification.
## The Problem

## The Conflict

1. **ops/TARGET.md says:** Gate 3, build hn-monitor runner (sub-PR A), `sdk/src/` code task
2. **ops/NEXT.md says:** Gate 3, cloud review-swarm preflight validation, `.github/workflows/` task
3. **These are completely different tasks** — one is SDK code (track A per TARGET), one is GitHub Actions (track D per NEXT)
ops/TARGET.md assigns work that was **already completed and merged** in PR #120 on 2026-09-01. The work package cannot proceed because the scope is internally contradictory.

## Evidence

**ops/TARGET.md line 1-5:**
**TARGET.md says (lines 1-7):**

```
# TARGET — gate 3

This run is pinned to **gate 3** and must not work on any other gate.

**Scope:** Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side.
**Scope:** Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling
runner in the SDK. CODE task, `sdk/src/`-side.
```

**ops/NEXT.md line 1-3:**
**ops/STATE.md says (lines 45-47):**

```
# NEXT — gate 3: complete cloud review-swarm preflight validation and documentation
- PR #120 (`201542a`, merged 2026-09-01 08:29 UTC) —
**`flows hn-monitor start`**, the CLI runner that turns the poller
into an unattended process.
```

**Verified in codebase:**

**Scope:** Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time
```bash
ls -l packages/sdk/src/cli/hn-monitor.ts packages/sdk/tests/cli-hn-monitor.test.ts
# -rw-r--r-- 1 daytona daytona 9468 Sep 15 12:59 packages/sdk/src/cli/hn-monitor.ts
# -rw-r--r-- 1 daytona daytona 18839 Sep 15 12:59 packages/sdk/tests/cli-hn-monitor.test.ts
```

## The Charter Says
The implementation at `packages/sdk/src/cli/hn-monitor.ts` contains all features TARGET.md requires:
- JournalClient construction + connection to relayflowd socket
- AgentWorker construction with workerAttach() before first poll
- Poll loop: pollHackerNewsOnce → sleep POLL_INTERVAL_MS → repeat
- Clean shutdown on AbortSignal (drain in-flight, close client, close worker)
- Fail-closed on journal errors (covenant 2), transient-tolerant on fetch errors
- All five findings from closed PR #83 addressed (per TARGET.md:9-22)

## Additional Contradictions

**1. Gate numbering confusion**

- TARGET.md:1 says "gate 3"
- TARGET.md:5 says "Gate 2 push"
- TARGET.md:71 acknowledges "the assessor on #83 confused itself" about gates

Per RFC-0001 §3:
- **Gate 2** = proactive agent (hn-monitor runs as relayflow) — AMBER per STATE.md
- **Gate 3** = Software Garden (factory DAG: discover → implement → review → merge) — RED/not started

Per charter/LEAD.md (the instruction I received):
- "Read ops/TARGET.md if it exists" — it does, says hn-monitor
- "Then read ops/STATE.md, ops/DIRECTIVES.md" — done
- "Then write ops/NEXT.md: the SINGLE highest-priority work package toward the current gate"
The hn-monitor runner is gate-2 work that's already complete.

But ops/NEXT.md ALREADY EXISTS with different work.
**2. Parallel execution risk**

## Additional Context Found
Per charter/LEAD.md:86-90, multiple drive runs execute in parallel, each pinned to a different gate. If I work on gate-2 follow-up while assigned to gate-3, I collide with any sibling run pinned to gate-2.

**ops/STATE.md gate 2 block (lines 39-81)** says:
- PR #120 merged 2026-09-01 — `flows hn-monitor start` CLI runner
- Gate 2 is AMBER, not GREEN
- Two clauses remain: trigger-plane liveness, analyze-agent execution
## What Remains for Gate 2

**Actual file check:**
- `packages/sdk/src/cli/hn-monitor.ts` exists (288 lines)
- Contains `runHnMonitor` function implementing all TARGET.md requirements
- Addresses all five findings from closed PR #83
ops/STATE.md:59-73 lists two unfinished gate-2 clauses (why gate 2 is AMBER, not GREEN):

1. **Trigger plane liveness-checked** — kernel does not detect when a poller stops. Needs RelayCron's deterministic-id claim + stale_after sweep pattern (RFC-0001 §3 gate 2: "a requirement, not an option").

2. **The analyze-agent step actually executing** — steps end in `worker_error` because AgentWorker has no user-supplied step handler. Dispatch works; analyzer doesn't. Whether this is gate-2 scope ("runs succeed") or gate-4 scope ("chief-as-relayflow supplies the runtime") is Khaliq's call.

These are valid gate-2 work, but they're NOT the work TARGET.md describes (which is the runner itself, now complete).

## The Question

**Which work package should this run execute?**
**What should this run work on?**

**Option A: Actual gate-3 work (Software Garden factory DAG migration)**
- Honor TARGET.md's gate label, ignore its description.
- Requires fresh assessment from RFC-0001 §3 gate-3 definition.
- No broken-down work package exists yet for gate 3.

**Option B: Gate-2 AMBER clause 1 (trigger plane liveness-checking)**
- Implement stale_after sweep + deterministic-id claim in kernel.
- This is gate-2 work, not gate-3 → collision risk with sibling gate-2 runs.

**Option C: Gate-2 AMBER clause 2 (analyze-agent step handler)**
- Add user-supplied step handler to AgentWorker for successful step completion.
- Also gate-2 work → same collision risk.

**Option D: Treat TARGET.md work as complete**
- The hn-monitor runner exists, PR merged, evidence filed.
- Report success, await new TARGET.md with fresh scope.
- If gate-3 is needed: scope Software Garden factory DAG.
- If gate-2 follow-up is needed: scope one of the two AMBER clauses explicitly.

## My Recommendation

**Option A: ops/TARGET.md wins (hn-monitor, sdk/ territory)**
- Overwrite ops/NEXT.md with hn-monitor work package
- But the code already exists (PR #120 merged per STATE.md)
- Task would be verification/testing, not building from scratch
**Option D** — The work TARGET.md describes is demonstrably complete. The confusion (gate label says 3, description says gate-2 work, references closed PR #83) suggests stale context.

**Option B: ops/NEXT.md wins (review-swarm, .github/ territory)**
- Ignore ops/TARGET.md
- Execute the review-swarm validation work
- But TARGET.md explicitly pins this run to gate 3 hn-monitor work
If continued gate-2 work is needed (AMBER → GREEN), the next TARGET.md should:
1. Say "gate 2" explicitly, not "gate 3"
2. Scope one of the two AMBER clauses by name
3. Clarify it's follow-up hardening, not new scaffolding

**Option C: TARGET.md is stale**
- The launcher wrote an outdated TARGET.md referencing closed PR #83
- Real work is in ops/NEXT.md (review-swarm)
- Proceed with review-swarm, update TARGET understanding
If gate-3 work is needed, the next TARGET.md should scope actual gate-3 work: Software Garden factory DAG migration per RFC-0001 §3.

**Option D: Both are stale**
- Neither accurately reflects current gate 3 needs
- Assessor should read RFC-0001 §3 gate 3 definition
- Write fresh work package from RFC requirements
## Why I Cannot Proceed

## Recommendation
Per charter/LEAD.md:86-90:

**Option C** — ops/TARGET.md appears stale (references closed PR #83 from earlier attempts, describes code that PR #120 already merged). The active work package is ops/NEXT.md (review-swarm). But I need human confirmation before overwriting NEXT.md or executing potentially wrong work.
> If gate 3 is genuinely unreachable from the current state, write
> ops/NEEDS_HUMAN.md saying exactly why and still end with ASSESS_DONE. Do not
> silently substitute different work: a run that reports progress on the wrong
> gate is worse than one that reports it is blocked.

## What I Need
The assigned gate-3 work is unreachable because it's already done. I cannot substitute gate-2 work without explicit authorization, or I collide with parallel runs. I cannot work on actual gate-3 (Software Garden) without a fresh scope breaking down that large, unstarted gate.

**Clear answer:** Which work package is correct for this run?
- If hn-monitor: shall I overwrite the review-swarm NEXT.md, or is there a different file I should write?
- If review-swarm: shall I proceed with ops/NEXT.md as-is and ignore TARGET.md?
- If neither: what is the actual gate 3 work I should assess?
**I need explicit clarification: which option above should this run execute?**
195 changes: 101 additions & 94 deletions ops/NEXT.md
Original file line number Diff line number Diff line change
@@ -1,123 +1,130 @@
# NEXT — gate 3 work package: document review-swarm secrets in README
# NEXT — BLOCKED: TARGET.md scopes already-completed work

**Scope (from TARGET.md):**
## Assessment

Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time, addressing every architectural finding from the walked-away #75/#77 attempts.
This run is **BLOCKED** on contradictory scope. Cannot proceed without human clarification.

## Objective
**Scope assigned in ops/TARGET.md (lines 1-5):**

Complete the final missing piece of gate 3's Definition of Done: document `RELAY_WORKSPACE_KEY` and `CLOUD_API_KEY` secrets in README.md with instructions on how to obtain them.
> # TARGET — gate 3
>
> This run is pinned to **gate 3** and must not work on any other gate.
>
> **Scope:** Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side.

## Current state assessment
**Problem 1: The work is already done.**

All 9 architectural requirements from TARGET.md are SATISFIED in the existing code:
Per ops/STATE.md:45-47, the hn-monitor CLI runner was merged in **PR #120** on 2026-09-01 08:29 UTC:

1. ✅ Immutable gate — two checkout steps (`.github/workflows/review-swarm.yml:32-53`)
2. ✅ Unified verdict logic — `swarm-verdict.sh` sourced by both callers
3. ✅ Auth secret validation — preflight validates all three secrets (lines 141-188)
4. ✅ Sticky marker + transcripts — HTML anchors with upsert_comment
5. ✅ No author whitelist — verified absent
6. ✅ Cloud sandbox fetch on GHA runner — `swarm-prepare.sh` with GH_TOKEN
7. ✅ Timeout ordering — 60m < 65m < 75m with comments
8. ✅ Wait step records status — swarm_status output, always() post step
9. ✅ Transcript freshness — run-start marker with stale detection

Verification commands all pass:
```
bash -n .github/workflows/scripts/swarm-post.sh && \
bash -n .github/workflows/scripts/swarm-prepare.sh && \
bash -n .github/workflows/scripts/swarm-verdict.sh && \
echo "All bash scripts parse OK"
# Output: All bash scripts parse OK

python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))" && \
python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))" && \
echo "YAML files parse OK"
# Output: YAML files parse OK

grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "No author whitelist found (GOOD)"
# Output: No author whitelist found (GOOD)

grep -c "actions/checkout@v4" .github/workflows/review-swarm.yml
# Output: 2
- PR #120 (`201542a`, merged 2026-09-01 08:29 UTC) —
**`flows hn-monitor start`**, the CLI runner that turns the poller
into an unattended process.
```

**The gap:** TARGET.md Definition of Done item 6 requires:
> README.md — document `RELAY_WORKSPACE_KEY` secret + how to obtain
The implementation exists at `packages/sdk/src/cli/hn-monitor.ts` (288 lines) with all features TARGET.md describes:
- Constructs JournalClient connected to relayflowd socket
- Constructs AgentWorker and calls workerAttach() before first poll
- Loops: pollHackerNewsOnce → sleep POLL_INTERVAL_MS → repeat
- Exits cleanly on AbortSignal (drain in-flight steps, close client, close worker)
- Fail-closed on journal errors (covenant 2), transient-tolerant on fetch errors
- All five findings from closed PR #83 addressed

Current reality:
```
grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md
# Output: 0
Test suite exists: `packages/sdk/tests/cli-hn-monitor.test.ts`

**Problem 2: Gate numbering is confused.**

TARGET.md:1 says "gate 3" but TARGET.md:5 says "Gate 2 push" and TARGET.md:71 acknowledges "the assessor on #83 confused itself" about gate numbers.

RFC-0001 §3 defines:
- **Gate 2** = proactive agent (hn-monitor runs as relayflow) — AMBER per STATE.md
- **Gate 3** = Software Garden (factory DAG: discover → implement → review → merge) — RED/not started

The hn-monitor runner is gate-2 work, not gate-3 work.

**Problem 3: Parallel execution collision.**

Per charter/LEAD.md:86-90:

> Several drive runs execute in parallel, each pinned to a different gate. Work outside this target collides with a sibling run, so staying inside it is not a preference — it is what makes parallel execution safe.

If I work on gate-2 follow-up (the two AMBER clauses in STATE.md:59-73), I'm outside my assigned gate-3 scope and collide with any sibling run pinned to gate 2.

If I work on actual gate-3 (Software Garden factory DAG), the TARGET.md description is nonsensical.

## What the previous run said

ops/NEEDS_HUMAN.md (from a previous assess run, same timestamp as TARGET.md) already identified this exact conflict and offered four options:
- A: TARGET.md wins (hn-monitor) — but code exists, task would be verification only
- B: Prior NEXT.md wins (review-swarm) — but TARGET pins to hn-monitor
- C: TARGET.md is stale (references closed PR #83) — NEXT.md is active work
- D: Both stale — write fresh from RFC-0001 gate-3 definition

That run ended blocked awaiting human decision. No decision was recorded.

## Current state verification

Verified the hn-monitor runner implementation is present and complete:

```bash
ls -l packages/sdk/src/cli/hn-monitor.ts packages/sdk/tests/cli-hn-monitor.test.ts
# -rw-r--r-- 1 daytona daytona 9468 Sep 15 12:59 packages/sdk/src/cli/hn-monitor.ts
# -rw-r--r-- 1 daytona daytona 18839 Sep 15 12:59 packages/sdk/tests/cli-hn-monitor.test.ts

grep -c "export async function runHnMonitor" packages/sdk/src/cli/hn-monitor.ts
# 1

grep -c "PR #120" ops/STATE.md
# 1
```

README.md does NOT document these secrets. The workflow comment (`.github/workflows/review-swarm.yml:21-24`) references a runbook in the `AgentWorkforce/cloud` repo, but README has no such documentation.
## The two remaining gate-2 AMBER clauses

ops/STATE.md:59-73 names two unfinished gate-2 requirements that COULD be valid work, but they're gate-2, not gate-3:

1. **Trigger plane liveness-checked** — kernel does not yet detect when a poller stops (needs RelayCron's deterministic-id claim + stale_after sweep pattern per RFC-0001 §3 gate 2).

2. **The analyze-agent step actually executing** — in the recorded run every step ended in `worker_error` because `hn-monitor start`'s AgentWorker has no user-supplied step handler. The dispatch loop works; the analyzer does not.

From `ops/NEEDS_HUMAN.md`, the secrets are stored and working (as of 2026-09-07), but gate 3 is blocked on Daytona CPU quota, not on implementation. The workflow WORKS; the documentation is missing.
## Question for operator

## Files in scope
**Which of these should this run work on?**

- `README.md` — add section documenting GitHub Actions secrets required for review-swarm
**Option A:** Actual gate-3 work (Software Garden factory DAG migration).
- Requires fresh assessment from RFC-0001 §3 gate-3 definition.
- No broken-down work package exists yet.

## Work package
**Option B:** Gate-2 AMBER clause 1 (trigger plane liveness-checking).
- Add stale_after sweep + deterministic-id claim to kernel.
- This is gate-2 work → collision risk with sibling gate-2 runs.

Add a "GitHub Actions Secrets" section to README.md documenting:
**Option C:** Gate-2 AMBER clause 2 (analyze-agent step handler).
- Add user-supplied step handler to AgentWorker so hn-monitor steps complete successfully.
- Also gate-2 work → same collision risk.

1. `RELAY_WORKSPACE_KEY` — Agent Relay workspace key for review swarm communication
- How to obtain: Contact repository administrator or see ops/NEEDS_HUMAN.md for historical context
- Why required: Enables agent coordination within review swarm workflow
**Option D:** Treat work as complete.
- The hn-monitor runner exists, PR merged, evidence filed.
- Report success, await new TARGET.md with different scope.

2. `CLOUD_API_KEY` — Agent Relay Cloud API credential for launching cloud workflows
- How to obtain: Minted per `AgentWorkforce/cloud → docs/runbooks/relay-ci-workflow-credential.md`
- Profile: `workflow-invoke`
- Scopes: `workflow:invoke:read` and `workflow:invoke:write`
- How to store: Repository Settings → Secrets and variables → Actions → New repository secret
## Recommendation

3. `CLOUD_API_URL` — Cloud API endpoint (typically `https://agentrelay.com/cloud`)
- Usually set as repository variable, not secret
- Defaults to production endpoint if not set
**Option D** — The described work (hn-monitor CLI runner) is complete and merged. The confusion between gate 2 and gate 3 suggests the TARGET.md was generated from stale context (references closed PR #83).

The section should be brief (10-15 lines) and reference the workflow files for implementation details.
If gate-3 work is genuinely needed, a fresh TARGET.md should scope actual gate-3 work (Software Garden factory DAG per RFC-0001 §3).

## Definition of done
If gate-2 follow-up is needed (AMBER → GREEN), a new TARGET.md should:
1. Explicitly say "gate 2" not "gate 3"
2. Scope one of the two AMBER clauses specifically
3. Note it's follow-up work, not new scaffolding

1. README.md contains a section documenting the three secrets/variables
2. Each entry states what it is and how to obtain it
3. Parse checks continue to pass:
```
bash -n .github/workflows/scripts/swarm-*.sh
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))"
python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))"
```
4. Verification remains true:
```
grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md
# Should return > 0
grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "GOOD"
# Should return "GOOD" or nothing (no whitelist)
```
5. As final action:
```
git status --porcelain
```
## Files in scope (if Option D rejected)

## Explicitly OUT of scope
Cannot be determined without operator decision on which option.

- `.github/workflows/review-swarm.yml` (already correct, all 9 requirements satisfied)
- `workflows/review-swarm.yaml` (already correct)
- `.github/workflows/scripts/swarm-*.sh` (all already correct)
- `.gitignore` (no .review-target mask exists, already correct)
- `sdk/` (Track A owns that)
- `kernel/` (gate 1 done)
- `ops/*` (chief owns briefs and state)
- Any other GHA workflow
- Resolving the Daytona CPU quota block (that's in ops/NEEDS_HUMAN.md, different issue)
- Actually testing the workflow end-to-end (blocked on Daytona capacity per ops/NEEDS_HUMAN.md)
## Definition of done (if Option D rejected)

## Why this is the work package
Cannot be determined without operator decision on which option.

TARGET.md's Definition of Done explicitly lists:
- Item 6: "PR body explicitly documents each of the 9 requirements above and shows where each is satisfied"
- Item 7: "`README.md` — document `RELAY_WORKSPACE_KEY` secret + how to obtain"
## Out of scope

The 9 requirements are satisfied in code. Item 7 is not satisfied. This is the remaining gap between current state and TARGET.md's done-when.
Everything, until scope is clarified.
Loading