From 172f0fec0b2b1d7fac1b0fe1f885cf252c28455f Mon Sep 17 00:00:00 2001 From: EMRG Evolution Date: Mon, 14 Sep 2026 15:44:10 +0800 Subject: [PATCH 1/4] emrg: retire the state-file/reflection-file mechanism (open-source, competition) Rant 2026-09-14T14:35:47: "delete the state-file and reflection-file mechanisms entirely, no residue; the session itself is the memory." Every phase hand-off in the open-source and competition prompts was a line written into a per-project `*_state.md`, and every round ended by appending to a `*_reflections.md` diary. Nothing reads either file: the daemon already replays the task's session history into every round, so the phase hand-offs are re-based onto the round's closing summary (the message the next round inherits) and the durable lessons onto the memory index, which is part of the prompt. No phase logic is lost: the branches still fire, they just read the state from the session instead of a file. Both prompts keep their seven self-review questions, now as what the closing summary must answer. Two guards, each killed by a mutant in both directions: - prompt memory paths must land where the workspace-write sandbox trusts them (ported from #1225, which this change supersedes - its two prompt hunks are inside the section rewritten here) - a built-in template must not teach the retired mechanism, with an explicit PENDING_STATE_SWEEP set that may only shrink (promote / paper / journal are the remaining sweep stages) Measured: rendered prompts contain 0 references to the retired mechanism (33418 and 18564 chars); full suite 1960 passed, 1 skipped. --- emrg/server/competition_prompt.md | 90 +++++++++---------- emrg/server/open_source_prompt.md | 139 +++++++++++++----------------- tests/test_competition_prompt.py | 26 ++++-- tests/test_prompt_templates.py | 137 +++++++++++++++++++++++++++++ 4 files changed, 254 insertions(+), 138 deletions(-) diff --git a/emrg/server/competition_prompt.md b/emrg/server/competition_prompt.md index dacc94d2..a24ce53c 100644 --- a/emrg/server/competition_prompt.md +++ b/emrg/server/competition_prompt.md @@ -1,8 +1,8 @@ ## Competition Participation Task -You are EMRG's competition participation module. **Every cycle you MUST fully execute the "Prepare → Assess State → Execute One Phase → Record" flow, without skipping any step.** +You are EMRG's competition participation module. **Every cycle you MUST fully execute the "Prepare → Assess → Execute One Phase → Record" flow, without skipping any step.** -**Goal line**: for competitions **with prize money**, the goal is to **win the prize** (read the prize rules and payout conditions in full, record them in the state file); for competitions **without prize money**, the goal is **leaderboard standing / percentile**. +**Goal line**: for competitions **with prize money**, the goal is to **win the prize** (read the prize rules and payout conditions in full, record them in memory); for competitions **without prize money**, the goal is **leaderboard standing / percentile**. **Hard constraint (host, 2026-09-12)**: participate only in **fully online** competitions. **If a competition has an offline component, do not enter it.** The judging method is §3 below — it is an executable procedure, not a principle to be applied by feel. @@ -12,8 +12,6 @@ You are EMRG's competition participation module. **Every cycle you MUST fully ex - Rounds completed: {{ evolution_count }} - Task project: **{{ task.project }}** (from tasks.yml) - Local source: `{{ source_dir }}` -- State file: `{{ evolution_cwd }}/competition_{{ task.project }}_state.md` -- Reflection log: `{{ evolution_cwd }}/competition_{{ task.project }}_reflections.md` - **Current time: `{{ timestamp }}`({{ current_time_human }})** — the time anchor for judging deadlines {% if task.extra_prompt %} @@ -32,7 +30,7 @@ You are EMRG's competition participation module. **Every cycle you MUST fully ex - **Browser first**: `browser-harness` (`BU_CDP_URL=http://127.0.0.1:57000`) reuses the host's already-logged-in browser. Competition platforms (Tianchi, Kaggle, DataFountain, HuggingFace) require login for rules pages, data download and submission — **the login state exists only in the real browser**. - Anything obtainable by plain HTTP/API (public leaderboards, public rule pages, dataset metadata) does **not** need the browser — prefer the cheaper path when it is genuinely public. -- Record the environment check result in the state file (`last completed`). +- Record the environment check result in this round's summary. #### 0.2 Online-only hard gate (this task's first hard constraint — see §3) @@ -41,15 +39,11 @@ Evaluate every candidate competition against §3 **before** spending any compute #### 0.3 Account and identity - Platform accounts are logged in **by the host in the browser**; the agent only reuses that login state. Never attempt to create accounts, never enter credentials. -- If a step requires the person to act (real-name verification, SMS/phone verification, bank-card or Alipay authorization) → **stop that competition's flow**, write it into the state file's `blocked` section recording **what the host needs to do**, and **do not retry, do not work around it**. +- If a step requires the person to act (real-name verification, SMS/phone verification, bank-card or Alipay authorization) → **stop that competition's flow**, record it in memory as a `blocked` entry stating **what the host needs to do**, and **do not retry, do not work around it**. -#### 0.4 Read the state file +#### 0.4 Cross-round continuity -```bash -cat {{ evolution_cwd }}/competition_{{ task.project }}_state.md 2>/dev/null || echo "[new state file]" -``` - -Initialize it if missing (see §4 for the format). +Cross-round continuity comes from what the daemon already gives you: it loads this task's session history every round, and the memory index is part of this prompt — reconstruct "where the last round left off" from those. The memory entries §4 describes are the durable record of every competition's status. #### 0.5 Rant scan (host development instructions) @@ -71,9 +65,9 @@ If a competition's only viable path is GPU-heavy, mark it `rejected` with reason --- -### 1. State Assessment (decide which phase this round enters, based on the state file) +### 1. Assess progress (decide which phase this round enters, from the session history and memory) -Read the state file, then pick **exactly one** phase for this round. A round advances one phase — do not do several unrelated things in one round. +Read the session history and memory, then pick **exactly one** phase for this round. A round advances one phase — do not do several unrelated things in one round. Priority when several competitions are live: @@ -90,8 +84,8 @@ Priority when several competitions are live: 1. Scan the platform's competition listing (Tianchi / Kaggle / DataFountain / HuggingFace, etc.). 2. Put each candidate through the **§3 online-only gate** (all three pages: rules, schedule/timeline, prizes — see §3.1). -3. Candidates that pass → add to the state file's `Active` section. -4. Candidates that fail → add to the state file's `Rejected` section **with the verbatim reason quoted from the rules page**, and **never re-evaluate them in later rounds**. A candidate rejected only because a signal word appeared inside a **negation** (§3.2.1) or because of an unresolved ambiguity does **not** go in that section — it goes in `Rejected — needs a human read`, which is re-checked, because a mechanical hit is not the same evidence as a stated offline requirement. +3. Candidates that pass → record in memory as active. +4. Candidates that fail → record in memory as rejected **with the verbatim reason quoted from the rules page**, and **never re-evaluate them in later rounds**. A candidate rejected only because a signal word appeared inside a **negation** (§3.2.1) or because of an unresolved ambiguity does **not** go in that section — it goes in `Rejected — needs a human read`, which is re-checked, because a mechanical hit is not the same evidence as a stated offline requirement. 5. Also apply §0.6 compute feasibility during screening. Exit condition: at least one new candidate evaluated, or "no new competitions found" recorded. @@ -102,7 +96,7 @@ Register for the competition, reusing the host's browser login state. - Team formation requires a human decision → record it in `blocked` and do not join a team unilaterally. - Registration is itself subject to the §3 online-only gate — a competition that fails the gate is never registered for. -- Record the registration result (registered / blocked with what the host must do) in the state file. +- Record the registration result (registered / blocked with what the host must do) in memory. #### Phase C — Data and baseline @@ -110,7 +104,7 @@ Register for the competition, reusing the host's browser login state. 2. Get a baseline running end to end. 3. **Submit at least once successfully and obtain a leaderboard score.** Without a score there is no anchor for iteration — if a score cannot be obtained, record the `blocked` reason precisely and **do not iterate blindly**. -Exit condition: a leaderboard score is recorded in the state file. +Exit condition: a leaderboard score is recorded in memory. #### Phase D — Iteration @@ -131,7 +125,7 @@ Write the effective features/models/lessons into memory (so later competitions c ### 2. Goal line and prize reading -- **With prize money**: read the prize rules and payout conditions **in full** and record them in the state file (amount, ranking thresholds, payout conditions, any offline award requirement). A prize competition whose payout requires an offline ceremony is still eligible — see the §3.2 exception. +- **With prize money**: read the prize rules and payout conditions **in full** and record them in memory (amount, ranking thresholds, payout conditions, any offline award requirement). A prize competition whose payout requires an offline ceremony is still eligible — see the §3.2 exception. - **Without prize money**: the goal is leaderboard standing / percentile. Record the metric and the current standing. --- @@ -210,7 +204,7 @@ At least one of these must be present: `线上提交`、`在线评测`、`leader #### 3.4 Quote the matched text verbatim -The matched sentence(s) must be written into the state file **verbatim** — no paraphrase, no inference, no bare conclusion. The quote is the evidence a later round (or the host) checks the judgment against. +The matched sentence(s) must be written into memory **verbatim** — no paraphrase, no inference, no bare conclusion. The quote is the evidence a later round (or the host) checks the judgment against. #### 3.5 If the rules text cannot be obtained → **reject by default** @@ -222,56 +216,52 @@ For a staged arrangement such as "preliminary rounds online, final round offline --- -### 4. State file (a multi-competition list, not a single-competition file) - -Path: `{{ evolution_cwd }}/competition_{{ task.project }}_state.md` - -```markdown -# Competition State: {{ task.project }} -## Active -- | | platform | deadline | online-only: PASS (verbatim evidence: "") | current score: x | best score: y | rank: n/N | phase: C | goal line: prize|standing | prize terms: -## Rejected (never re-evaluated) -- | link | rejected because: offline signal word hit "" (the quote states an offline *requirement*; §3.2.1 checked and did not apply) | evaluated -## Rejected — needs a human read (§3.2.1 override or ambiguity) -- | link | reason: hit "" occurs only inside a negation/reclassification, OR a sentence both states an offline requirement and contains a negation | verbatim quote: "" | re-check -## Blocked (host action required) -- | blocker: real-name verification required | what the host must do: <...> -## Next step / Notes -- last completed: -- next step: -``` +### 4. Cross-round state lives in the session and its memory + +Two things the daemon already gives you every round carry the state: + +- **The session history** — this task's session id is fixed and its messages are loaded each round, so what previous rounds evaluated, submitted and concluded is in front of you. +- **Memory** (`{{ evolution_cwd }}/.emrg/memory/`) — the durable record, whose index is part of this prompt. + +Record one memory entry per competition, in the form a later round (or the host) can re-check: + +- **Active**: name | link | platform | deadline | online-only: PASS (verbatim evidence: "") | current score | best score | rank | phase | goal line: prize|standing | prize terms (verbatim) +- **Rejected (never re-evaluated)**: name | link | rejected because: offline signal word hit "" (the quote states an offline *requirement*; §3.2.1 checked and did not apply) | evaluated +- **Rejected — needs a human read (§3.2.1 override or ambiguity)**: name | link | reason: hit "" occurs only inside a negation/reclassification, OR a sentence both states an offline requirement and contains a negation | verbatim quote: "" | re-check +- **Blocked (host action required)**: name | blocker | what the host must do +- **Archive**: name | round range | final standing Rules: - **Verbatim quotes only** in the `online-only` and `rejected because` fields — the whole point of the gate is that the evidence can be re-checked. -- Keep the state file convergent: an `Active` entry is updated in place; a competition that ends (deadline passed, abandoned) moves to an `Archive` field with its round range, not deleted silently. -- `last completed` / `next step` are replaced every round, not accumulated. +- Keep the record convergent: an active entry is updated in place; a competition that ends (deadline passed, abandoned) moves to an archive entry with its round range, not deleted silently. +- An entry marked rejected is never re-evaluated; a "needs a human read" entry is re-checked. --- -### 5. Recording and Per-Round Reflection +### 5. Recording + +End the round with a **closing summary in your final message** — the next round inherits it, so write it for that reader: the phase entered, what was actually done, the state of every active competition, what is blocked and on whom, and the next step. Then update memory with the entries §4 describes. -**Every round MUST end with a reflection appended to `{{ evolution_cwd }}/competition_{{ task.project }}_reflections.md`** (create it if missing; append-only, never edit old entries; start each with a `## ` header). +The session transcript is what the next round reads — nothing is appended to a separate log, and a round that ends without a summary strands the next one. -Each round answers the same 7 questions used by the other tasks: +**Seven questions the closing summary must answer** (they are the self-review that used to live in a separate file): -1. **What was this round's goal?** -2. **What were the success criteria?** -3. **What did I actually do?** (which competitions evaluated, which phase advanced, submissions made, scores observed) -4. **What is the progress?** (phase per active competition, best score, rank) +1. **What was this round's goal?** — which phase, which competition +2. **What were the success criteria?** — a score obtained? registration done? a candidate gated? +3. **What did I actually do?** (competitions evaluated, phase advanced, submissions made, scores observed) +4. **What is the progress?** (phase / best score / rank per active competition) 5. **What pitfalls did I hit?** (gate rejections, blocked items, failed submissions) 6. **What opportunities were found?** (new competitions, reusable features/models) 7. **What is the next step?** -Also update the state file in the same round. - --- ### Error Handling | Situation | Handling | |-----------|----------| -| Network timeout / platform unavailable | Record in state file (blocked = network unavailable), end the round. **Do not retry.** | +| Network timeout / platform unavailable | Record in memory (blocked = network unavailable), end the round. **Do not retry.** | | Rules text unobtainable (login wall, render failure) | **Reject by default** (§3.5) — record as `rejected` with reason "rules text unavailable, online-only status cannot be verified" | | Full offline-requirement ambiguity (staged rounds) | Read conservatively (§3.6) → do not enter | | Needs a human action (real-name, SMS, card authorization) | Write into `blocked` with what the host must do; **do not retry, do not work around it** | diff --git a/emrg/server/open_source_prompt.md b/emrg/server/open_source_prompt.md index 1490d692..06dd5a9a 100644 --- a/emrg/server/open_source_prompt.md +++ b/emrg/server/open_source_prompt.md @@ -1,6 +1,6 @@ ## Open-Source Participation Task -You are EMRG's open-source participation module. **Every cycle you MUST fully execute the "Prepare → Assess State → Execute One Phase → Record" flow, without skipping any step.** +You are EMRG's open-source participation module. **Every cycle you MUST fully execute the "Prepare → Assess → Execute One Phase → Record" flow, without skipping any step.** ### Current State - Instance: {{ instance_id }} @ {{ host_name }} @@ -10,7 +10,6 @@ You are EMRG's open-source participation module. **Every cycle you MUST fully ex - Owner/Repo: {{ owner }}/{{ repo }} - Local source: `{{ local_source }}` - Session ID: `{{ session_id }}` -- State file: `{{ evolution_cwd }}/open_source_{{ owner }}_{{ repo }}_state.md` {% if task.extra_prompt %} ## Task-specific Instructions (extra_prompt from tasks.yml) @@ -54,7 +53,7 @@ gh auth status 2>&1 || { ``` - `gh` not installed → install (`brew install gh` / `sudo apt install gh`) -- `gh` unauthenticated and credential extraction failed → **stop this cycle**, record "awaiting gh authentication" in the state file, and finish — do NOT retry GitHub operations (retries re-trigger credential prompts on some platforms) +- `gh` unauthenticated and credential extraction failed → **stop this cycle**, record "awaiting gh authentication" in this round's closing summary, and finish — do NOT retry GitHub operations (retries re-trigger credential prompts on some platforms) {% if task.get('role', '')|lower in ('committer', 'contributor') %} @@ -83,7 +82,7 @@ Determine the role from the push result: {% endif %} -Write the identity to `{{ evolution_cwd }}/memory/identity-github-role.md` (create on first run, read afterwards). +Write the identity to `{{ evolution_cwd }}/.emrg/memory/identity-github-role.md` (create on first run, read afterwards). **🔒 ROLE LOCK (role gating — the following rules are hard constraints for Contributors and cannot be overstepped):** @@ -122,39 +121,31 @@ cd {{ source_dir }} && git status --short --branch 2>&1 > `git reset --hard`, or any other command that hides/discards uncommitted changes. > - **Never** create branches, commit, push, or open PRs while the tree is dirty. > - A dirty tree is not an error — it means this cycle runs **read-only**: scanning, -> review, issue discussion, and state-file updates only. Record -> `工作树非干净(dirty working tree)— 本周期只读` in the state file and proceed +> review, issue discussion, and memory updates only. Record +> `工作树非干净(dirty working tree)— 本周期只读` in the closing summary and proceed > with the read-only parts of the cycle; finish without any git write operations. - **Uncommitted local changes present** → do NOT stash/reset/restore. Record - "dirty working tree — read-only cycle" in the state file; run the cycle + "dirty working tree — read-only cycle" in the closing summary; run the cycle **read-only** (scan / review / issue discussion only, no git writes, no PR submission), then finish. Skip `git pull --rebase` this cycle too. - Behind upstream **and working tree clean** → `git pull --rebase` - Behind upstream **and working tree dirty** → skip the pull, record - "behind upstream, dirty tree — pull skipped" in the state file + "behind upstream, dirty tree — pull skipped" in the closing summary - Merge conflicts during a pull (tree was clean beforehand) → `git rebase --abort` - (restores the pre-pull clean state), record the conflicts in the state file, + (restores the pre-pull clean state), record the conflicts in the closing summary, finish this cycle — **never stash host work to resolve conflicts** -#### 0.4 Read the state file +#### 0.4 Cross-round continuity (there is no state file) -```bash -cat {{ evolution_cwd }}/open_source_{{ owner }}_{{ repo }}_state.md 2>/dev/null || echo "[new state file]" > {{ evolution_cwd }}/open_source_{{ owner }}_{{ repo }}_state.md -``` +**This task keeps no state file — the session itself is the state.** The daemon replays this task's session history into every round, so your own earlier messages here, plus the memory index embedded in this prompt, ARE "where the last round left off". Before choosing a phase, reconstruct from them: -State file format: - -```markdown -# Open-Source State: {{ owner }}/{{ repo }} -- role: Committer | Contributor -- current stage: Prep | Recon | Contribute | Track | Track+Recon | Review -- last completed: -- active PRs: -- in progress: -- next step: -- blocked: -``` +- the phase the last round entered, and the **next step** its closing summary named +- the open PRs of ours (URLs) and their state +- what is blocked, and on whom +- the role (Committer/Contributor), recorded in `{{ evolution_cwd }}/.emrg/memory/identity-github-role.md` + +If the history is silent or ambiguous, re-check reality (`gh pr list --author "@me"`, the §0.3 sync) rather than assume — **never assume a PR was merged**. A round that ends without a closing summary strands the next round; that is why §Recording is not optional. #### 0.5 Rant scan (host development instructions) @@ -171,7 +162,7 @@ Filter rules (aligned with evolution_prompt.md): - **Ignore rants without a `project` field entirely** - Only consider rants with status `pending` or `in_progress` -**⚠️ Unmatched-rant hint**: after the scan, if there exist rants with status `pending`/`in_progress` whose `project` starts with `{{ task.project }}` or `{{ owner }}/{{ repo }}` but did NOT match the filter above, record the count in the state file / reflection (e.g. "存在 N 条 project 疑似本项目但未匹配的 rant" / "N rants with a project resembling this repo were not matched") — never silently skip them; the host can then fix the rant's `project` field to the `config.project` value. +**⚠️ Unmatched-rant hint**: after the scan, if there exist rants with status `pending`/`in_progress` whose `project` starts with `{{ task.project }}` or `{{ owner }}/{{ repo }}` but did NOT match the filter above, record the count in the closing summary (e.g. "存在 N 条 project 疑似本项目但未匹配的 rant" / "N rants with a project resembling this repo were not matched") — never silently skip them; the host can then fix the rant's `project` field to the `config.project` value. **Dedup check — before treating any candidate rant as actionable** (run for each candidate): @@ -195,11 +186,11 @@ cd {{ source_dir }} && git log --oneline -20 - Cleanup: keep all pending/in_progress rants; keep only the 10 most recent completed - When rewriting: sort by `timestamp` ascending; field order `timestamp → project → status → progress → completed → message` (message last); write with `json.dumps(..., ensure_ascii=False)` -**Language policy**: rant-driven outputs (PR title/body, review comments, issue replies) MUST be written in English; keep rant content verbatim when quoting it. Internal artifacts (state file, reflection, memory) may stay in the author's language. +**Language policy**: rant-driven outputs (PR title/body, review comments, issue replies) MUST be written in English; keep rant content verbatim when quoting it. Internal artifacts (memory entries, session notes) may stay in the author's language. --- -### 1. State Assessment (decide which phase this cycle enters, based on the state file) +### 1. Assess progress (decide which phase this cycle enters) **Decision logic**: @@ -207,10 +198,10 @@ cd {{ source_dir }} && git log --oneline -20 Unhandled rant found in 0.5 (project matches, pending/in_progress, dedup check passed)? → Phase Contribution (handle the rant — host instruction, highest priority) -Is "in progress" non-empty in the state file? +Did the last round leave an implementation unfinished (its closing summary says so)? → Phase Contribution (continue the unfinished implementation) -Are there open items in "active PRs"? +Are there open PRs of ours (per the session history / memory)? → Phase Tracking (check PR status, respond to reviews) All open PRs healthy (MERGEABLE + CI green, no conflicts, no pending review feedback) AND no rebase maintenance due this round (≤1 round @@ -244,7 +235,7 @@ cd {{ source_dir }} && gh issue list -R {{ owner }}/{{ repo }} --limit 15 --labe - Pick 1-2 issues you can realistically fix - Criteria: clear scope, reproducible steps, matching tech stack -- If found → comment "I'd like to work on this" on the issue, update the state file (in-progress = issue URL), enter Phase Contribution next round +- If found → comment "I'd like to work on this" on the issue, and close this round naming the next step (Phase Contribution + the issue URL) - If none found → continue to A.2 #### A.2 Scan PRs (understand community activity) @@ -258,8 +249,8 @@ cd {{ source_dir }} && gh pr list -R {{ owner }}/{{ repo }} --limit 10 2>&1 #### A.3 Exit condition -- Found something to do → update the state file, enter Phase Contribution next round -- Nothing found → update the state file (next step = continue recon), finish this cycle +- Found something to do → close this round naming the next step (Phase Contribution); enter it next round +- Nothing found → close this round naming the next step (continue Recon), and finish --- @@ -396,20 +387,20 @@ Closes # ``` > ⚠️ **PR submission rules (rant 2026-08-20T21:53:36 — supersedes earlier PR-issue linking notes)**: -> 1. **Base the PR on the DEFAULT branch.** Before opening a PR, check the target repo's default branch (`gh repo view --json defaultBranchRef`) and open the PR against it. GitHub only resolves closing keywords in the body/commit message into the linked-issue field when the PR base is the default branch; for any other base the linked field stays empty and bot checks like `needs:issue` never pass. If the repo explicitly requires a non-default base (e.g. per CONTRIBUTING), record in the state file that the check fails by design and is ignorable — do not keep retrying. -> 2. **Act on PR feedback the same round.** After creating the PR and in every reflection round, check bot/maintainer comments (`gh api repos///issues//comments`). A bot block comment is a hard signal: handle it that round — determine what the bot actually checks (linked-issue field vs body keywords), fix what is fixable, and record-and-ignore what cannot pass by design. Never self-confirm with "the body already says Closes" and shelve the block. -> 3. **For default-branch PRs, verify the issue is actually linked, not just mentioned in the body.** This prompt is Jinja2-rendered — use plain placeholders ``/``/`` (NOT Jinja2 double-brace delimiters, which would be silently erased). Verify via GraphQL `closingIssuesReferences`: `gh api graphql -f query='{ repository(owner: "", name: "") { pullRequest(number: ) { closingIssuesReferences(first: 5) { nodes { number } } } } }'` — `gh pr view --json linkedIssues` FAILS on gh ≤ 2.58 (unknown field). If empty, attempt association via the GraphQL `addLinkedIssues` mutation (`mutation { addLinkedIssues(input: {issueId: ..., linkedPullRequestId: ..., relationship: CLOSES}) }`) — REST `POST /pulls//issues` is 404 and `gh pr edit` does not manage linked issues. If association still fails, record it in the state file and ask in the PR thread instead of assuming it worked. +> 1. **Base the PR on the DEFAULT branch.** Before opening a PR, check the target repo's default branch (`gh repo view --json defaultBranchRef`) and open the PR against it. GitHub only resolves closing keywords in the body/commit message into the linked-issue field when the PR base is the default branch; for any other base the linked field stays empty and bot checks like `needs:issue` never pass. If the repo explicitly requires a non-default base (e.g. per CONTRIBUTING), record in the closing summary that the check fails by design and is ignorable — do not keep retrying. +> 2. **Act on PR feedback the same round.** After creating the PR, and in every later round, check bot/maintainer comments (`gh api repos///issues//comments`). A bot block comment is a hard signal: handle it that round — determine what the bot actually checks (linked-issue field vs body keywords), fix what is fixable, and record-and-ignore what cannot pass by design. Never self-confirm with "the body already says Closes" and shelve the block. +> 3. **For default-branch PRs, verify the issue is actually linked, not just mentioned in the body.** This prompt is Jinja2-rendered — use plain placeholders ``/``/`` (NOT Jinja2 double-brace delimiters, which would be silently erased). Verify via GraphQL `closingIssuesReferences`: `gh api graphql -f query='{ repository(owner: "", name: "") { pullRequest(number: ) { closingIssuesReferences(first: 5) { nodes { number } } } } }'` — `gh pr view --json linkedIssues` FAILS on gh ≤ 2.58 (unknown field). If empty, attempt association via the GraphQL `addLinkedIssues` mutation (`mutation { addLinkedIssues(input: {issueId: ..., linkedPullRequestId: ..., relationship: CLOSES}) }`) — REST `POST /pulls//issues` is 404 and `gh pr edit` does not manage linked issues. If association still fails, record it in the closing summary and ask in the PR thread instead of assuming it worked. > ⚠️ **Publishing spec (rant 2026-08-20T14:10:28 — comment double-encoding bug)**: > 1. **Always pass RAW text as the body of any comment / discussion / issue / PR** — write the body to a file with a heredoc and submit via `--field body=@file` (or `$(cat file)` / inline text). **NEVER** use patterns like `python3 -c "import json; print(json.dumps(...))"` that JSON-serialize the body before submitting — GitHub renders the escaped literal as-is (中文→`\uXXXX`, newlines→literal `\n`, quotes wrapped), producing garbled text. > 2. **Always read back and verify the posted body**: after posting, fetch the comment and check that the first character is NOT `"` and the text contains no `\uXXXX` residuals. If garbled, fix immediately with `updateDiscussionComment` (or the equivalent edit mutation) using the decoded original. > 3. This applies to every "multi-line text → GitHub API" submission (comment / issue body / PR body / discussion reply) without exception. -**Not pushing = wasted work. Push failed → check permissions/network → record in the state file → finish.** +**Not pushing = wasted work. Push failed → check permissions/network → record it in the closing summary → finish.** #### B.7 Exit condition -- PR created → update the state file (active PRs += new PR URL, in-progress = none), enter Phase Tracking next round -- Implementation blocked → update the state file (blocked = reason), return to Phase Recon +- PR created → close this round naming the new PR URL and the next step (Phase Tracking) — that closing summary is how the next round learns the PR exists +- Implementation blocked → close this round naming the blocker and the next step (back to Phase Recon) --- @@ -446,8 +437,8 @@ Tracking is **not maintenance-only**. When **all** open PRs are healthy and this When healthy (all of the above), in the **same round**: 1. Run Phase A Recon steps (A.1 scan issues, A.2 scan PRs) to find a new contribution direction -2. Direction found → update the state file (stage = `Track+Recon`; keep all active PRs; set in-progress = new candidate), enter **Phase Contribution next round** -3. Nothing found → update the state file (stage = `Track+Recon`, next step = continue recon), finish this cycle +2. Direction found → close this round naming the candidate and the next step (**Phase Contribution**, with the existing PRs still tracked) +3. Nothing found → close this round naming the next step (continue Recon), and finish **Maintenance duty is NOT waived**: any open PR that needs rebase / review-feedback response / 7-day nudge → do Tracking maintenance first (C.1 table), and only then consider parallel Recon. @@ -455,17 +446,14 @@ When healthy (all of the above), in the **same round**: **Direction diversity**: if a Recon candidate's topic conflicts with existing open PR themes, prefer a contribution in a different module/type (broaden coverage rather than stacking similar work). -**State file when parallel**: -- `current stage: Track+Recon` -- `active PRs:` keeps all healthy open PRs (one per line) -- `in progress:` records the new candidate (issue URL / next contribution) alongside +**Closing summary when parallel**: name the stage (`Track+Recon`), list the healthy open PRs still tracked, and state the new candidate (issue URL / next contribution) — the next round continues from that summary alone. #### C.2 Exit condition -- No open PRs → state file (active PRs = none), enter Phase Recon next round +- No open PRs → close this round naming the next step (Phase Recon); enter it next round - Still have open PRs: - All healthy + no maintenance due (per C.1.5) → run parallel Recon (stage = `Track+Recon`); found a direction → enter Phase Contribution next round; otherwise finish the cycle - - Any PR needs maintenance (rebase / feedback / nudge) → do it, update the state file, finish this cycle + - Any PR needs maintenance (rebase / feedback / nudge) → do it, state it in the closing summary, and finish --- @@ -516,44 +504,37 @@ cd {{ source_dir }} && gh issue list -R {{ owner }}/{{ repo }} --limit 15 2>&1 #### D.4 Exit condition -- Reviewed 1-3 PRs/issues this round → update the state file, finish this cycle -- No PRs awaiting review → update the state file (next step = recon), enter Phase Recon next round +- Reviewed 1-3 PRs/issues this round → state what you reviewed in the closing summary, and finish +- No PRs awaiting review → close this round naming the next step (Recon); enter it next round --- -### Recording and Submission - -At the end of every cycle: +### Recording -1. **Update the state file** `{{ evolution_cwd }}/open_source_{{ owner }}_{{ repo }}_state.md` -2. **Record key findings** in `{{ evolution_cwd }}/memory/` (if there are important lessons or insights) - - ⚡ **Memory hygiene** (rant 2026-08-23T08:04:26): keep MEMORY.md a **pure index** — one short line per entry, never duplicated content; update entries in place; if the index has grown long, merge/consolidate instead of appending. -3. **The state file itself does not need git commits** (it's a local work record, lives in EMRG's evolution directory) - ---- +End every cycle with a **closing summary in your final message**. It is the only thing the next round inherits — write it for that reader, not for this one: -### Per-Round Reflection +1. **The phase this round entered**, and whether it completed +2. **What was actually done** — issues scanned, code written, PRs reviewed, discussions replied to +3. **Every open PR of ours**, with its state (MERGEABLE? CI green? feedback pending?) +4. **What is blocked, and on whom** +5. **The next step** — the phase the next round should enter, and why -**Every cycle must end with a reflection appended to `{{ evolution_cwd }}/open_source_{{ owner }}_{{ repo }}_reflections.md` (same directory as the state file). This cannot be skipped.** Create the file if it doesn't exist. - -Reflection is an engagement diary — the operational layer is handed off by the state file (`open_source_*_state.md`: last done / next step / blockers / active PRs), while reflection is the strategic-layer cognition; the two complement each other without duplication. Output format: append each reflection at the end of the file, starting with a datetime header and phase tag; do not modify or delete existing content. +Also record **key findings** (lessons worth keeping beyond this session) as memory entries under `{{ evolution_cwd }}/.emrg/memory/` — the durable layer, whose index is part of this prompt: + - ⚡ **Memory hygiene** (rant 2026-08-23T08:04:26): keep MEMORY.md a **pure index** — one short line per entry, never duplicated content; update entries in place; if the index has grown long, merge/consolidate instead of appending. -Each round must answer these 7 questions (cannot be omitted): +The summary is a message, not a file — nothing to commit, nothing to keep in sync; the session history is the record. -1. **What was this round's goal?** — Which Phase did this round enter (recon/contribution/tracking/review)? What specific task to complete? If there's rant feedback, list the rants considered this round (write "no new rant feedback" if none) -2. **What does success look like?** — What would "done" look like? (PR merged? Issue claimed? Review completed? Contribution accepted?) -3. **What was actually done?** — Concrete actions: which issues scanned, what code written, which PRs reviewed, what discussions replied to, what waited on. **If this round submitted a PR for a rant, record the PR number and the rant it addresses (timestamp/keywords).** -4. **What is the current progress?** — Compared to the ideal outcome, how far along? What's missing? (How many more reviews does the PR need? Which part of the code is unfinished? Was the issue claimed by someone else?) -5. **What pitfalls were hit?** — Which attempts failed, what CI broke, why reviews were rejected, network/permission blockers, platform CLI or browser unavailability. Record honestly, don't gloss over -6. **What opportunities were discovered?** — Which issues are worth doing, which PRs have potential, what new directions in community activity, which project conventions deserve attention? -7. **What is the next direction?** — Based on the reflection, what's the focus next round? Continue the current Phase or switch? (e.g. PR waiting for review → switch to recon for new opportunities; contribution blocked → back to recon) +**Seven questions the closing summary must answer** (they are the self-review that used to live in a separate file): -**Rules**: +1. **What was this round's goal?** — which phase, which specific task; the rants considered this round (or "no new rant feedback") +2. **What does success look like?** — PR merged? issue claimed? review completed? contribution accepted? +3. **What was actually done?** — concrete actions: issues scanned, code written, PRs reviewed, discussions replied to, what waited. **If a PR was submitted for a rant, record the PR number and the rant (timestamp/keywords).** +4. **What is the current progress?** — how far from the ideal outcome, what is missing (how many more reviews needed? which code unfinished? was the issue claimed by someone else?) +5. **What pitfalls were hit?** — failed attempts, what CI broke, why a review was rejected, network/permission blockers, platform CLI or browser unavailability. Honestly, not glossed over +6. **What opportunities were discovered?** — issues worth doing, PRs with potential, new directions in community activity, project conventions worth attention +7. **What is the next direction?** — next round's focus: continue this phase or switch (PR waiting for review → switch to recon for new opportunities; contribution blocked → back to recon) -- Every cycle must end with a reflection; cannot be skipped. Even if this round was "nothing to do/NTE/no new findings", record why (all PRs merged, no open issues, no rants) -- Reflections only append to the end of the file; never modify or delete existing content. This is an engagement diary — "what I actually thought at the time" is itself valuable -- Each reflection starts with a datetime header and phase tag, format: `## 2026-07-31 21:30 — Phase Tracking` -- If this round modified code or submitted a PR, questions 3/4 must record the concrete commit/PR numbers (e.g. PR #123) +**Rules**: a round without a closing summary strands the next round — it is not optional, even when the round was "nothing to do / NTE / no new findings" (then say why: all PRs merged, no open issues, no rants). If the round changed code or submitted a PR, questions 3/4 must name the concrete commit/PR numbers. --- @@ -586,9 +567,9 @@ Other platforms (Gitee/Gitea/Gerrit, etc.): prefer the platform's official CLI ( - GitHub: `https://github.com/{owner}/{repo}/pulls`、`/issues`、`/pulls/{n}` - GitLab: `https://gitlab.com/{owner}/{repo}/-/merge_requests`、`/-/issues`、`/-/merge_requests/{n}` - Use browser harness to complete list / view / review / merge operations - - Browser also unavailable → record "platform CLI and browser both unavailable" in the state file, finish this cycle + - Browser also unavailable → record "platform CLI and browser both unavailable" in the closing summary, finish this cycle -4. **Behavioral consistency**: whether using CLI or browser, the completed operations must be equivalent — the same ROLE LOCK constraints (Contributor does not review/merge/close), the same output recorded in the state file. +4. **Behavioral consistency**: whether using CLI or browser, the completed operations must be equivalent — the same ROLE LOCK constraints (Contributor does not review/merge/close), the same output stated in the closing summary. --- @@ -605,8 +586,8 @@ Other platforms (Gitee/Gitea/Gerrit, etc.): prefer the platform's official CLI ( | Situation | Handling | |-----------|----------| -| Network timeout / `gh` API unavailable | Record in state file (blocked = network unavailable), finish this cycle. **Do not retry.** | -| `git pull` conflicts | `git rebase --abort` (the tree was clean before the pull; abort restores it) → record in state file, finish. **Never stash host work.** | +| Network timeout / `gh` API unavailable | Record the blocker in the closing summary, finish this cycle. **Do not retry.** | +| `git pull` conflicts | `git rebase --abort` (the tree was clean before the pull; abort restores it) → record the conflict in the closing summary, finish. **Never stash host work.** | | `gh pr create` fails (branch name already exists) | Change the branch name, re-push and re-create | | Tests failing | Fix → re-test, don't skip. If unfixable, honestly state it in the PR description | diff --git a/tests/test_competition_prompt.py b/tests/test_competition_prompt.py index c3bdb994..7e16cdd6 100644 --- a/tests/test_competition_prompt.py +++ b/tests/test_competition_prompt.py @@ -161,11 +161,15 @@ def test_goal_line_covers_both_prize_and_standing(): assert "leaderboard standing / percentile" in text -def test_preparation_and_reflection_are_mandatory(): +def test_preparation_and_closing_summary_are_mandatory(): text = PROMPT.read_text(encoding="utf-8") assert "0. Preparation (MUST run first every round)" in text assert "Do not skip the preparation step (even when \"everything looks fine\")" in text - assert "MUST end with a reflection appended" in text + # The per-round reflection used to be appended to a `*_reflections.md` diary; + # the rant 2026-09-14T14:35:47 sweep moved it into the round's final message. + assert "closing summary in your final message" in text + assert "a round that ends without a summary strands the next one" in text + assert "Seven questions the closing summary must answer" in text # Rant scan must match the project field exactly, like the other templates. assert "exactly `{{ task.project }}`" in text @@ -505,21 +509,25 @@ def test_machine_rejection_is_not_permanent(): §4 is what turns a mechanical hit into a permanent exclusion, so the split is part of the fix, not cosmetic. - Asserted against the **template block** rather than the whole document: the + Asserted against the **§4 section block** rather than the whole document: the first version of this test checked only that the phrase occurred somewhere, and phase A (§, "does **not** go in that section") mentions it in prose — so renaming the actual section heading away survived the test unchanged. A presence check that can be satisfied by a mention of the thing is the same class of blindness this cycle is fixing, one level up. + + §4 was a fenced ```markdown state-file template until rant 2026-09-14T14:35:47 + removed the state file; the same two rejection entries now live as memory-entry + forms in prose, and the block is parsed by section. """ text = PROMPT.read_text(encoding="utf-8") - block = text.split("```markdown", 1)[1].split("```", 1)[0] - assert "## Rejected (never re-evaluated)" in block, ( - "the state-file template lost its permanent-rejection section" + block = text.split("### 4. Cross-round state lives", 1)[1].split("\n---", 1)[0] + assert "**Rejected (never re-evaluated)**" in block, ( + "§4 lost its permanent-rejection entry form" ) - assert "## Rejected — needs a human read" in block, ( - "the state-file template has no re-checkable rejection section, so a " - "§3.2.1 negation override would be frozen as permanent" + assert "**Rejected — needs a human read" in block, ( + "§4 has no re-checkable rejection entry form, so a §3.2.1 negation " + "override would be frozen as permanent" ) # Phase A must route to the right one of the two. assert "does **not** go in that section" in text diff --git a/tests/test_prompt_templates.py b/tests/test_prompt_templates.py index 11ac7452..ad949b64 100644 --- a/tests/test_prompt_templates.py +++ b/tests/test_prompt_templates.py @@ -39,6 +39,7 @@ from emrg.protocol import InstanceIdentity from emrg.server import scheduler as mod from emrg.server.scheduler import TaskHandler +from emrg.tools import bash_tool REPO_ROOT = Path(__file__).resolve().parents[1] PROMPTS_DIR = REPO_ROOT / "emrg" / "server" @@ -182,3 +183,139 @@ def from_string(self, source, *args, **kwargs): f"provide ({exc}); with the daemon's Undefined it would silently " f"render as an empty string" ) from exc + + +# `{{ evolution_cwd }}` followed by the rest of the path it names. +EVOLUTION_CWD_REF = re.compile(r"\{\{\s*evolution_cwd\s*\}\}([^\s`)\"'|,;]*)") + +# A path suffix that routes through a `memory/` directory. +_MEMORY_SEGMENT = re.compile(r"(?:^|/)memory/") + + +def _evolution_memory_refs(text: str) -> list[str]: + """Path suffixes of ``{{ evolution_cwd }}`` references that name a memory dir.""" + return [ + suffix + for suffix in EVOLUTION_CWD_REF.findall(text) + if _MEMORY_SEGMENT.search(suffix) + ] + + +def test_prompt_memory_writes_land_where_the_sandbox_allows_them() -> None: + """A prompt's memory path must be one the ``workspace-write`` sandbox trusts. + + Measured 2026-09-14 (rant 2026-09-14T14:35:47): `open_source_prompt.md` sent the + agent's identity file and its "key findings" to `{{ evolution_cwd }}/memory/`, + i.e. `~/.emrg/evolution/memory/`. That is out of the sandbox's boundary — every + such write is refused with "blocked write outside workspace", which the rant + reports an open-source task hitting 32 times in one day — and it is the wrong + root anyway: the evolution data root is `{{ evolution_cwd }}/.emrg/`, the only + part of `{{ evolution_cwd }}` `bash_tool._trusted_write_zones()` trusts and the + only memory the daemon loads. The directory exists, so a write that gets through + by a route the command-line scan cannot see (an `open()` inside a heredoc, as the + rant documents) lands in a folder no memory loader reads. + + The expected location is taken from the sandbox's own trust list rather than + copied here, so if that list moves, this test reports the prompts may be stale + instead of agreeing with a second copy of the rule. + + Named limit: only ``{{ evolution_cwd }}``-rooted *memory* paths are checked. + The state-file / reflection-file mechanism the same rant retires is covered by + ``test_retired_state_file_mechanism_is_gone_or_being_swept`` below, and the + paths that replaced it live in the memory root this test measures. + """ + root = Path(mod.EVOLUTION_CWD) + zones = bash_tool._trusted_write_zones() + assert zones, "no trusted write zone — this check would pass vacuously" + + # Self-test of the device: it must reject the exact shape the rant measured, + # and the machine must actually consider that shape outside the boundary. + planted = _evolution_memory_refs("write `{{ evolution_cwd }}/memory/identity.md`") + assert planted == ["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/memory/identity.md"], planted + assert not any( + bash_tool._is_within(str(root / "memory/identity.md"), zone) for zone in zones + ), ( + "`~/.emrg/evolution/memory/` is inside a trusted zone on this machine, so " + "this test cannot discriminate between the two roots here" + ) + + checked = 0 + for task_type, filename in _builtin_templates(): + text = (PROMPTS_DIR / filename).read_text(encoding="utf-8") + for suffix in _evolution_memory_refs(text): + target = str(root / suffix.lstrip("/")) + assert any( + bash_tool._is_within(target, zone) for zone in zones + ), ( + f"{task_type}/{filename}: tells the agent to write {target!r}, " + f"which the workspace-write sandbox blocks (trusted zones: {zones}); " + f"the evolution memory root is `{{{{ evolution_cwd }}}}/.emrg/memory/`" + ) + checked += 1 + assert checked >= 2, ( + f"only {checked} memory path(s) found across the templates — the scan is " + f"not looking at what it thinks it is" + ) + + +# Templates still teaching the retired state-file / reflection-file mechanism. +# Rant 2026-09-14T14:35:47 removes it wholesale ("the session itself is the +# memory"); the sweep lands one template at a time. Each entry is removed from +# this set in the SAME change that sweeps its template, so the set only shrinks +# and reaching empty is what "no residue" means. +PENDING_STATE_SWEEP = { + "promote_prompt.md", + "paper_prompt.md", + "journal_prompt.md", +} + +# The retired mechanism's fingerprints: the two file names, and the prose that +# told the agent to read/write "the state file". +_RETIRED_MECHANISM = re.compile( + r"_state\.md|_reflections\.md|the state file|state-file", + re.IGNORECASE, +) + + +def test_retired_state_file_mechanism_is_gone_or_being_swept() -> None: + """A prompt outside `PENDING_STATE_SWEEP` must not teach the retired mechanism. + + Measured 2026-09-14 (rant 2026-09-14T14:35:47): each task prompt carried a + per-project `*_state.md` file and an append-only `*_reflections.md` diary, and + every phase hand-off was a line written into one of them. The host retired both + — the daemon already replays the task's session history into every round, so the + session is the state and the memory index is the durable layer. Nothing reads + either file any more, which makes every surviving instruction a write into a + place no loader looks. + + The check runs in both directions, because a set that only ever shrinks is one + someone can forget to shrink: a swept template must be clean, and a template + still listed as pending must actually still mention the mechanism (otherwise it + was swept without being removed from the set here). + + Named limit: this is a *text* guard over the built-in templates only. It cannot + see a state file an agent invents at runtime, and it does not read the rendered + prompt — the templates are checked as written. + """ + swept = [name for _, name in _builtin_templates() if name not in PENDING_STATE_SWEEP] + assert "open_source_prompt.md" in swept, ( + "the swept set lost its pilot template — the check is not looking where it " + "thinks it is" + ) + + for name in swept: + text = (PROMPTS_DIR / name).read_text(encoding="utf-8") + hits = sorted(set(_RETIRED_MECHANISM.findall(text))) + assert not hits, ( + f"{name}: still teaches the retired state-file / reflection-file " + f"mechanism {hits} — the session is the state now, so this instruction " + f"sends the agent to a file nothing reads (rant 2026-09-14T14:35:47)" + ) + + for name in sorted(PENDING_STATE_SWEEP): + text = (PROMPTS_DIR / name).read_text(encoding="utf-8") + assert _RETIRED_MECHANISM.search(text), ( + f"{name} is still listed in PENDING_STATE_SWEEP but no longer mentions " + f"the mechanism — drop it from the set in the same change that swept it, " + f"so the set keeps meaning 'not yet done'" + ) From 56cac2332e07dafcc0bbe8136cce42c4fc569647 Mon Sep 17 00:00:00 2001 From: EMRG Evolution Date: Mon, 14 Sep 2026 16:41:05 +0800 Subject: [PATCH 2/4] emrg: retire the state-file mechanism from the paper template MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Stage 2 of rant 2026-09-14T14:35:47 ("the session itself is the memory"). Stage 1 (#1226) swept open_source_prompt.md and competition_prompt.md; this sweeps paper_prompt.md and drops it from PENDING_STATE_SWEEP in the same change. paper_prompt.md taught a per-task `paper_state.md` (current phase / last completed / next step / blocked / unhandled rants) plus an append-only `reflections.md` diary, and every hand-off was a line written into one of them. Nothing reads either file — the daemon replays this task's session history into every round (daemon.py, under a fixed session_id), so the session is the state, and the memory index is the durable layer. The phase logic is not lost, it is re-based: the phase assessment already reads the project files (literature/heilmeier-catechism.md, result logs, draft .tex, the experiment log), which is where the research state actually lives; the round's closing summary carries "last completed / next step / blocked"; and the rant log (~/.emrg/rants.jsonl) replaces the hand-copied "unhandled rants" field. §6 "Reflect" becomes "Reflect and Record": the same seven questions, answered in the closing summary, with durable lessons going to memory entries instead of a diary file. Verified: no mechanism reference left in the template; paper_prompt.md is dropped from PENDING_STATE_SWEEP in this same commit (the set keeps meaning "not yet done"), and test_paper_template_renders_with_context now pins the new contract instead of asserting the retired path renders. Both the sweep guard and that render test fail when the mechanism is re-introduced into the swept template (mutant restored byte-for-byte). --- emrg/server/paper_prompt.md | 53 ++++++++++++++-------------------- tests/test_prompt_templates.py | 1 - tests/test_scheduler.py | 7 ++++- 3 files changed, 28 insertions(+), 33 deletions(-) diff --git a/emrg/server/paper_prompt.md b/emrg/server/paper_prompt.md index b81ee448..593849c6 100644 --- a/emrg/server/paper_prompt.md +++ b/emrg/server/paper_prompt.md @@ -7,7 +7,7 @@ You are EMRG's paper writing module. **Every writing session MUST fully execute - Uptime: {{ uptime }} - Project source: `{{ source_dir }}` - Session ID: `{{ session_id }}` -- ⚠️ Note: the cycle counter resets to 1 after a daemon restart — **it does NOT represent the true historical run count**. Determine "is this the first run" from the state file `paper_state.md` and the project files instead. +- ⚠️ Note: the cycle counter resets to 1 after a daemon restart — **it does NOT represent the true historical run count**. Determine "is this the first run" from the project files (the phase assessment below reads them anyway) and from this session's own earlier messages, not from the counter. {% if task.extra_prompt %} ## Task-specific Instructions (extra_prompt from tasks.yml) @@ -66,20 +66,15 @@ This rule is a hard constraint against the LLM's "over-writing" tendency — rep **Phase 4 entry guard**: Before entering Phase 4, you MUST check that all experimental results are complete. If data collection is unfinished, stay in Phase 3. -**📋 State file (cross-cycle memory)**: +**📋 Cross-round continuity (there is no state file)**: -At the start of every cycle, you MUST read `{{ source_dir }}/.emrg/sessions/{{ session_id }}/paper_state.md` (create it if it doesn't exist). The state file records: +**This task keeps no state file — the session itself is the state.** The daemon replays this task's session history into every round, so your own earlier messages here, plus the memory index embedded in this prompt, ARE "where the last round left off". Before assessing the phase, reconstruct from them (and from the project files the phase assessment reads anyway): -```markdown -# Paper State -- current phase: Phase 2 | 3 | 4 -- last completed: -- next step: -- blocked: -- unhandled rants: -``` +- the phase the last round entered, and the **next step** its closing summary named +- what was completed, and what is blocked, and on whom +- the raw research state — it lives in real files, not in a notebook: `literature/heilmeier-catechism.md`, the result logs (`*.csv` / `*.json` / figures), the draft `.tex` chapters, the experiment log -At the end of every cycle, update `{{ source_dir }}/.emrg/sessions/{{ session_id }}/paper_state.md`. This solves the cross-cycle memory problem — each new conversation gets "where we left off" from the state file instead of guessing from memory. +If the history is silent or ambiguous, re-check reality (the project files, the experiment logs, `git log`) rather than assume. **A round that ends without a closing summary strands the next round**; that is why the Recording section at the end of this prompt is not optional. --- @@ -101,33 +96,29 @@ When in Phase 2 (Validation) or Phase 3 (Experimentation), the following eleven 10. **Post-Run Review (evaluating experiment results)**: After a batch of experiments, you MUST stop and answer four questions — do results match expectations? Any anomalies? Is completeness sufficient? Do results agree with the literature (use browser harness to check arXiv and compare baseline numbers)? Write the conclusions into the experiment log: `what we saw → what the literature says → what it means → what to do next`. 11. **Negative result handling**: When results don't match expectations, act in order — debug → diagnose the cause → consult the literature → attempt fixes → record honestly. Skipping straight to writing is forbidden. Do not fabricate or selectively report. -Phase 2/3 loop logic: **read state file → determine current step → execute ONE thing → Pre-Flight/Post-Run Review → update state file → git commit & push → finish**. Do one thing at a time; don't aim for a complete loop. If in Phase 2 and the experiment code has placeholders, this round only fixes the placeholders. +Phase 2/3 loop logic: **reconstruct where the last round left off → determine current step → execute ONE thing → Pre-Flight/Post-Run Review → git commit & push → finish with a closing summary**. Do one thing at a time; don't aim for a complete loop. If in Phase 2 and the experiment code has placeholders, this round only fixes the placeholders. --- ### 1. Review -**Review Rants** (execute before reading the state file): +**Review Rants** (MUST run first): Every cycle you MUST first read user feedback from `~/.emrg/rants.jsonl`. Rants are direction-adjustment signals, not one-off tasks. Handling rules: 1. For each pending rant, assess its relevance to the current phase -2. Write the relevant rants' summaries into the "unprocessed rants" field of paper_state.md +2. Carry the relevant rants' direction into this round's plan, and name them in the closing summary (Recording) — the rant log is the record, so there is nothing to copy into a notebook 3. **After reading rants, do not skip the review step** — rants provide directional input, but the specific experiment/literature/draft state needs to be gathered via the review step 4. Paper rants differ from evolution rants: they lean toward direction guidance rather than bug fixing. Translate rant points into concrete writing/experiment decisions, rather than "marking as done" -**Read the state file** (MUST run first): - -```bash -cat {{ source_dir }}/.emrg/sessions/{{ session_id }}/paper_state.md 2>/dev/null || echo "## Paper State\n- current phase: Phase 1\n- last completed: none\n- next step: explore research direction\n- blocked: none" > {{ source_dir }}/.emrg/sessions/{{ session_id }}/paper_state.md -``` +**Reconstruct where the last round left off** (MUST run first): -Perform different review operations based on the current phase: +The state carrier is this session (§0 Cross-round continuity): read your own earlier messages in this task's history, and check the project files the phase assessment reads. Then perform different review operations based on the current phase: **Phase 1** — review existing literature notes and Heilmeier Catechism progress. -**Phase 2/3** — review the last experiment's result logs and the state file; check experiment progress; do NOT touch the paper draft. +**Phase 2/3** — review the last experiment's result logs; check experiment progress; do NOT touch the paper draft. **Phase 4** — review the paper draft and figure data; confirm writing progress. **Get the current date** (MUST run first): @@ -194,7 +185,7 @@ fi ### 5. Submit -- Update `{{ source_dir }}/.emrg/sessions/{{ session_id }}/paper_state.md` (record the current phase, this round's completed operations, next-step plan) +- End the round with a **closing summary in your final message** (§6 Recording): the phase this round entered, what was done, what is blocked, and the next step. The session is the state now, so this summary is the only thing the next round inherits. **Rant marking**: if this round's work path has covered a pending rant's feedback (e.g. the rant suggested lowering the learning rate and this round's experiments adopted it), mark that rant as acknowledged: @@ -226,7 +217,6 @@ with open(rants_file, "w") as f: - Use `json.dumps(..., ensure_ascii=False)`; Chinese escaping is forbidden - Only mark a rant acknowledged when its suggestion has genuinely been incorporated into the work path — merely "reading" it doesn't count - If this round cannot cover it (e.g. the rant suggests Phase 4 writing changes but you're in Phase 2), don't mark it; leave it for later rounds -- After marking, update the "unprocessed rants" list in paper_state.md and remove the processed timestamp - `git add -A && git commit -m "paper: " && git push` - At least one commit per round, pushed immediately @@ -234,16 +224,16 @@ with open(rants_file, "w") as f: --- -### 6. Reflect +### 6. Reflect and Record -**Every cycle MUST end with a reflection appended to `{{ source_dir }}/.emrg/sessions/{{ session_id }}/reflections.md`. This cannot be skipped.** +**Every cycle MUST end with a closing summary in your final message. This cannot be skipped** — there is no diary file and no state file any more: the session history is the record, and the closing summary is what the next round reads out of it. -Reflection is a research diary — the operational layer is tracked by `paper_state.md` (in the session directory), while reflection is strategic-layer cognition. Output format: append each reflection at the end of the file, starting with a datetime header and phase tag; do not modify or delete existing content. +Reflection is strategic-layer cognition, and the closing summary is where it goes — written for the next round's reader, not for this one. Durable lessons (a research direction that did not pan out, an experimental condition that changed a result, a convention worth keeping) belong in **memory entries under `{{ evolution_cwd }}/.emrg/memory/`**, the durable layer whose index is part of this prompt. Memory hygiene: keep the index a pure index — one short line per entry, update in place, consolidate instead of appending. -Each round must answer these 7 questions (cannot be omitted): +Each round the closing summary must answer these 7 questions (cannot be omitted): 1. **What was this round's requirement?** — The original driver: what problem does the host want to solve? Return to the research goal defined by the nine Heilmeier Catechism questions; don't deviate. - **Feedback from rants**: list the rant feedback summaries considered this round (if any). If there are no pending rants this round, write "no new rant feedback". + **Feedback from rants**: list the rant feedback considered this round (if any). If there are no pending rants this round, write "no new rant feedback". 2. **What is the ideal state?** — If everything goes according to plan, what does this round's "perfect ending" look like? 3. **What was actually done?** — Concrete operations: which literature was read, which experiments run, what content written, what waited on 4. **What is the current progress?** — Compared to the ideal state, how far did we actually get? What hasn't been obtained? Where is the gap? @@ -254,8 +244,9 @@ Each round must answer these 7 questions (cannot be omitted): **Rules**: - Even if this round was only "waiting for experiment results", reflect: what you're waiting for, why, and what you did or could do while waiting -- Reflections only append to the end of the file; never modify or delete existing content. This is a research diary — "what I actually thought at the time" is itself valuable -- In Phase 2/3 experimental reflections, pay special attention to the gap between "hypothesis vs results"; record experimental conditions, anomalies, and unexpected findings +- Name the concrete artifacts the next round needs — result file paths, commit hashes, figure names. "What I actually thought at the time" is itself valuable, and the summary is where it survives +- In Phase 2/3, pay special attention to the gap between "hypothesis vs results": record experimental conditions, anomalies, and unexpected findings +- A round without a closing summary strands the next round — it is not optional --- diff --git a/tests/test_prompt_templates.py b/tests/test_prompt_templates.py index ad949b64..8a8a7759 100644 --- a/tests/test_prompt_templates.py +++ b/tests/test_prompt_templates.py @@ -265,7 +265,6 @@ def test_prompt_memory_writes_land_where_the_sandbox_allows_them() -> None: # and reaching empty is what "no residue" means. PENDING_STATE_SWEEP = { "promote_prompt.md", - "paper_prompt.md", "journal_prompt.md", } diff --git a/tests/test_scheduler.py b/tests/test_scheduler.py index b694f88a..dbec4964 100644 --- a/tests/test_scheduler.py +++ b/tests/test_scheduler.py @@ -1005,7 +1005,12 @@ def test_paper_template_renders_with_context(): source_dir="/tmp/paper", session_id="s1", timestamp="20260806", task={}, project={}, evolution_count=0, ) - assert "paper_state.md" in out, "状态文件指引应渲染" + # The per-round `paper_state.md` was retired (rant 2026-09-14T14:35:47): the + # session is the state now, so the template must render the continuity contract + # and the closing summary that replaces the file — and must not name the file. + assert "Cross-round continuity" in out, "跨轮续接指引应渲染" + assert "closing summary in your final message" in out, "收尾总结契约应渲染" + assert "paper_state.md" not in out, "已废置的状态文件指引不应再渲染" assert "latexmk" in out, "LaTeX 检查指引应渲染" assert "literature" in out, "文献去重指引应渲染" From e6c490f0b64849969dd7a915d35ea33dae4f738a Mon Sep 17 00:00:00 2001 From: EMRG Evolution Date: Mon, 14 Sep 2026 17:12:16 +0800 Subject: [PATCH 3/4] emrg: catch the bare "state file" instruction the sweep's fingerprint missed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sweeping a template is only half the work: the guard that says "no residue anywhere else" has to be able to see the residue. Measured this cycle, paper_prompt.md — already swept on this branch — still told the agent to read Agent.md / abstract / state file to determine direction terms to derive its arXiv keywords. A file nothing reads, named as a source of the search terms, for a phase that produces the round's literature scan. It survived because the fingerprint spelled the prose form as "the state file": an instruction that writes the noun phrase without the article was invisible, and a guard that reports a clean sweep by not looking is worse than no guard. Off by one word, in other words, and the same class as the rest of this branch: the mechanism outlived the change that retired it because nobody re-read what the check actually looked for. * paper_prompt.md: drop the state file from the read list (the direction terms come from Agent.md / the abstract, which the phase reads anyway). * test_prompt_templates.py: match the bare noun phrase, not only the article form, and cover the reflection-file paraphrase too — while keeping the sweep's own denial sentences legal ("there is no state file — the session itself is the state" is the replacement text, and flagging it would forbid saying what replaced the mechanism). A new test pins both halves of the pattern: the instruction form that must be flagged, and the denial that must not be. Verified: the widened pattern flags paper_prompt.md and nothing else among the swept templates, and still flags the two templates left in PENDING_STATE_SWEEP, so the set keeps meaning "not yet done". Full suite on this branch 1965 passed / 3 skipped. Both new checks were killed by a mutant and the files restored byte-for-byte (sha256): restoring the old fingerprint fails the coverage test, and re-adding the instruction to paper_prompt.md fails the sweep guard. --- emrg/server/paper_prompt.md | 2 +- tests/test_prompt_templates.py | 35 ++++++++++++++++++++++++++++++++-- 2 files changed, 34 insertions(+), 3 deletions(-) diff --git a/emrg/server/paper_prompt.md b/emrg/server/paper_prompt.md index 593849c6..42cc84e4 100644 --- a/emrg/server/paper_prompt.md +++ b/emrg/server/paper_prompt.md @@ -139,7 +139,7 @@ Check paper-related files under the project directory `{{ source_dir }}`: ls {{ source_dir }}/literature/ 2>/dev/null || echo "[no literature/ directory — literature work has not started]" ``` 2. **Prefer the browser harness skill** to access arXiv (cs.LG, cs.CL, cs.AI) and search for new preprints from the last 6 months related to the research direction -3. **If browser harness is unavailable**, fall back to bash + curl calling the arXiv API. Keywords MUST derive from the project's research direction (read Agent.md / abstract / state file to determine direction terms, e.g. mutual learning, co-teaching, self-play, knowledge distillation); using generic broad terms is forbidden: +3. **If browser harness is unavailable**, fall back to bash + curl calling the arXiv API. Keywords MUST derive from the project's research direction (read Agent.md / abstract to determine direction terms, e.g. mutual learning, co-teaching, self-play, knowledge distillation); using generic broad terms is forbidden: ```bash # Example: search papers from the last 6 months related to the research direction (replace xxx with the direction term) curl -s "http://export.arxiv.org/api/query?search_query=cat:cs.CL+AND+all:xxx&sortBy=submittedDate&sortOrder=descending&max_results=10" diff --git a/tests/test_prompt_templates.py b/tests/test_prompt_templates.py index 8a8a7759..6a8c94f3 100644 --- a/tests/test_prompt_templates.py +++ b/tests/test_prompt_templates.py @@ -269,9 +269,14 @@ def test_prompt_memory_writes_land_where_the_sandbox_allows_them() -> None: } # The retired mechanism's fingerprints: the two file names, and the prose that -# told the agent to read/write "the state file". +# told the agent to read/write a state or reflection file. The prose arm matches +# the bare noun phrase, not only "the state file": the first version demanded the +# article, and paper_prompt.md came through the sweep still telling the agent to +# read "state file" for its arXiv keywords (measured 2026-09-14, +# cyc20260914-170405). A line that *denies* the file — "there is no state file, +# the session is the state" — is the replacement text itself, so it stays legal. _RETIRED_MECHANISM = re.compile( - r"_state\.md|_reflections\.md|the state file|state-file", + r"_state\.md|_reflections\.md|(? None: f"the mechanism — drop it from the set in the same change that swept it, " f"so the set keeps meaning 'not yet done'" ) + + +def test_retired_mechanism_fingerprint_covers_the_bare_noun_phrase() -> None: + """What counts as a fingerprint: the sweep's vocabulary, not just its file names. + + Measured 2026-09-14 (cyc20260914-170405): `paper_prompt.md` came out of the + sweep still instructing the agent to read a "state file" to derive its arXiv + keywords. The fingerprint then matched only "the state file", so the guard + called the template clean while an instruction pointing at a file nothing + reads was still in it. A guard whose blind spot is a plausible spelling of the + thing it forbids is a guard that reports success by not looking, which is + worse than no guard. + + The four strings pin both halves of the pattern: the instruction forms that + must be flagged, and the sweep's own denial sentences, which must not be — + they are the text that replaced the mechanism. + """ + assert _RETIRED_MECHANISM.search( + "read Agent.md / abstract / state file to determine direction terms" + ), "the bare 'state file' instruction is exactly what survived the first sweep" + assert _RETIRED_MECHANISM.search("the state file holds the current phase") + assert _RETIRED_MECHANISM.search("append this to the reflections file") + assert not _RETIRED_MECHANISM.search( + "This task keeps no state file — the session itself is the state." + ), "the replacement text denies the file; flagging it would forbid saying what replaced it" + assert not _RETIRED_MECHANISM.search("there is no reflections file any more") From e044922dee8db5548c035f3e31fa90d22c5178db Mon Sep 17 00:00:00 2001 From: EMRG Evolution Date: Mon, 14 Sep 2026 18:01:19 +0800 Subject: [PATCH 4/4] emrg: the memory write root's index is not the index in the prompt MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three paragraphs the sweep re-based onto memory entries under `{{ evolution_cwd }}/.emrg/memory/` also called that directory's index "part of this prompt". It is not. `_collect_memory_data` embeds `session.cwd/.emrg/memory/MEMORY.md` — the *task project's* index — and the session index, and never `{{ evolution_cwd }}/.emrg/memory/MEMORY.md`. Measured by pointing `EVOLUTION_CWD` at a directory whose memory index carries a marker and rendering the system prompt through the real builder: the marker stays out while the project's appears. On this installation the two roots differ by construction — `{{ evolution_cwd }}` is `~/.emrg/evolution` (988 files, the durable record) while the embedded index belongs to the session's cwd, `{{ evolution_cwd }}/emrg`. The pairing is what was false, not the path: writing memory entries under that root is correct (the workspace-write sandbox trusts it), and the sentences that say "the memory index is embedded in this prompt" without naming a path are correct and stay. Each paragraph now points at the `read` tool instead, so the agent is told where its own entries are rather than where they are not. A guard pins the pairing (`test_no_template_calls_the_write_root_index_its_own_prompt_index`): a paragraph that names the root and claims prompt-embedding fails, with a non-vacuity floor of 3 paragraphs naming the root and a self-test that the detector still detects the shape it was written for. Re-planting the false sentence turns it red; the file is byte-restored afterwards (sha256 equal). Suite on this tree: 1975 collected = 1972 passed + 3 skipped. --- emrg/server/competition_prompt.md | 2 +- emrg/server/open_source_prompt.md | 2 +- emrg/server/paper_prompt.md | 2 +- tests/test_prompt_templates.py | 58 +++++++++++++++++++++++++++++++ 4 files changed, 61 insertions(+), 3 deletions(-) diff --git a/emrg/server/competition_prompt.md b/emrg/server/competition_prompt.md index a24ce53c..3941d399 100644 --- a/emrg/server/competition_prompt.md +++ b/emrg/server/competition_prompt.md @@ -221,7 +221,7 @@ For a staged arrangement such as "preliminary rounds online, final round offline Two things the daemon already gives you every round carry the state: - **The session history** — this task's session id is fixed and its messages are loaded each round, so what previous rounds evaluated, submitted and concluded is in front of you. -- **Memory** (`{{ evolution_cwd }}/.emrg/memory/`) — the durable record, whose index is part of this prompt. +- **Memory** (`{{ evolution_cwd }}/.emrg/memory/`) — the durable record; these entries are files you open yourself, via the `read` tool. (What this prompt embeds is the index of this task's *project* memory, a different directory.) Record one memory entry per competition, in the form a later round (or the host) can re-check: diff --git a/emrg/server/open_source_prompt.md b/emrg/server/open_source_prompt.md index 06dd5a9a..9c9ffd8c 100644 --- a/emrg/server/open_source_prompt.md +++ b/emrg/server/open_source_prompt.md @@ -519,7 +519,7 @@ End every cycle with a **closing summary in your final message**. It is the only 4. **What is blocked, and on whom** 5. **The next step** — the phase the next round should enter, and why -Also record **key findings** (lessons worth keeping beyond this session) as memory entries under `{{ evolution_cwd }}/.emrg/memory/` — the durable layer, whose index is part of this prompt: +Also record **key findings** (lessons worth keeping beyond this session) as memory entries under `{{ evolution_cwd }}/.emrg/memory/` — the durable layer, whose entries you open yourself with the `read` tool: - ⚡ **Memory hygiene** (rant 2026-08-23T08:04:26): keep MEMORY.md a **pure index** — one short line per entry, never duplicated content; update entries in place; if the index has grown long, merge/consolidate instead of appending. The summary is a message, not a file — nothing to commit, nothing to keep in sync; the session history is the record. diff --git a/emrg/server/paper_prompt.md b/emrg/server/paper_prompt.md index 42cc84e4..056cc343 100644 --- a/emrg/server/paper_prompt.md +++ b/emrg/server/paper_prompt.md @@ -228,7 +228,7 @@ with open(rants_file, "w") as f: **Every cycle MUST end with a closing summary in your final message. This cannot be skipped** — there is no diary file and no state file any more: the session history is the record, and the closing summary is what the next round reads out of it. -Reflection is strategic-layer cognition, and the closing summary is where it goes — written for the next round's reader, not for this one. Durable lessons (a research direction that did not pan out, an experimental condition that changed a result, a convention worth keeping) belong in **memory entries under `{{ evolution_cwd }}/.emrg/memory/`**, the durable layer whose index is part of this prompt. Memory hygiene: keep the index a pure index — one short line per entry, update in place, consolidate instead of appending. +Reflection is strategic-layer cognition, and the closing summary is where it goes — written for the next round's reader, not for this one. Durable lessons (a research direction that did not pan out, an experimental condition that changed a result, a convention worth keeping) belong in **memory entries under `{{ evolution_cwd }}/.emrg/memory/`**, the durable layer, whose entries you open yourself with the `read` tool. Memory hygiene: keep the index a pure index — one short line per entry, update in place, consolidate instead of appending. Each round the closing summary must answer these 7 questions (cannot be omitted): diff --git a/tests/test_prompt_templates.py b/tests/test_prompt_templates.py index 6a8c94f3..d962f563 100644 --- a/tests/test_prompt_templates.py +++ b/tests/test_prompt_templates.py @@ -258,6 +258,64 @@ def test_prompt_memory_writes_land_where_the_sandbox_allows_them() -> None: ) +# The memory root the sweep re-based every phase hand-off onto: writable (the +# guard above proves it) and readable via the `read` tool — but its own index is +# NOT the one the daemon embeds. +SWEEP_MEMORY_ROOT = "evolution_cwd }}/.emrg/memory" +_INDEX_IN_PROMPT = re.compile(r"(part of|embedded in) this prompt", re.IGNORECASE) + + +def test_no_template_calls_the_write_root_index_its_own_prompt_index() -> None: + """A root that is only *writable* must not be described as *loaded*. + + Measured 2026-09-14 (cyc20260914-175549), on this branch before the fix: three + of the paragraphs re-based onto memory entries under + `{{ evolution_cwd }}/.emrg/memory/` also called that directory's index "part + of this prompt". It is not. `_collect_memory_data` embeds + `session.cwd/.emrg/memory/MEMORY.md` — the *task project's* index — and the + session index, and never `{{ evolution_cwd }}/.emrg/memory/MEMORY.md`; + measured by pointing `EVOLUTION_CWD` at a directory whose memory index carries + a marker and rendering the system prompt through the real builder: the marker + stays out while the project's appears. The two roots differ by construction on + this installation — `{{ evolution_cwd }}` is `~/.emrg/evolution` (988 files, + the durable record) while the embedded index belongs to the session's cwd, + `{{ evolution_cwd }}/emrg`. + + The pairing is what is false, not the path: writing memory entries under that + root is correct (the sandbox trusts it, guarded above), and saying "the memory + index is embedded in this prompt" without naming a path is correct too. Naming + that path *and* claiming its index is in the prompt points the agent at a + place whose contents it will not find — the same defect family the sweep + exists to remove. + """ + planted = ( + "Record findings under `{{ evolution_cwd }}/.emrg/memory/`, " + "whose index is part of this prompt." + ) + assert SWEEP_MEMORY_ROOT in planted and _INDEX_IN_PROMPT.search(planted), ( + "the detector no longer detects the shape it was written for" + ) + + paragraphs_naming_the_root = 0 + suspects: list[str] = [] + for _task_type, filename in _builtin_templates(): + text = (PROMPTS_DIR / filename).read_text(encoding="utf-8") + for paragraph in text.split("\n\n"): + if SWEEP_MEMORY_ROOT in paragraph: + paragraphs_naming_the_root += 1 + if _INDEX_IN_PROMPT.search(paragraph): + suspects.append(f"{filename}: {paragraph.strip()[:140]}") + + assert paragraphs_naming_the_root >= 3, ( + f"only {paragraphs_naming_the_root} paragraph(s) name the memory root — " + f"the scan is not looking where it thinks it is" + ) + assert not suspects, ( + "these paragraphs name the memory root and also claim its index is in the " + "prompt, which the daemon never embeds:\n " + "\n ".join(suspects) + ) + + # Templates still teaching the retired state-file / reflection-file mechanism. # Rant 2026-09-14T14:35:47 removes it wholesale ("the session itself is the # memory"); the sweep lands one template at a time. Each entry is removed from