Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions docker/OpenShell-RSI.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,12 @@ RUN apt-get update \
&& find /opt/headlesscode/src -type d -name __tests__ -prune -exec rm -rf {} + \
&& find /opt/headlesscode/src -type f \( -name '*.test.ts' -o -name '*.spec.ts' \) -delete \
&& rm -rf /opt/headlesscode/scripts/eval-suite \
&& rm -rf /opt/headlesscode/scripts/rsi-benchmark \
&& rm -rf /opt/headlesscode/fixtures/rsi-benchmark \
&& test ! -e /opt/headlesscode/src/rsi

RUN test ! -e /opt/headlesscode/scripts/rsi-benchmark/task-spec.mjs \
&& test ! -e /opt/headlesscode/fixtures/rsi-benchmark \
&& test ! -e /opt/headlesscode/.git

USER sandbox
4 changes: 4 additions & 0 deletions docker/OpenShell.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,10 @@ COPY src ./src
COPY shared ./shared
COPY scripts/spawn-parallel-worktrees.sh scripts/run-worker.sh scripts/headlesscode-answer.sh ./scripts/
RUN chmod -R a+rX /opt/headlesscode/src /opt/headlesscode/shared /opt/headlesscode/scripts
RUN test ! -e /opt/headlesscode/scripts/rsi-benchmark/task-spec.mjs \
&& test ! -e /opt/headlesscode/scripts/rsi-benchmark/manual-tests \
&& test ! -e /opt/headlesscode/fixtures/rsi-benchmark \
&& test ! -e /opt/headlesscode/.git
RUN printf '%s\n' "$HEADLESSCODE_HARNESS_COMMIT" > /opt/headlesscode/.headlesscode-harness-commit
LABEL org.capsize.headlesscode.harness-commit="$HEADLESSCODE_HARNESS_COMMIT"
RUN chmod +x scripts/spawn-parallel-worktrees.sh scripts/run-worker.sh scripts/headlesscode-answer.sh
Expand Down
74 changes: 64 additions & 10 deletions docs/recursive-self-improvement.md
Original file line number Diff line number Diff line change
Expand Up @@ -246,7 +246,10 @@ OpenRouter is available for RSI fleet mutations and the development benchmark
when the worker role is configured with
`HEADLESSCODE_RSI_ROLE_WORKER_PROVIDER=openrouter` and
`HEADLESSCODE_RSI_ROLE_WORKER_MODEL=deepseek/deepseek-v4-flash-0731`. The model
is pinned to DeepSeek's OpenRouter endpoint with fallback routing disabled.
is routed through OpenRouter's available DeepInfra fp8 endpoint
(`deepinfra/fp8`) with fallback routing disabled. The broker checks endpoint
metadata and requires the returned selected-provider metadata to identify
DeepInfra before qualifying usage.
Include `REMOTE_API` in `HEADLESSCODE_RSI_WORKER_RESOURCE_CLASSES` for workers
that claim brokered calls; Ollama workers continue to claim `LOCAL_GPU`.
Workers need the host-side `HEADLESSCODE_OPENROUTER_API_KEY` and
Expand Down Expand Up @@ -311,15 +314,66 @@ benchmark inference use distinct broker campaign IDs derived from the shared
proof campaign ID, so their allocation caps remain separate while their
reservations count against the one PostgreSQL proof-campaign root ceiling.

The current controller completes only the development phase. It records the
arm as `evaluated` and releases its lease without finishing the proof campaign.
A later confirmation phase can reacquire the same arm and use only the
predeclared cells. Every selected final candidate and each root-control
comparison must complete; each unused final candidate slot must be recorded
with the zero-spend `not-selected` disposition. Only then may `finishArm`
mark that arm terminal. Confirmation execution is not yet wired into the
controller; until that work is added, these campaigns remain development-only
and are not complete proof results.
The demonstration launcher separates engineering from holdout planning. The
engineering run freezes development-only cells and runs both arms for two
generations. It must record an accepted generation-zero self-hosted candidate
and an accepted generation-one child whose worker execution used that exact
parent commit. The fixed-control arm must also reach a terminal generation-one
decision. Only after these checks pass may the operator create the final plan;
that plan binds the engineering evidence and freezes the three confirmation
campaigns. Engineering-only manifests contain no confirmation task IDs or
confirmation schedule cells.

Use the staged commands below from a clean checkout with the exact fixed and
self-hosted harness image identities configured. Estimate and plan fetch only
the pinned endpoint's read-only price metadata. Execution requires PostgreSQL,
OpenShell, the broker, the model key on the supervisor, finite per-role
allocations, and a separately reviewed operator cap. No production inference
has been run as part of this implementation.

```sh
# Review the engineering-only schedule and live price before choosing a cap.
npm run rsi:proof-demonstration -- engineering --repo "$PWD" --seed "$ENGINEERING_SEED"

# Run the two-generation engineering pair. The cap must cover this root's
# frozen reservation and fit the configured mutation + benchmark allocations.
npm run rsi:proof-demonstration -- engineering --repo "$PWD" --seed "$ENGINEERING_SEED" \
--operator-cap-usd "$ENGINEERING_CAP_USD" --execute

# Freeze three new campaign seeds only after the engineering evidence passes.
npm run rsi:proof-demonstration -- plan --repo "$PWD" \
--engineering-evidence .headlesscode/rsi-proof/engineering-evidence.json \
--operator-cap-usd "$CAMPAIGN_CAP_USD" --seed "$ENGINEERING_SEED" \
--campaign-seed "$CAMPAIGN_SEED_A" --campaign-seed "$CAMPAIGN_SEED_B" \
--campaign-seed "$CAMPAIGN_SEED_C" --out .headlesscode/rsi-proof/frozen-plan.json

# Verify the frozen identities and cap, execute all three paired development
# and confirmation campaigns, then write JSON and Markdown analysis reports.
npm run rsi:proof-demonstration -- verify --repo "$PWD" --plan .headlesscode/rsi-proof/frozen-plan.json
npm run rsi:proof-demonstration -- execute --repo "$PWD" --plan .headlesscode/rsi-proof/frozen-plan.json \
--operator-cap-usd "$CAMPAIGN_CAP_USD" --execute
npm run rsi:proof-demonstration -- analyze --repo "$PWD" --plan .headlesscode/rsi-proof/frozen-plan.json \
--archive .headlesscode/rsi
```

The engineering command binds each campaign arm to a deterministic
campaign-and-arm run ID. A restart reuses a finished arm's verified archive
record or resumes the same active identity; an archived incomplete arm is
rejected. Final campaign execution likewise reuses terminal arm records and
will not dispatch a second attempt for an already completed cell. A crash
after broker dispatch but before a verified result artifact remains uncertain
and fails closed under the PostgreSQL ledger; it is not retried automatically.
Each final campaign's cap is checked against its frozen root reservation, and
the command's supplied cap must exactly match the value in the plan. The
engineering cap is separate from the final plan's cap for the three campaigns.

Analysis verifies the three paired terminal campaign records, frozen
manifest/config/model/runtime identities, each confirmation schedule cell,
the selected-parent lineage, broker receipts, artifact digests, and PostgreSQL
root accounting. Missing, unqualified, or altered evidence yields
`inconclusive`; a completed set that misses the frozen improvement gates yields
`not-demonstrated`. A demonstration verdict is not a release or promotion
action.

To exercise the worker image and guest network boundary without a production
provider credential or paid request, build the exact current harness image
Expand Down
1 change: 1 addition & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,7 @@
"cli": "tsx src/cli.ts",
"monitor:pilot": "tsx scripts/monitor-pilot/run.ts",
"rsi:promote-curriculum": "tsx src/rsi/promote-curriculum.ts",
"rsi:proof-demonstration": "tsx scripts/rsi-proof-demonstration.ts",
"test": "node scripts/run-tests.mjs",
"prepublishOnly": "npm ci && npm test"
},
Expand Down
5 changes: 3 additions & 2 deletions scripts/rsi-benchmark/mock-openrouter-guest-smoke.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ import * as os from "node:os"
import * as path from "node:path"
import { createServer } from "node:http"
import { OpenShellBenchmarkAgent, initializeTaskRepository } from "../../src/rsi/benchmark/openshell-agent.js"
import { RSI_OPENROUTER_MODEL, startOpenRouterBroker } from "../../src/rsi/inference-broker.js"
import { RSI_OPENROUTER_ENDPOINT_PROVIDER, RSI_OPENROUTER_ENDPOINT_TAG, RSI_OPENROUTER_MODEL, startOpenRouterBroker } from "../../src/rsi/inference-broker.js"
import { OpenShellSessionProvider } from "../../src/cloud/openshell-provider.js"

const repoRoot = path.resolve(process.argv[2] ?? process.cwd())
Expand Down Expand Up @@ -60,7 +60,7 @@ try {
upstreamFetch: async (input, init) => {
const url = String(input)
if (url.endsWith(`/models/${RSI_OPENROUTER_MODEL}/endpoints`)) {
return new Response(JSON.stringify({ data: { endpoints: [{ provider_slug: "deepseek", pricing: { prompt: "0.000001", completion: "0.000001" } }] } }), { status: 200, headers: { "content-type": "application/json" } })
return new Response(JSON.stringify({ data: { endpoints: [{ provider_name: "DeepInfra", tag: RSI_OPENROUTER_ENDPOINT_TAG, quantization: "fp8", status: 0, supported_parameters: ["tools", "tool_choice", "temperature", "seed", "max_tokens"], supports_tool_choice: { auto: true, required: true, function: true }, pricing: { prompt: "0.000001", completion: "0.000001" } }] } }), { status: 200, headers: { "content-type": "application/json" } })
}
assert.equal(url, "https://openrouter.ai/api/v1/chat/completions")
assert.equal(new Headers(init?.headers).get("authorization"), `Bearer ${key}`)
Expand All @@ -76,6 +76,7 @@ try {
assert.ok(body.messages?.length)
return new Response(JSON.stringify({
id: `mock-response-${upstreamCalls}`, model: RSI_OPENROUTER_MODEL,
openrouter_metadata: { endpoints: { available: [{ provider: RSI_OPENROUTER_ENDPOINT_PROVIDER, selected: true }] } },
choices: [{ message, finish_reason: "tool_calls" }],
usage: { prompt_tokens: 400, completion_tokens: 24, cost: 0.000424 },
}), { status: 200, headers: { "content-type": "application/json" } })
Expand Down
5 changes: 3 additions & 2 deletions scripts/rsi-benchmark/self-hosted-harness-smoke.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import { createHash } from "node:crypto"
import * as fs from "node:fs/promises"
import * as os from "node:os"
import * as path from "node:path"
import { FileOpenRouterInferenceLedger, RSI_OPENROUTER_MODEL, startOpenRouterBroker, type OpenRouterBrokerLimits } from "../../src/rsi/inference-broker.js"
import { FileOpenRouterInferenceLedger, RSI_OPENROUTER_ENDPOINT_PROVIDER, RSI_OPENROUTER_ENDPOINT_TAG, RSI_OPENROUTER_MODEL, startOpenRouterBroker, type OpenRouterBrokerLimits } from "../../src/rsi/inference-broker.js"
import { createMutationSnapshotBundle, runOpenShellMutation } from "../../src/rsi/openshell.js"
import { workerHarnessSentinelDiff } from "../../src/rsi/worker-harness-sentinel.js"
import type { CandidateRecord, RsiConfig, WorkerHarnessMode, WorkerHarnessRuntimeIdentity } from "../../src/rsi/types.js"
Expand Down Expand Up @@ -54,7 +54,7 @@ function mockUpstream(stage: "seed" | "bridge" | "sentinel", observed: { sentine
return async (input: RequestInfo | URL, init?: RequestInit): Promise<Response> => {
const url = String(input)
if (url.endsWith(`/models/${RSI_OPENROUTER_MODEL}/endpoints`)) {
return new Response(JSON.stringify({ data: { endpoints: [{ provider_slug: "deepseek", pricing: { prompt: "0.000001", completion: "0.000001" } }] } }), { status: 200, headers: { "content-type": "application/json" } })
return new Response(JSON.stringify({ data: { endpoints: [{ provider_name: "DeepInfra", tag: RSI_OPENROUTER_ENDPOINT_TAG, quantization: "fp8", status: 0, supported_parameters: ["tools", "tool_choice", "temperature", "seed", "max_tokens"], supports_tool_choice: { auto: true, required: true, function: true }, pricing: { prompt: "0.000001", completion: "0.000001" } }] } }), { status: 200, headers: { "content-type": "application/json" } })
}
assert.equal(url, "https://openrouter.ai/api/v1/chat/completions")
assert.equal(new Headers(init?.headers).get("authorization"), "Bearer mock-only-not-a-provider-key")
Expand All @@ -81,6 +81,7 @@ function mockUpstream(stage: "seed" | "bridge" | "sentinel", observed: { sentine
return new Response(JSON.stringify({
id: `mock-self-hosted-${stage}-${calls}`,
model: RSI_OPENROUTER_MODEL,
openrouter_metadata: { endpoints: { available: [{ provider: RSI_OPENROUTER_ENDPOINT_PROVIDER, selected: true }] } },
choices: [{ message, finish_reason: "tool_calls" }],
usage: { prompt_tokens: 10_000, completion_tokens: 200, cost: 0.0102 },
}), { status: 200, headers: { "content-type": "application/json" } })
Expand Down
Loading
Loading