Skip to content

labs: tini as PID 1, kill the whole process group on a timeout (Closes #7090) - #7351

Merged
gHashTag merged 1 commit into
masterfrom
claude/lab-init-reaper-7090
Oct 7, 2026
Merged

gHashTag merged 1 commit into
masterfrom
claude/lab-init-reaper-7090

Conversation

@gHashTag

@gHashTag gHashTag commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Closes #7090. Covers both labs (t27c-lab and t27b-lab), and the #7255 comments' suggestions: tini, and killing the process group on a timeout.

Cause

Both Railway labs ran lab.py as PID 1. PID 1 never waited on orphans reparented to it, so every timed-out t27c test-report left its spec_tests/zig behind as zombies.

  • t27c-lab, 07:35Z: 763 zombies, pids.current 769 of 1000.
  • t27b-lab, 07:35Z: 866 zombies, pids.current 912 of 1000, plus 37 orphaned spec_tests still in state R (up to 3 h old).

Change

  • Dockerfiles (both): install tini and make it the ENTRYPOINT. tini is PID 1 and reaps every orphan.
  • lab.py (both): each command runs with start_new_session=True, and a timeout SIGKILLs the whole group (os.killpg). Once the command ends, whatever is left of its group is swept too, so t27c's children no longer outlive it.
    • t27b: run() and the reference path (run_group, which replaces subprocess.run(timeout=...)).
    • t27c: sh(). Its limit used to be checked only when a line of output arrived, so a gate that hung silently was never stopped. It is now a timer.
  • t27c lab, the stuck queue: since about 06:41Z the queue held 6 commits and nothing ran. The worker thread had died. Railway's log shows OSError: [Errno 28] No space left on device, raised while recording a failed run on a full /data. It was not pid exhaustion. The worker now catches errors, publishes them as error, and continues.

Checks (local, no build)

  • On a tree that starts a grandchild and then hangs, under a 2 s limit, run_group, run and sh each leave 0 grandchildren. Master's subprocess.run(timeout=) leaves 1.
  • A worker whose first job raises ENOSPC keeps running, and its error is published.
  • t27c lab --self-check passes, as do test_a_t27b_spec_cannot_move_silently.py and test_the_t27b_lab_heals_its_clone.py.
  • Negative control on both labs before the fix: a child that orphans a grandchild and exits. The grandchild stays Z with PPID 1.

Lab measurements after the deploy follow in a comment.

🤖 Generated with Claude Code

…#7090)

Both Railway labs ran lab.py as PID 1, which never waits on orphans
reparented to it, so zombies filled the 1000-pid cgroup (t27c lab 763,
t27b lab 866 at 07:35Z) and spawns failed with EAGAIN.

- Dockerfiles: install tini and make it the ENTRYPOINT; it reaps every
  orphan.
- lab.py (both): each command runs in a session of its own and a
  timeout SIGKILLs the whole group, so t27c's spec_tests and zig no
  longer outlive the t27c that started them. The t27c lab's limit is
  now a timer, so a gate that hangs without printing is also stopped.
- t27c lab: the queue worker survives an exception. ENOSPC on /data,
  raised while recording a failed run, ended the thread at about
  06:41Z and the queue stood still with nothing running.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@gHashTag gHashTag added the owner-approved-foreign Owner-approved exception to the only-t27 rule: hand-written foreign code allowed in this PR label Oct 7, 2026
@gHashTag
gHashTag enabled auto-merge (squash) October 7, 2026 07:43
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-10-07 07:45:20 UTC

Summary

Status Count
Total Open PRs 50
PRs with Failing Checks 47
PRs with All Checks Green 3
READY 2
FAILING 47
PENDING 0
NO CHECKS YET 0

These columns do not partition: 2 + 47 + 0 + 0 = 49, and there are 50 open PRs. A PR is being counted twice or not at all.

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=1aa228450491 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@gHashTag

gHashTag commented Oct 7, 2026

Copy link
Copy Markdown
Owner Author

Measured on both labs after the redeploy (2026-10-07). Zombies were counted with ps -eo stat | grep -c ^Z; pids come from /sys/fs/cgroup/pids.current, which has a cap of 1000.

lab before (07:35Z) after a full run
t27c PID 1 python3 /app/lab.py: 763 zombies, 769 pids PID 1 tini: 0 zombies, 9 pids (08:06Z, after 3 full gate runs: 1b74b1b, 223c28c, e45ba56)
t27b PID 1 python3 -u lab.py: 866 zombies, 912 pids PID 1 tini: 0 zombies, 111 pids (08:14Z, after its full reference and fuzz run on e45ba56; the 111 pids are the threads of a lane's live t27c test-report)

Negative control. A child orphans a grandchild and exits.

  • Before the redeploy, the grandchild stayed on both labs as ppid=1 stat=Z.
  • After the redeploy, it is reaped on both labs (empty stat).

Why the queue stopped. The t27c queue had stood still since about 06:41Z. The cause was not the pid cap. /data was full, an OSError: [Errno 28] was raised while a failed run was being recorded, and that error ended the worker thread. The thread now survives that error and publishes it.

Disk. /data was cleaned by hand from 99% (694M free) to 61% (18G free). #7358 adds an automatic reap of idle target dirs.

Persistent /data and /srv survived both deploys.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

owner-approved-foreign Owner-approved exception to the only-t27 rule: hand-written foreign code allowed in this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

t27c-lab: PID 1 never reaps orphans -- 417 zombies of pids.max 1000

1 participant