Repository navigation
labs: tini as PID 1, kill the whole process group on a timeout (Closes #7090) - #7351
Merged
Merged
Conversation
…#7090) Both Railway labs ran lab.py as PID 1, which never waits on orphans reparented to it, so zombies filled the 1000-pid cgroup (t27c lab 763, t27b lab 866 at 07:35Z) and spawns failed with EAGAIN. - Dockerfiles: install tini and make it the ENTRYPOINT; it reaps every orphan. - lab.py (both): each command runs in a session of its own and a timeout SIGKILLs the whole group, so t27c's spec_tests and zig no longer outlive the t27c that started them. The t27c lab's limit is now a timer, so a gate that hangs without printing is also stopped. - t27c lab: the queue worker survives an exception. ENOSPC on /data, raised while recording a failed run, ended the thread at about 06:41Z and the queue stood still with nothing running. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gHashTag
enabled auto-merge (squash)
October 7, 2026 07:43
Contributor
This was referenced Oct 7, 2026
Owner
Author
|
Measured on both labs after the redeploy (2026-10-07). Zombies were counted with
Negative control. A child orphans a grandchild and exits.
Why the queue stopped. The t27c queue had stood still since about 06:41Z. The cause was not the pid cap. Disk. Persistent |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #7090. Covers both labs (t27c-lab and t27b-lab), and the #7255 comments' suggestions: tini, and killing the process group on a timeout.
Cause
Both Railway labs ran
lab.pyas PID 1. PID 1 never waited on orphans reparented to it, so every timed-outt27c test-reportleft itsspec_tests/zig behind as zombies.spec_testsstill in state R (up to 3 h old).Change
tiniand make it the ENTRYPOINT. tini is PID 1 and reaps every orphan.start_new_session=True, and a timeout SIGKILLs the whole group (os.killpg). Once the command ends, whatever is left of its group is swept too, so t27c's children no longer outlive it.run()and the reference path (run_group, which replacessubprocess.run(timeout=...)).sh(). Its limit used to be checked only when a line of output arrived, so a gate that hung silently was never stopped. It is now a timer.OSError: [Errno 28] No space left on device, raised while recording a failed run on a full /data. It was not pid exhaustion. The worker now catches errors, publishes them aserror, and continues.Checks (local, no build)
run_group,runandsheach leave 0 grandchildren. Master'ssubprocess.run(timeout=)leaves 1.erroris published.t27c lab --self-checkpasses, as dotest_a_t27b_spec_cannot_move_silently.pyandtest_the_t27b_lab_heals_its_clone.py.Zwith PPID 1.Lab measurements after the deploy follow in a comment.
🤖 Generated with Claude Code