Skip to content

fix(layout): grow the filesystem on a thread of its own - #342

Open
ci-robbot wants to merge 1 commit into
mudler:masterfrom
ci-robbot:fix/grow-fs-thread-namespace
Open

ci-robbot wants to merge 1 commit into
mudler:masterfrom
ci-robbot:fix/grow-fs-thread-namespace

Conversation

@ci-robbot

Copy link
Copy Markdown
Contributor

What

ephemeralMount unshares the mount namespace so the temporary mount the grow
needs is not visible to the rest of the system:

_ = unix.Unshare(unix.CLONE_NEWNS)
_ = unix.Mount("", "/", "", unix.MS_REC|unix.MS_PRIVATE, "")

A mount namespace belongs to an OS thread, not to a goroutine, and Go gives no
goroutine a thread of its own. So the unshare lands on whichever thread the
caller happens to be scheduled on, and when GrowFSToMax returns, the runtime
puts that thread back in the pool with its frozen view of the mount tree still
on it. The next goroutine to ask for a thread can get it.

That is a problem for anything that mounts in the same process afterwards. A
mount made from another thread does not exist as far as the stale thread is
concerned, so code running there sees whatever was underneath.

In immucore this is a boot bug. The Grow persistent stage in
00_rootfs.yaml is an expand_partition layout stage, it runs from the
rootfs yip stage in-process, and the mount DAG that follows it in the same
process then mounts COS_PERSISTENT on /sysroot/usr/local and the /etc
overlay. On a fraction of boots the bind step reports, for mounts it made
itself a moment earlier:

* mkdir /sysroot/usr/local/.state: read-only file system
* mkdir /sysroot/etc/cni: read-only file system
* mounting ".state/home.bind" at "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/sysroot/home": no such file or directory

The read-only active image showing through at a path that was just mounted
read-write is exactly the shape of a stale mount namespace, and which entries
are hit varies from boot to boot, which is exactly what thread scheduling gives
you. Reported as kairos-io/kairos#4837, and the matching CI failure is
kairos-io/kairos#4743.

Fix

Run the grow on a goroutine that locks its OS thread and never unlocks it, so
the runtime destroys the thread when the goroutine returns and the namespace
dies with it. Two smaller fixes in the same path:

  • Only MS_REC|MS_PRIVATE the mount tree when the unshare actually succeeded.
    The error was discarded, and in the shared namespace that line turns every
    mount on the machine private and breaks propagation for everything else.
  • Drop the runtime.SetFinalizer. It would run the unmount on an arbitrary
    thread, in a namespace where the mount does not exist. Every caller already
    defers cleanup.

Test

TestRunOnDedicatedThreadRetiresTheThread records the thread fn ran on and
then asks the runtime for 256 locked goroutines. None may land on that thread.
It fails 5 times out of 5 when the helper is changed to defer runtime.UnlockOSThread(), and passes 5 out of 5 with the fix:

--- FAIL: TestRunOnDedicatedThreadRetiresTheThread (0.00s)
    a later goroutine ran on thread 11492, the one fn mutated

TestRunOnDedicatedThreadContainsUnshare does the real check, that the
caller's /proc/thread-self/ns/mnt is unchanged after fn unshares. It needs
CAP_SYS_ADMIN and skips without it. It is skipped, not run, on the machine
this was prepared on: user namespaces are blocked there, so unshare -Urm
cannot give it the capability either.

go test ./pkg/plugins/ gives 84 passed / 31 failed both before and after this
change. The 31 are the pre-existing user and layout specs that need root.


AI was used to write this change. No human has read it yet.

ephemeralMount unshares the mount namespace so the temporary mount it needs
is not visible to the rest of the system. A mount namespace belongs to an OS
thread, and Go gives no goroutine a thread of its own, so the unshare landed
on whatever thread the caller happened to sit on and the runtime handed that
thread to the next goroutine that asked for one.

Any caller that mounts something afterwards, from another thread, is then
invisible to code running on the stale one: the read-only image underneath
shows through instead. In immucore the Grow persistent rootfs stage runs in
the same process as the boot mount DAG, so the bind-mount step intermittently
stopped seeing COS_PERSISTENT and /etc, and the node booted with a varying
subset of the persistent binds missing (kairos-io/kairos#4837, #4743).

Run the grow on a goroutine locked to its thread and never unlocked, so the
runtime destroys the thread, and the namespace, when it returns. Two smaller
fixes in the same path: only turn the mount tree private when the unshare
actually succeeded, since in the shared namespace that breaks propagation for
everything on the machine, and drop the finalizer, which would have run the
unmount on an arbitrary thread in the wrong namespace.

Signed-off-by: Ettore Di Giacinto <mudler@kairos.io>
@ci-robbot

Copy link
Copy Markdown
Contributor Author

For context on priority: this is what reds test-core across kairos PRs right now (kairos-io/kairos#4743 and #4837). v1.26.3 did not fix it, the thread-affinity bug here is the actual cause.

@ci-robbot

Copy link
Copy Markdown
Contributor Author

@mudler review ping: this is still the live fix for kairos-io/kairos#4837 (boots come up with no persistence, /home empty and SSH host keys rotated), and v1.26.4 does not contain it. Green on all four checks since 2026-09-21.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants