Conversation
ephemeralMount unshares the mount namespace so the temporary mount it needs is not visible to the rest of the system. A mount namespace belongs to an OS thread, and Go gives no goroutine a thread of its own, so the unshare landed on whatever thread the caller happened to sit on and the runtime handed that thread to the next goroutine that asked for one. Any caller that mounts something afterwards, from another thread, is then invisible to code running on the stale one: the read-only image underneath shows through instead. In immucore the Grow persistent rootfs stage runs in the same process as the boot mount DAG, so the bind-mount step intermittently stopped seeing COS_PERSISTENT and /etc, and the node booted with a varying subset of the persistent binds missing (kairos-io/kairos#4837, #4743). Run the grow on a goroutine locked to its thread and never unlocked, so the runtime destroys the thread, and the namespace, when it returns. Two smaller fixes in the same path: only turn the mount tree private when the unshare actually succeeded, since in the shared namespace that breaks propagation for everything on the machine, and drop the finalizer, which would have run the unmount on an arbitrary thread in the wrong namespace. Signed-off-by: Ettore Di Giacinto <mudler@kairos.io>
This was referenced Sep 21, 2026
Merged
Contributor
Author
|
For context on priority: this is what reds |
Contributor
Author
|
@mudler review ping: this is still the live fix for kairos-io/kairos#4837 (boots come up with no persistence, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
ephemeralMountunshares the mount namespace so the temporary mount the growneeds is not visible to the rest of the system:
A mount namespace belongs to an OS thread, not to a goroutine, and Go gives no
goroutine a thread of its own. So the unshare lands on whichever thread the
caller happens to be scheduled on, and when
GrowFSToMaxreturns, the runtimeputs that thread back in the pool with its frozen view of the mount tree still
on it. The next goroutine to ask for a thread can get it.
That is a problem for anything that mounts in the same process afterwards. A
mount made from another thread does not exist as far as the stale thread is
concerned, so code running there sees whatever was underneath.
In
immucorethis is a boot bug. TheGrow persistentstage in00_rootfs.yamlis anexpand_partitionlayout stage, it runs from therootfsyip stage in-process, and the mount DAG that follows it in the sameprocess then mounts
COS_PERSISTENTon/sysroot/usr/localand the/etcoverlay. On a fraction of boots the bind step reports, for mounts it made
itself a moment earlier:
The read-only active image showing through at a path that was just mounted
read-write is exactly the shape of a stale mount namespace, and which entries
are hit varies from boot to boot, which is exactly what thread scheduling gives
you. Reported as kairos-io/kairos#4837, and the matching CI failure is
kairos-io/kairos#4743.
Fix
Run the grow on a goroutine that locks its OS thread and never unlocks it, so
the runtime destroys the thread when the goroutine returns and the namespace
dies with it. Two smaller fixes in the same path:
MS_REC|MS_PRIVATEthe mount tree when the unshare actually succeeded.The error was discarded, and in the shared namespace that line turns every
mount on the machine private and breaks propagation for everything else.
runtime.SetFinalizer. It would run the unmount on an arbitrarythread, in a namespace where the mount does not exist. Every caller already
defers
cleanup.Test
TestRunOnDedicatedThreadRetiresTheThreadrecords the threadfnran on andthen asks the runtime for 256 locked goroutines. None may land on that thread.
It fails 5 times out of 5 when the helper is changed to
defer runtime.UnlockOSThread(), and passes 5 out of 5 with the fix:TestRunOnDedicatedThreadContainsUnsharedoes the real check, that thecaller's
/proc/thread-self/ns/mntis unchanged afterfnunshares. It needsCAP_SYS_ADMINand skips without it. It is skipped, not run, on the machinethis was prepared on: user namespaces are blocked there, so
unshare -Urmcannot give it the capability either.
go test ./pkg/plugins/gives 84 passed / 31 failed both before and after thischange. The 31 are the pre-existing user and layout specs that need root.
AI was used to write this change. No human has read it yet.