Skip to content

[tailscale] runtime, runtime/metrics: add goroutine stack size histogram - #186

Merged
bradfitz merged 1 commit into
tailscale.go1.27from
bradfitz/stack-size-hist
Sep 15, 2026
Merged

bradfitz merged 1 commit into
tailscale.go1.27from
bradfitz/stack-size-hist

Conversation

@bradfitz

@bradfitz bradfitz commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

TailscaleReadStackStats tells a goroutine about its own stack, but
finding out how the stacks of all the goroutines in a server are
spread across size classes still meant walking every goroutine, and
nothing reports how often stacks are being copied. The runtime itself
only exposes aggregate stack bytes, in runtime/metrics and MemStats.
For a server holding millions of idle connections, a walk on every
metrics scrape is exactly the O(goroutines) cost we want to avoid.

Add four metrics to runtime/metrics, all maintained incrementally so
that reading them never touches the goroutines.

/tailscale/sched/goroutines-by-stack-size:bytes is a histogram whose
buckets are the power-of-two stack size classes, each counting the
live goroutines currently in that class. It includes runtime system
goroutines, so at rest the buckets sum to /sched/goroutines:goroutines.

/tailscale/sched/stacks/growths:events and
/tailscale/sched/stacks/shrinks:events count stack copies since
program start in each direction, and
/tailscale/sched/stacks/copied:bytes counts the bytes of in-use stack
those copies moved. Their rates show whether stacks are thrashing,
growing on each request and being shrunk back by the GC while idle,
and what that costs.

The histogram is dead-reckoned the way the GC pacer tracks its
scannable stack total, at the same three sites (newproc1, gdestroy,
and copystack), plus oneNewExtraM so that a cgo callback goroutine
growing its stack does not push a bucket negative. Each P keeps an
int32 counter per size class, so goroutine creation and exit touch no
shared atomics; a goroutine created on one P and exiting on another
leaves offsetting entries whose sum is still correct. The counters are
int32 rather than int64 so that other Ps can read them atomically on
32-bit platforms without the alignment requirements that
TestAtomicAlignment enforces, and, like the pacer's maxStackScanDelta,
they are flushed to a global atomic once they drift a small slack from
zero so they cannot overflow. Destroying a P flushes its counters too,
like goroutinesCreated. The three cumulative counters are plain global
atomics, since a stack copy is rare and expensive next to an atomic
add.

Reading the histogram sums the global and per-P arrays under
sched.lock, so it costs O(GOMAXPROCS): about 240ns whether there are
zero or ten thousand goroutines.

The test compares the incrementally maintained histogram against a
stop-the-world walk of every goroutine, which must match exactly,
after creating goroutines, growing stacks, shrinking them via GC,
changing GOMAXPROCS in both directions, and letting the goroutines
exit. It also checks that the cumulative counters advance by at least
what one goroutine's TailscaleReadStackStats reports for itself. A
manual cgo check confirmed that a callback on a C-created thread that
grows its stack to 512 KiB is accounted for as well.

Updates #184
Updates https://github.com/tailscale/corp/issues/48120
Updates golang#81527

TailscaleReadStackStats tells a goroutine about its own stack, but
finding out how the stacks of all the goroutines in a server are
spread across size classes still meant walking every goroutine, and
nothing reports how often stacks are being copied. The runtime itself
only exposes aggregate stack bytes, in runtime/metrics and MemStats.
For a server holding millions of idle connections, a walk on every
metrics scrape is exactly the O(goroutines) cost we want to avoid.

Add four metrics to runtime/metrics, all maintained incrementally so
that reading them never touches the goroutines.

/tailscale/sched/goroutines-by-stack-size:bytes is a histogram whose
buckets are the power-of-two stack size classes from the smallest
stack (2 KiB on most platforms) through 16 MiB, plus a final bucket
for anything larger, since nobody needs to tell a 1 GiB stack from a
512 MiB one. Each bucket counts the live goroutines currently in that
class. It includes runtime system goroutines, so at rest the buckets
sum to /sched/goroutines:goroutines.

/tailscale/sched/stacks/growths:events and
/tailscale/sched/stacks/shrinks:events count stack copies since
program start in each direction, and
/tailscale/sched/stacks/copied:bytes counts the bytes of in-use stack
those copies moved. Their rates show whether stacks are thrashing,
growing on each request and being shrunk back by the GC while idle,
and what that costs.

The histogram is dead-reckoned the way the GC pacer tracks its
scannable stack total, at the same three sites (newproc1, gdestroy,
and copystack), plus oneNewExtraM so that a cgo callback goroutine
growing its stack does not push a bucket negative. Each P keeps an
int16 counter per size class, so goroutine creation and exit touch no
shared memory; a goroutine created on one P and exiting on another
leaves offsetting entries whose sum is still correct. Like the pacer's
maxStackScanDelta, a P's counter is flushed to a global atomic once it
drifts 1024 from zero, so it cannot overflow. Destroying a P flushes
its counters too, like goroutinesCreated.

The per-P array is deliberately small: 15 int16 entries, 30 bytes,
placed in the padding after p.preempt so that it shares a cache line
with p.goroutinesCreated, which newproc1 already writes. Maintaining
the histogram therefore dirties no cache line that goroutine creation
did not dirty already, and p's malloc size class is unchanged. A test
checks that layout on 64-bit platforms. The counters are int16 rather
than int64 partly for that reason and partly because
TestAtomicAlignment forbids 64-bit atomics on struct fields it cannot
prove aligned; the reader loads them without synchronization, as
gcount does with per-P fields.

The three cumulative counters are plain global atomics, since a stack
copy is rare and expensive next to an atomic add.

Reading the histogram sums the global and per-P arrays under
sched.lock, so it costs O(GOMAXPROCS): about 170ns whether there are
zero or ten thousand goroutines.

The test compares the incrementally maintained histogram against a
stop-the-world walk of every goroutine, which must match exactly,
after creating goroutines, growing stacks, growing one past the
largest exact bucket, shrinking them via GC, changing GOMAXPROCS in
both directions, and letting the goroutines exit. It also checks that
the cumulative counters advance by at least what one goroutine's
TailscaleReadStackStats reports for itself. A manual cgo check
confirmed that a callback on a C-created thread that grows its stack
to 512 KiB is accounted for as well.

Updates #184

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I197e27ca2cd17bd144a3154c25e0766033d9e845
@bradfitz
bradfitz force-pushed the bradfitz/stack-size-hist branch from c8e7b58 to 35793ad Compare September 15, 2026 02:53
@bradfitz
bradfitz merged commit a940030 into tailscale.go1.27 Sep 15, 2026
5 checks passed
@bradfitz
bradfitz deleted the bradfitz/stack-size-hist branch September 15, 2026 19:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants