Skip to content

runtime: add metrics for zombie timers, add Timer.TailscaleRelease - #189

Merged
bradfitz merged 2 commits into
tailscale.go1.27from
bradfitz/timer-zombies
Sep 17, 2026
Merged

bradfitz merged 2 commits into
tailscale.go1.27from
bradfitz/timer-zombies

Conversation

@bradfitz

Copy link
Copy Markdown
Member
  • [tailscale] runtime, runtime/metrics, time: add zombie timer metrics and Timer.TailscaleRelease
  • [tailscale] context: optionally release the timer when a timeout context is canceled

Updates #141

…and Timer.TailscaleRelease

Stopping a timer does not remove it from the runtime's timer heap.
The timer may be on another P's heap, so Stop only marks it as a
zombie, and the owning P removes it later: when it reaches the top of
the heap, when a timer add finds it at the heap's tail, or when
zombies exceed a quarter of that heap. Until then the timer stays
reachable, and for a time.AfterFunc timer that means its func and
everything the func captures stay reachable too. For a server whose
per-request timeouts are stopped long before they would have fired,
that is a lot of retained memory with nothing to show for it.
TailscaleNumTimers has exposed the total zombie count since Go 1.22,
but it cannot say which zombies matter.

Add four gauges to runtime/metrics. /tailscale/sched/timers/tracked
is the number of timers in the per-P heaps, zombies included.
/tailscale/sched/timers/zombies is the zombie count, split into
zombies/chan and zombies/func. Channel timers become zombies
constantly, since a select with a timer case that finishes through
another case leaves the timer heaped and marked, but those pin only
the timer and its channel. Func zombies, from AfterFunc and from
runtime-internal timers such as netpoll deadlines, are the ones that
pin closures. The split needs one more per-P counter next to
timers.zombies, maintained by a zombieAdd helper at the sites that
previously adjusted zombies directly. TailscaleNumTimers now reads
the same counters.

Add time.(*Timer).TailscaleRelease. It stops the timer and, for a
func timer, also drops the runtime's reference to the func, under the
same lock hold, so that the func is collectable while the zombie
waits its turn in the heap. That is only safe if the timer is never
reset again, which Stop cannot assume, so this is a separate method
whose contract forbids it: a released timer records a state bit and
Reset panics. A channel timer is only stopped and flagged, because
the runtime finds the channel through the same arg field. The name
is deliberately ugly so that it is obviously fork-specific and stays
out of open source code that must build with upstream Go.

Adding the counter grows the timers struct embedded in p by eight
bytes, which pushed p.goroutinesCreated out of the cache line it
shares with p.tsStackHist. Swap it ahead of gcStopTime to restore
that.

Updates #141

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I4c7be1d3b8a9f5e206d1c2a48e7f39b0d5a6c1e2
…ext is canceled

A timeout or deadline context that is canceled before its deadline
stops its timer, but a stopped timer sits in the runtime's timer heap
as a zombie until the owning P gets around to removing it, which can
take until the original deadline. The timer's func holds the
timerCtx, and through it the parent context chain and every value on
it. In a server that puts a timeout on each request, every request
that completes early leaves its whole context tree pinned for as long
as its timeout was. The Go 1.22 runtime comment on clearing deleted
timers named context.WithTimeout as the motivating case.

timerCtx never resets its timer after cancel, so it can use the
fork's TailscaleRelease, which stops the timer and drops the func
reference immediately.

Make that opt-in for now: it is used only when the process starts
with TS_RELEASE_CONTEXT_TIMER=1 in its environment, so that the
effect can be A/B tested in production before it becomes the
default. The context package sits below os in the standard library
dependency order, so it reads the variable through syscall.Getenv.

The test needs the variable set on the test process and skips
otherwise. It pads the heap with live timers so that the runtime's
own zombie sweep does not remove the canceled context's timer first;
with plain Stop in place of TailscaleRelease it fails.

Updates #141

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I9e2d5a7f1c3b48e6a0d7f2c5b8e1a4d6c3f0b9e7
@bradfitz
bradfitz merged commit 32e8826 into tailscale.go1.27 Sep 17, 2026
5 checks passed
@bradfitz
bradfitz deleted the bradfitz/timer-zombies branch September 17, 2026 20:47
bradfitz pushed a commit to tailscale/tailscale that referenced this pull request Sep 17, 2026
* Go toolchain: tailscale/go@d030173...32e8826

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/go#189

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
bradfitz added a commit to tailscale/tailscale that referenced this pull request Sep 19, 2026
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.

While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.

Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.

Updates #21300
Updates tailscale/go#189
Updates golang/go#75935

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
bradfitz added a commit to tailscale/tailscale that referenced this pull request Sep 20, 2026
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.

While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.

Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.

Updates #21300
Updates tailscale/go#189
Updates golang/go#75935

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
bradfitz added a commit to tailscale/tailscale that referenced this pull request Sep 20, 2026
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.

While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.

Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.

Updates #21300
Updates tailscale/go#189
Updates golang/go#75935

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants