Repository navigation
Bump Go 1.27, cherry-pick memory wins - #182
Merged
Merged
Conversation
For golang#81250 Fixes golang#81252 Change-Id: I973ead54e02cbc86b488c446c9d581f86a6a6964 Reviewed-on: https://go-review.googlesource.com/c/go/+/822805 Auto-Submit: Filippo Valsorda <filippo@golang.org> Reviewed-by: Dmitri Shuralyov <dmitshur@google.com> Reviewed-by: David Chase <drchase@google.com> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Roland Shoemaker <roland@golang.org> (cherry picked from commit e656c14) Reviewed-on: https://go-review.googlesource.com/c/go/+/824904 Auto-Submit: Michael Pratt <mpratt@google.com> Reviewed-by: Michael Pratt <mpratt@google.com>
…reemption on windows Async preemption on Windows includes some synchronous components: one thread pauses another and inspects and operates on its state. If the paused thread is holding locks, the thread doing the preemption won't be able to acquire them, just as if it were a signal handler running on the paused thread. Mutex profiling results in acquiring a mutex within runtime.unlock, if the code within unlock2 determines that the current M doesn't hold any other locks. Usually signal handlers aren't allowed to acquire any locks at all, but Windows' preemptM does. Communicate the restriction to unlock2 with the same mechanism we use when a gscan bit is held: with a call to acquireLockRankAndM. Fixes golang#81145 Change-Id: I722bb9ce8e15e01ec70bb0efc1bec4eca823cfa6 Reviewed-on: https://go-review.googlesource.com/c/go/+/818940 Auto-Submit: Rhys Hiltner <rhys.hiltner@gmail.com> Reviewed-by: Dmitri Shuralyov <dmitshur@google.com> Reviewed-by: David Chase <drchase@google.com> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Zxilly Chou <zxilly@outlook.com> Reviewed-by: Quim Muntal <quimmuntal@gmail.com> Reviewed-by: Russ Cox <rsc@golang.org> (cherry picked from commit 3d727ce) Reviewed-on: https://go-review.googlesource.com/c/go/+/824104 Reviewed-by: Keith Randall <khr@google.com> Reviewed-by: Keith Randall <khr@golang.org>
…entfd type confusion
netpoll must distinguish the netpollBreak eventfd from a socket fd. It does so by comparing ev.Data against &netpollEventFd: if they are equal the event is for the eventfd, otherwise it is for a socket and ev.Data holds a tagged pointer made up of a *pollDesc and an fdseq.
On 64-bit systems this works as expected, but on 32-bit little-endian systems it can lead to type confusion.
netpollinit stores the address of netpollEventFd as a uintptr:
*(**uintptr)(unsafe.Pointer(&ev.Data)) = &netpollEventFd
so for the eventfd only the lower half of ev.Data is written: ev.Data[0:4] holds &netpollEventFd and ev.Data[4:8] is left untouched. netpollopen, in contrast, stores a full tagged pointer, so that ev.Data[0:4] holds fdseq and ev.Data[4:8] holds the *pollDesc.
Because a uintptr is only 4 bytes on 32-bit systems, the comparison in netpoll:
if *(**uintptr)(unsafe.Pointer(&ev.Data)) == &netpollEventFd {
tests only ev.Data[0:4], so for a socket it effectively compares fdseq against the address of netpollEventFd. In the vast majority of cases the two differ and the fd types are distinguished correctly.
However, in a program that runs for a long time or opens many sockets, fdseq can eventually grow large enough to equal &netpollEventFd. When that happens netpoll mistakes a socket for the eventfd and the runtime aborts with a fatal error such as:
runtime: netpoll: eventfd ready for 5
fatal error: runtime: netpoll: eventfd ready for something unexpected
Fix this by using the tagged-pointer encoding for both cases and detecting the eventfd by its nil *pollDesc.
Updates golang#72900
Fixes golang#81226
Change-Id: I9779dc038f437eeb23fbd1b7057eee5f5951ac13
GitHub-Last-Rev: 521495e
GitHub-Pull-Request: golang#81037
Reviewed-on: https://go-review.googlesource.com/c/go/+/819800
Reviewed-by: Ian Lance Taylor <iant@golang.org>
Reviewed-by: David Chase <drchase@google.com>
LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com>
Reviewed-by: Cherry Mui <cherryyz@google.com>
(cherry picked from commit b9ba134)
Reviewed-on: https://go-review.googlesource.com/c/go/+/824304
Auto-Submit: Michael Pratt <mpratt@google.com>
Reviewed-by: Alan Donovan <adonovan@google.com>
Reviewed-by: Michael Pratt <mpratt@google.com>
…into bradfitz/picks
tomhjp
approved these changes
Sep 9, 2026
…r error channel pool The move of HTTP/2 into std made errChanPool a field on http2.Server so that pooled channels wouldn't be reused across synctest bubbles. But that regressed memory for servers using the ServeConn-per-conn pattern (HTTP/2 over hijacked or tunneled conns): every sync.Pool re-pins after each GC, allocating a GOMAXPROCS-sized poolLocal array (4 KB at GOMAXPROCS=32) per pool, and a per-conn pool provides no reuse anyway. On a production proxy with ~400k conns, the poolLocal arrays accounted for ~2 GB, ~10% of heap. Instead, go back to a single global pool and skip pooling entirely when the current goroutine is in a synctest bubble. Updates #174 Change-Id: I8a8df547507b185b7943ab8d227da8262acf8a49 Reviewed-on: https://go-review.googlesource.com/c/go/+/825424 LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Damien Neil <dneil@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org> Reviewed-by: Nicholas Husin <husin@google.com>
…coder table before SETTINGS ack A server configured with a MaxDecoderHeaderTableSize below the initial 4096 bytes (RFC 7540, section 6.5.2) applied it to its HPACK decoder at connection start. But the server doesn't know when the client has received and applied a SETTINGS value until the client acknowledges the SETTINGS frame (RFC 7540, section 6.5.3), and the client's encoder signals a table size reduction with a dynamic table size update at the beginning of the first header block following the settings acknowledgment (RFC 7541, section 4.2). Until then, the client may keep using the initial 4096-byte table. A spec-compliant client (including Go's own http2 client) that sent a header block referencing a dynamic table entry in the first RTT was rejected with a connection-fatal COMPRESSION_ERROR. Instead, start the decoder at the initial table size and apply the configured smaller size once the client acks our SETTINGS. All header blocks after the ack were encoded by a client that has processed the setting, and a compliant encoder begins its next header block with a dynamic table size update. Memory during the window is bounded by the same 4096 bytes as today's default. Updates #180 Change-Id: Idecd25a1e0473c3cff314922ddc97ace14cb2e5e Reviewed-on: https://go-review.googlesource.com/c/go/+/827024 Reviewed-by: Nicholas Husin <husin@google.com> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Damien Neil <dneil@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org> Auto-Submit: Brad Fitzpatrick <bradfitz@golang.org>
…le blocked in Read A Conn's rawInput buffer grows to the maximum record size (~17 kB after receiving full-sized records) and, being a bytes.Buffer, never shrinks. A connection blocked in Read waiting for a new record, often for minutes or hours on long-polling or mostly idle connections, pinned that memory the whole time. Servers with millions of open TLS connections strand gigabytes of heap in buffers holding no data. The hand buffer similarly retained the peer's largest handshake flight for the life of the connection. Instead, return an empty grown rawInput buffer to a sync.Pool before blocking to wait for a new record, mirroring outBufPool on the write side. The record header is read into a small per-connection buffer, and readFromUntil switches back to a pooled record-sized buffer only once the header arrives and the payload length is known. The hand buffer is likewise pooled: it is returned once the handshake completes and after buffered post-handshake messages are consumed. Measured with 1000 idle TLS 1.3 server connections that had each received 16 kB records before blocking in a 4-byte Read (heap bytes per connection, runtime.MemProfileRate=1): HeapAlloc/conn: 23020 B -> 4468 B HeapInuse/conn: 25509 B -> 5726 B The remainder is mostly the AES-GCM cipher states (~1.8 kB), the Conn struct itself (~1 kB), and the small header buffer (~0.6 kB). Throughput is unchanged within noise (geomean +0.4%): benchstat of -bench=Throughput -count=6: no change in 17 of 28 cases, worst case +1.2%, best case -1.8%. Pooling the hand buffer also drops an allocation and ~1% of bytes per server handshake (geomean of -bench=HandshakeServer -benchmem). Fixes golang#47672 Updates golang#81348 Updates #174 Change-Id: Idb6f9361d888977a2d655fb2d57da0fa61f61ca5 Reviewed-on: https://go-review.googlesource.com/c/go/+/827524 LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Roland Shoemaker <roland@golang.org> Reviewed-by: Daniel McCarney <daniel@binaryparadox.net> Reviewed-by: Michael Pratt <mpratt@google.com>
…f func field into a method The field was never configurable: it was unexported and only ever set to its one default implementation by NewFramer, even in its previous life in x/net/http2. Make it a plain method, removing the per-Framer closure allocation and the indirection, and drop the stale TODO about making it configurable. No behavior change. Updates golang#80735 Updates #174 Change-Id: I694312b683a40d441e02faf089c08922f30bf946 Reviewed-on: https://go-review.googlesource.com/c/go/+/828204 LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Nicholas Husin <husin@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org> Auto-Submit: Damien Neil <dneil@google.com> Reviewed-by: Damien Neil <dneil@google.com>
…/1 keep-alive connections As part of my recent mission to reduce pinned memory during idle blocked reads, this addresses the bulk of HTTP/1.1 memory wastage. We already had bufio memory pools in this package, but they weren't being used in this case. After we do the first HTTP/1 request on a connection and are about to read the second, we already had a read timeout on the connection (the idle timeout). That next request might be ready ~right away, or it might time out. This splits up that idle timeout into two phases: a quick little timeout to see if the server's hot enough to have a new request ready, and the the remainder of the time we would've waited anyway. If we don't get a new HTTP request ~right away (50 ms, currently), then we release our bufios (read+write) to the existing pools and enter our happy memory-released long idle. But on hot servers, no new pool machinery is involved and the same bufios remain. Measured with 1000 idle HTTP/1.1 keep-alive TLS connections after POST exchanges with 16 kB uploads and 4 kB responses (heap bytes per connection, runtime.MemProfileRate=1): HeapAlloc/conn: 13852 B -> 5335 B Updates golang#80735 Updates #174 Change-Id: I64b27bb7f1a65d9cb65cc0c38ebc1faa6c5a7682 Reviewed-on: https://go-review.googlesource.com/c/go/+/828504 Reviewed-by: Nicholas Husin <husin@google.com> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Auto-Submit: Damien Neil <dneil@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org> Reviewed-by: Damien Neil <dneil@google.com>
… read buffers on idle conns A Framer's readBuf grows to the size of the largest frame received (16 kB with the default max frame size) and was retained for the life of the connection, including while blocked in ReadFrame waiting, often for minutes or hours on long-polling or mostly idle connections, for the next 9-byte frame header. Servers with many thousands of open HTTP/2 connections strand most of their per-connection memory in these buffers holding no data. The previous frame is already documented as invalid as of the next ReadFrame call, so release the read buffer to a sync.Pool at the top of ReadFrame, before blocking in ReadFrameHeader, and get one back only once a frame header with a large payload length arrives. Buffers up to 1 kB stay attached to their Framer, so connections receiving only small control frames don't pay for the pooling, and if the Framer's reader is a bufio.Reader with data already buffered (the Transport's case, mid-stream), the buffer is kept since no blocking wait is about to happen. Measured with 1000 idle HTTP/2 connections on separate TCP flows, each blocked reading a frame header after POST exchanges with 16 kB bodies (heap bytes per connection, runtime.MemProfileRate=1): HeapAlloc/conn: 32715 B -> 16363 B HeapInuse/conn: 40779 B -> 22216 B The largest remaining consumers are the Framer's write buffer (~4.8 kB), the serverConn struct (~1.8 kB), TLS cipher state (~1.8 kB), and the HPACK tables (~1 kB). Updates golang#80735 Updates #174 Change-Id: I3d3fd1a237b90e716b708789dea40e3eeb0dbd25 Reviewed-on: https://go-review.googlesource.com/c/go/+/828205 Reviewed-by: Damien Neil <dneil@google.com> Reviewed-by: Nicholas Husin <husin@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com>
… write buffers on idle conns A Framer's wbuf grows to the size of the largest frame written (typically a full DATA frame) and was retained for the life of the connection, pinning several kB per idle connection. The frame is fully written out by the end of endWrite, so release large write buffers to a sync.Pool there and have startWrite adopt a pooled buffer for the next frame. Buffers up to 1 kB stay attached to their Framer, so connections writing only small control frames don't pay for the pooling. Measured with 1000 idle HTTP/2 connections on separate TCP flows after POST exchanges with 4 kB responses (heap bytes per connection, runtime.MemProfileRate=1, HPACK dynamic tables disabled, and with the corresponding read buffer fix already applied): HeapAlloc/conn: 15265 B -> 10377 B HeapInuse/conn: 20357 B -> 15114 B http2 benchmarks are unchanged within noise. The largest remaining consumers are the serverConn struct (~1.8 kB), TLS cipher state (~1.8 kB), and the tls.Conn struct (~0.9 kB). Updates golang#80735 Updates #174 Change-Id: I7c1e08e2f6df944d34a2ef7a5f5091b1d3648b95 Reviewed-on: https://go-review.googlesource.com/c/go/+/828206 LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Damien Neil <dneil@google.com> Reviewed-by: Nicholas Husin <husin@google.com> Reviewed-by: Nicholas Husin <nsh@golang.org>
bradfitz
force-pushed
the
bradfitz/picks
branch
from
September 9, 2026 21:23
816f610 to
f147da0
Compare
bradfitz
added a commit
to tailscale/tailscale
that referenced
this pull request
Sep 9, 2026
Pulls in tailscale/go#182 Updates tailscale/corp#29053 Updates #21064 Change-Id: If028b1eec6f978593e000e5f961e7d2acc488f35 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
bradfitz
added a commit
to tailscale/tailscale
that referenced
this pull request
Sep 10, 2026
Pulls in tailscale/go#182 Updates tailscale/corp#29053 Updates #21064 Change-Id: If028b1eec6f978593e000e5f961e7d2acc488f35 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
bradfitz
added a commit
to tailscale/tailscale
that referenced
this pull request
Sep 10, 2026
Pulls in tailscale/go#182 Updates tailscale/corp#29053 Updates #21064 Change-Id: If028b1eec6f978593e000e5f961e7d2acc488f35 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Updates #174
Updates #180
Updates https://github.com/tailscale/corp/issues/29053
Updates tailscale/tailscale#21064