Skip to content

fix: goroutine leaks in pubsub broker and job signaler - #982

Closed
nokernel wants to merge 6 commits into
leg100:masterfrom
Optable:otfd-kube-memory
Closed

nokernel wants to merge 6 commits into
leg100:masterfrom
Optable:otfd-kube-memory

Conversation

@nokernel

@nokernel nokernel commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Problem

otfd steadily grows in memory until it's OOMKilled. It's most visible when running with the kubernetes executor.

Broker.Subscribe starts a goroutine that blocks on ctx.Done() to remove the subscriber. If the caller unsubscribes before its context is done, that goroutine stays blocked, holding the subscription channel
and its pre-allocated event buffer alive.

The server runner hits this on every loop iteration: it calls AwaitAllocatedJobs, which subscribes and unsubscribes, and because the runner is in-process it passes a context that lives for the whole process. So
each job permanently leaks one goroutine plus one buffer — around 484 KB per job with OTF_SUB_BUFFER_SIZE=10000, less at the default but never reclaimed.

jobSignaler leaks the same way. It also keys subscriptions by job ID alone, so when an operation re-establishes an interrupted signal connection, the previous subscriber is orphaned and never released — and the
stale context cancellation then closes the new subscriber's channel, which the operation reads as a cancellation signal for a job that is still running.

Fix

Let the watcher goroutine exit when the subscription is removed, and track job-signal subscriptions individually so a job can have more than one.

Also, while in the kubernetes executor:

  • currentJobs dereferenced the job list before checking the error, panicking whenever listing jobs failed.
  • The job-token secret is created before its Job and only then given an owner reference. If either step failed, the secret was left with no owner to collect it — now deleted.
  • Owner reference kind is Job, matching the kubernetes kind.

Verification

New tests cover both leaks; the job-signaler test hangs against the old code, and 100 subscribe/unsubscribe cycles leak 100 goroutines before the change and none after.

Measured on a running deployment: before, 218 Mi after 388 job cycles and still climbing; after, memory plateaus at ~46 Mi across 333 cycles, where the leak would have added ~157 Mi.

dependabot Bot and others added 6 commits July 27, 2026 18:37
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.81.1 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](grpc/grpc-go@v1.81.1...v1.82.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.82.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Add --kubernetes-node-selector and --kubernetes-tolerations flags so
run jobs can be pinned to a dedicated/tainted node pool. Tolerations
use the kubectl-taint syntax key[=value]:effect (Exists when no value).
Expose the new --kubernetes-node-selector and --kubernetes-tolerations
flags through the otfd and otf-agent helm charts (via the runner
library), regenerate READMEs, bump chart versions, and document the
flags in flags.md and executors.md.
kubernetes executor: support node selector and tolerations
…g.org/grpc-1.82.1

chore(deps): bump google.golang.org/grpc from 1.81.1 to 1.82.1
Broker.Subscribe leaked a goroutine, and the subscription channel it held,
whenever a caller unsubscribed before its context was done. The server
runner subscribes on every iteration of its job-processing loop with a
process-lifetime context, leaking ~484KB per job at OTF_SUB_BUFFER_SIZE
10000. Let the goroutine exit when the subscription is removed.

jobSignaler leaked the same way, and keyed subscriptions by job ID alone,
so re-subscribing to a job orphaned the previous subscriber and could
signal cancelation to a job that was still running. Track subscriptions
individually.

Also in the kubernetes executor:

  - currentJobs panicked instead of returning 0 when listing jobs failed.
  - delete the job token secret when it is left without an owner to
    garbage collect it.
  - the owner reference kind is "Job", matching the kubernetes kind.

Read the broker's subscriptions map through an accessor in tests, which
the race detector flags otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@nokernel nokernel closed this Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants