Repository navigation
Conversation
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.81.1 to 1.82.1. - [Release notes](https://github.com/grpc/grpc-go/releases) - [Commits](grpc/grpc-go@v1.81.1...v1.82.1) --- updated-dependencies: - dependency-name: google.golang.org/grpc dependency-version: 1.82.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com>
Add --kubernetes-node-selector and --kubernetes-tolerations flags so run jobs can be pinned to a dedicated/tainted node pool. Tolerations use the kubectl-taint syntax key[=value]:effect (Exists when no value).
Expose the new --kubernetes-node-selector and --kubernetes-tolerations flags through the otfd and otf-agent helm charts (via the runner library), regenerate READMEs, bump chart versions, and document the flags in flags.md and executors.md.
kubernetes executor: support node selector and tolerations
…g.org/grpc-1.82.1 chore(deps): bump google.golang.org/grpc from 1.81.1 to 1.82.1
Broker.Subscribe leaked a goroutine, and the subscription channel it held,
whenever a caller unsubscribed before its context was done. The server
runner subscribes on every iteration of its job-processing loop with a
process-lifetime context, leaking ~484KB per job at OTF_SUB_BUFFER_SIZE
10000. Let the goroutine exit when the subscription is removed.
jobSignaler leaked the same way, and keyed subscriptions by job ID alone,
so re-subscribing to a job orphaned the previous subscriber and could
signal cancelation to a job that was still running. Track subscriptions
individually.
Also in the kubernetes executor:
- currentJobs panicked instead of returning 0 when listing jobs failed.
- delete the job token secret when it is left without an owner to
garbage collect it.
- the owner reference kind is "Job", matching the kubernetes kind.
Read the broker's subscriptions map through an accessor in tests, which
the race detector flags otherwise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
otfd steadily grows in memory until it's OOMKilled. It's most visible when running with the kubernetes executor.
Broker.Subscribe starts a goroutine that blocks on ctx.Done() to remove the subscriber. If the caller unsubscribes before its context is done, that goroutine stays blocked, holding the subscription channel
and its pre-allocated event buffer alive.
The server runner hits this on every loop iteration: it calls AwaitAllocatedJobs, which subscribes and unsubscribes, and because the runner is in-process it passes a context that lives for the whole process. So
each job permanently leaks one goroutine plus one buffer — around 484 KB per job with OTF_SUB_BUFFER_SIZE=10000, less at the default but never reclaimed.
jobSignaler leaks the same way. It also keys subscriptions by job ID alone, so when an operation re-establishes an interrupted signal connection, the previous subscriber is orphaned and never released — and the
stale context cancellation then closes the new subscriber's channel, which the operation reads as a cancellation signal for a job that is still running.
Fix
Let the watcher goroutine exit when the subscription is removed, and track job-signal subscriptions individually so a job can have more than one.
Also, while in the kubernetes executor:
Verification
New tests cover both leaks; the job-signaler test hangs against the old code, and 100 subscribe/unsubscribe cycles leak 100 goroutines before the change and none after.
Measured on a running deployment: before, 218 Mi after 388 job cycles and still climbing; after, memory plateaus at ~46 Mi across 333 cycles, where the leak would have added ~157 Mi.