Skip to content

Add scale down to HistoryPubnetParallelCatchupV2 - #433

Open
Jonathan-Eid wants to merge 22 commits into
stellar:mainfrom
Jonathan-Eid:jonathan/catchup-scale-down-v2
Open

Jonathan-Eid wants to merge 22 commits into
stellar:mainfrom
Jonathan-Eid:jonathan/catchup-scale-down-v2

Conversation

@Jonathan-Eid

@Jonathan-Eid Jonathan-Eid commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Currently, the CatchupV2 mission does not scale down workers when the amount of remaining/in-progress jobs becomes less than the amount of parallel workers.

Implementation

Before, each parallel worker was a replica of the same statefulset, which prevented picking out an idle worker to scale down.

  • Workers are now bare Pods (hostname/subdomain for headless-service DNS) instead of statefulset replicas: one pod per worker, no controller.
  • The marking decision lives in job_monitor.py (phase 1c). It marks surplus idle workers into the Redis set retiring, where surplus is measured against outstanding = queued + in_progress, keeping a reserve of MIN_UNMARKED_WORKERS (default 8) counted against all unmarked workers.
  • /status publishes retirable: workers marked on an earlier pass and still idle.
  • The F# driver only reads retirable, collects logs, deletes those pods, and tracks livePods. No Redis access from the driver; readyPods and the redis-cli exec are gone.
  • worker.sh checks SISMEMBER retiring $POD_NAME before claiming a job.
  • The orphan sweep discovers pc-v2 releases via worker pods and reaps ownerless pods.

Testing

  • Local ssc-test runs at 40 and 900 workers.
  • Jenkins config-test build # 194 (40 workers, 48 ranges): reserve invariant held (retirable capped at 32 = 40 - 8), 48/48 ranges completed.

Caveats

  • Up to 64 workers can have their logs collected and scaled down in one pass; keeping the log collection synchronous and blocking the loop. This keeps the code cleaner for not much affect on the loop time and avoids managing threads. A log collection failure doesn't blocks retiring the batch -- any error skips that pod's logs.
  • Worker pods are not recreated if lost (accepted).
  • retiring is never pruned, so names stay in retirable after deletion; the driver filters them via livePods.
  • A dead pod that still owns a job occupies a reserve slot until the end-of-run requeue, hence the floor of 8.
  • The job monitor image must be rebuilt and pushed before merge, since the driver reads retirable from /status.

🤖 Generated with Claude Code

A run holds its full fleet for its whole duration, including the tail
where almost nothing is left: 37.7% of prod worker-hours are idle, and at
the extreme 1020 workers wait 4.5 hours on 4 remaining jobs.

Each worker becomes its own single-replica StatefulSet, so any idle
worker can be removed without regard to position; scaling one shared
StatefulSet only ever deletes a suffix of its ordinals, which a single
long range on a high ordinal can pin for hours.

Each poll marks up to `ready - outstanding` idle workers into a Redis set
that worker.sh checks before claiming, and deletes the ones marked on an
earlier pass that are still idle. The one-pass gap is the safety
property: a marked worker stops claiming within its 10s loop, so by the
next poll it provably takes no new work. Busy workers come from reading
job_owners directly rather than from the status snapshot, which the
monitor only republishes after serially pinging every owner and so can be
older than the mark.

Capacity is Running pods, not the replica count: a Pending pod owns no
job and so reads as idle capacity it cannot supply.

Logs are collected before each deletion, since /data is emptyDir, and a
failed collection skips the removal rather than losing them.
A pod being deleted still reports phase Running, so readyPods counted
pods on their way out. That inflated capacity, marked more workers than
the queue could spare, and re-selected the pods just deleted -- whose
StatefulSets were already gone, so the delete threw NotFound and aborted
the whole pass. A run lost a pass after every successful wave.

The ready set also has to drop the workers removed earlier in the same
pass, or the marking step reads a fleet size that is one wave stale and
re-marks what it just deleted.
Both reads ran on every poll of a multi-hour run even while the queue
outran the fleet and nothing was marked, and the pod list at 1024 workers
is not cheap. They are now skipped unless something is marked or the
queue has dropped below the fleet, which also stops the ramp logging a
skipped pass every poll: with no pod Running yet, taking the head of an
empty ready set threw and the handler reported it as a real failure.

The guard reads livePods rather than the ready set, since consulting the
ready set is the cost being avoided. livePods is always the larger of the
two, so the guard errs toward doing the work rather than skipping it.

redisIn now takes the ready set and picks a host itself, returning
nothing when there is none, so callers stop reaching for Seq.head.
Log collection is serial and runs inside the poll loop, so an unbounded
batch blocks it for as long as the batch takes. A 400-worker run retired
213 in one pass and spent 103 seconds collecting, and that was with tiny
archives: each worker had run only 2 ranges. A prod worker churns dozens
over hours, so per-pod cost is seconds rather than half a second, and an
uncapped pass late in a 1024-worker run would block for tens of minutes
with no timeout on any single exec.

Capped at 32 per pass, so the worst case is bounded by the cap rather
than by the fleet size. Later passes drain the rest.
helm was left to use the kubeconfig's current namespace while every other
call in the mission honours context.namespaceProperty. A run launched
with --namespace ssc-config-test therefore installed its chart into
stellar-supercluster: 1024 workers and a monitor landed in the
production namespace, while the driver polled and scaled an empty one,
so scale-down silently did nothing for the whole run.

Both uninstall paths get it too. They were previously consistent with
the install only by accident -- both defaulted to the same current
namespace -- which would have broken the moment the install was fixed
alone.
Once every worker is marked retiring, a job the monitor puts back on the
queue can never be claimed: worker.sh refuses to claim while marked, and
there is no way to unmark. A single recoverable orphan then hangs the run.

Observed on a 1024-worker run: all 4009 ranges were accounted for, one job
lost its owner upstream, the monitor correctly requeued it, and nothing was
left willing to take it. The run had to be aborted.

Two changes, both in the mark step. The reserve is now counted against
unmarked workers rather than `ready`, which includes marked pods that have
not been deleted yet and so let the reserve erode away over successive
passes. And the reserve never drops below minUnmarkedWorkers, so a run whose
queue empties still has somewhere to put returned work.

Three rather than one so the reserve survives a worker restarting or its
node being disrupted, and so several returned jobs drain in parallel. The
cost is three idle workers, about two nodes, against a tail that otherwise
idles the whole fleet.
RunRemoteCommandAndCaptureOutput returned void and wrote the k8s status
channel to the console, so neither caller could tell a failed exec from a
command that produced no output. Two fail-open paths followed.

`redisIn` reads job_owners to decide which workers are idle. A failed exec
left it with zero lines, which reads as "no worker is busy" and makes every
worker look retirable -- the exact check that is supposed to stop a busy
worker being retired. Two such passes in a row would mark a worker and then
delete it mid-range.

`collectLogsFromPods` treats an exception as the only failure, but tar
exiting non-zero raises nothing, so a failed archive counted as collected
and the worker was deleted with its logs.

The exit code was already available: channel 3 carries a V1Status whose
causes hold it, and Kubernetes.GetExitCodeOrThrow already parses exactly
that for RunRemoteCommand in the same file. Renamed the locals and fixed
the comments too -- a variable called `stderr` holding the status channel
is what made this invisible.

Both callers now fail closed: redisIn raises into the poll loop's handler,
which skips the pass with no state mutated, and a failed tar puts the pod
on the failed list so retirement is skipped for it.
Copilot AI balanced review requested due to automatic review settings August 27, 2026 17:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds dynamic worker scale-down to parallel catchup by assigning each worker its own StatefulSet.

Changes:

  • Introduces worker retirement through Redis.
  • Collects logs before deleting idle workers.
  • Returns remote command exit codes and scopes Helm commands to the namespace.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.

File Description
catchup_workers.yaml Creates one StatefulSet per worker.
worker.sh Prevents retiring workers from claiming jobs.
MissionHistoryPubnetParallelCatchupV2.fs Implements worker selection, log collection, and retirement.
RemoteCommandRunner.cs Returns remote command exit codes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs Outdated
Comment thread src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs Outdated
The reserve was subtracted from the idle pool alone, but `outstanding`
counts in-progress jobs too, and those already have a worker each. With
1000 workers, 700 busy and 700 jobs outstanding, the surplus came out as
300 - 700 and nothing was marked at all: 300 idle workers were held to
cover work that was already being served.

Counting the surplus against every unmarked worker gives 1000 - 700 = 300,
and those 300 idle workers are marked. Only idle workers are ever marked,
and the reserve still keeps `max outstanding minUnmarkedWorkers` workers
unmarked, so a requeued job still has somewhere to go -- an unmarked busy
worker becomes claimable as soon as it finishes.

Run stellar#192 never sat in that state: in-progress fell from 1022 to 80 in three
minutes, so the largest instantaneous gap was about 80 workers and
retirement was capped at 32 a pass throughout. A run whose ranges vary
enough for in-progress to decay gradually would sit there and mark nothing.
Copilot AI review requested due to automatic review settings August 27, 2026 17:40

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated no new comments.

Suppressed comments (2)

Previously missed (1) — in code that hasn't changed since the last review.

src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs:255

  • This updated signature takes explicit pod names, but the function documentation still says it automatically derives them from pubnetParallelCatchupNumWorkers. Update the comment so callers are not given an obsolete contract.
// Returns the pods whose collection raised; an empty archive is success.

src/MissionParallelCatchup/parallel_catchup_helm/templates/catchup_workers.yaml:28

  • The new StatefulSet name makes every pod end in -0 (...-stellar-core-<worker-index>-0), but worker.sh derives core_id from the final hyphen-separated segment and stores it in each Redis metric. Consequently, every worker is now reported as core 0. Pass the worker index to the container or update the script to extract the penultimate segment.
  name: {{ $.Release.Name }}-stellar-core-{{ $i }}

A worker that never claimed a job has an empty /data, so the tar glob
matches nothing and tar exits 2. Checking that exit code put those workers
on the failed list, which blocks retirement for the whole batch and never
clears, because the batch is rebuilt the same way every pass.

It lands on precisely the workers scale-down exists to remove: a surplus
worker that never got work. Seen on ssc-test with 12 workers and 8 ranges,
where the 4 that never claimed blocked retirement indefinitely.

redisIn keeps its exit-code check. There a failed exec yields an empty busy
set, which reads as "no worker is busy" and makes every worker look
retirable, so failing closed there is worth having.
Copilot AI review requested due to automatic review settings August 27, 2026 17:49

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.

Suppressed comments (2)

src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs:305

  • Running is not equivalent to a ready pod: for example, a pod in CrashLoopBackOff normally retains phase Running. Because redisIn always selects the first member of this set, one such lower-sorted pod can make every Redis exec fail and disable scale-down indefinitely even when other workers are healthy. Filter on the pod's Ready=True condition as well.
    |> Seq.filter (fun pod -> pod.Status.Phase = "Running" && isNull (box pod.Metadata.DeletionTimestamp))

src/MissionParallelCatchup/parallel_catchup_helm/templates/catchup_workers.yaml:28

  • This new StatefulSet name makes each pod end in -<worker-index>-0. The worker still derives core_id from the final segment (worker.sh:88), so every worker now reports ID 0 in the raw catchup metrics instead of its actual worker index. Update that extraction (or pass the worker-index label through the downward API) alongside this naming change.
  name: {{ $.Release.Name }}-stellar-core-{{ $i }}

Comment thread src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs
Per-worker StatefulSets renamed pods to <release>-stellar-core-<i>-0, so the
last hyphen-separated segment is the StatefulSet's own -0 suffix rather than
the worker index. Every metrics record in run stellar#192 reported core_id 0.

Nothing consumes the field -- the monitor discards it and keeps only the two
durations -- but a value that is always 0 reads as real data, so pass the
index the chart already has.
Copilot AI review requested due to automatic review settings August 27, 2026 18:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.

Suppressed comments (2)

Previously missed (2) — in code that hasn't changed since the last review.

src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs:535

  • If a later chunk fails, earlier SADD calls have already committed in Redis, but marked is updated only after every chunk succeeds. Those workers stop claiming while the driver still treats them as unmarked; if the outstanding count then rises, they may never be selected again for tracking/removal. Record each chunk in marked immediately after its successful SADD.
                    for chunk in List.chunkBySize 30 toMark do
                        let names = chunk |> List.map (sprintf "'%s'") |> String.concat " "

                        redisIn context ready (sprintf "SADD \"%s-retiring\" %s" helmReleaseName names)

src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs:256

  • The function no longer derives pod names from pubnetParallelCatchupNumWorkers, so the preceding numbered contract is stale and can mislead callers. Document that the caller supplies the pod list.
// Returns the pods whose collection raised; an empty archive is success.
let collectLogsFromPods (context: MissionContext) (podNames: string list) : string list =

Comment thread src/MissionParallelCatchup/parallel_catchup_helm/files/worker.sh Outdated
Comment thread src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs Outdated
Comment thread src/MissionParallelCatchup/parallel_catchup_helm/files/worker.sh
Copilot AI review requested due to automatic review settings August 27, 2026 18:28
The sweep helm-uninstalls a PCv2 release before deleting resources
individually, so the release secret does not end up dangling at deleted
workloads. It found the release by matching a StatefulSet name ending in
-stellar-core, which no longer happens: per-worker StatefulSets end in
-stellar-core-<index>. Abandoned runs therefore had their workloads deleted
one by one with the release left behind.

Matching on the -stellar-core segment instead, and collapsing the names to a
set, so a 1024-worker release is uninstalled once rather than 1024 times.
@Jonathan-Eid
Jonathan-Eid force-pushed the jonathan/catchup-scale-down-v2 branch from 333f80e to e178fc2 Compare August 27, 2026 18:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated 1 comment.

Suppressed comments (2)

Previously missed (1) — in code that hasn't changed since the last review.

src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs:305

  • Running does not imply that the stellar-core container is Ready; CrashLoopBackOff pods commonly retain phase Running. Because redisIn always chooses Seq.tryHead, one low-index crash-looping pod can be selected every pass and prevent scale-down even while other workers are healthy. Restrict this set to pods whose worker container reports Ready (and ideally retry another member if an exec races with termination).
    pods.Items
    |> Seq.filter (fun pod -> pod.Status.Phase = "Running" && isNull (box pod.Metadata.DeletionTimestamp))

src/MissionParallelCatchup/parallel_catchup_helm/files/worker.sh:30

  • The retirement check is separate from both LMOVE and the ownership HSET, leaving a deletion race: a worker can observe “not retiring,” be paused, then get marked; on the next poll the driver sees no owner and deletes it just as it claims a range. Make the membership check, claim, and ownership registration one atomic Redis operation (for example, a Lua script) so every claimed range is visible before a marked worker can be considered removable.
# Stop claiming once the driver marks us, so it can remove us without interrupting a range.
if [ "$(redis-cli -h "$REDIS_HOST" -p "$REDIS_PORT" SISMEMBER "$RELEASE_NAME-retiring" "$POD_NAME")" = "1" ]; then
    echo "$(date) $POD_NAME is retiring; not claiming."
    sleep $SLEEP_INTERVAL
    continue
fi

Comment thread src/FSLibrary/StellarOrphanSweep.fs Outdated
The function no longer derives pod names from pubnetParallelCatchupNumWorkers;
callers supply them. The remaining bullets restated the tar command and output
path, both already commented inline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 8, 2026 15:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated 1 comment.

Comment thread src/FSLibrary/MissionHistoryPubnetParallelCatchupV2.fs Outdated

@jayz22 jayz22 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for proposing this change. I agree with the direction of automatic scaling down workers and saving on idle costs. However, I'm not quite comfortable with this change as is, since it breaks the existing pattern and the role boundaries of the components.
The existing design is the monitor being a thin layer that manages states and surfaces information to the mission where it performs simple actions (logging, cleanup). This PR breaks that relation by adding a bunch of state tracking and probing logic inside the mission itself.
Is it possible to have the job monitor perform the state tracking and labeling, return the exact list of pods to be reclaimed, and the mission just claims them?

Also to point out the main structure change: going from a mission having 1 statefulset with 1024 (num_workers) pods, to 1024 statefulsets with 1 pods each. While this should not have a performance impact, it does make the sts feel a little bit redundant. With additional Kubernetes objects (sts, controller revision), simple mission level queries like get Pods and the list StatefulSet will run slower. One of the main motivations for doing parallel catchup v2 was to reduce the Kubernetes overhead and API server loads. and as a result we've adopted this simple design where api calls are minimized. This reverts some of it back. Don't think this is a deal breaker, but just something to be aware.

@Jonathan-Eid

Jonathan-Eid commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

The existing design is the monitor being a thin layer that manages states

We can move the state logic to the job monitor and the mission can cleanup the pod logs and delete them.

Also to point out the main structure change: going from a mission having 1 statefulset with 1024 (num_workers) pods, to 1024 statefulsets with 1 pods each.

We could just go with bare pods, they can restart on their own and be deterministically named. They just won't reschedule if they're evicted due to node loss, etc, and we'll just be down a few workers during the run. Anything else would need controller logic (putting a k8s client in the job_monitor). I'm more confident in this approach since your orphaned jobs fix.

@jayz22

jayz22 commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

We could just go with bare pods, they can restart on their own and be deterministically named. They just won't reschedule if they're evicted due to node loss, etc, and we'll just be down a few workers during the run.

Yeah, that's why I didn't suggest it as a fix. Not sure if the node loss issue is worth the overhead improvement. Will leave this up to you though.

@jayz22

jayz22 commented Sep 14, 2026

Copy link
Copy Markdown
Contributor
image

@Jonathan-Eid I also saw you posted this figure on the thread. Did this finish in 3.4 hours? I didn't see anything in this PR that attribute to perf improvement but want to make sure I'm not missing anything

@Jonathan-Eid

Jonathan-Eid commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

Not that that screenshot is revealing, but I didn't feel like posting any internal systems screenshot in a Public PR.

This PR change does not have any performance improvements other than log cleanup happening sooner. That's all moore's law by upgrading the machine type : p

If we ever want to squeeze out the last hour of performance, we'd have to discuss the K8s job per ledger architecture.
Job throttling is still gonna happen in this architecture, but m8a.2xl improvements mask most of it.

- job_monitor.py marks surplus idle workers (reserve counted against all unmarked, MIN_UNMARKED_WORKERS=8), skips Pending/dead pods via DNS, publishes `retirable`
- driver only collects logs and deletes pods listed in `retirable`; cap 64/pass; redis exec, readyPods and marking math removed
- workers are bare Pods with hostname/subdomain instead of per-worker StatefulSets
- worker.sh: retiring key unprefixed; RELEASE_NAME injection dropped
- orphan sweep discovers pc-v2 releases via pods

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 22, 2026 15:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Worker retirement spans Redis coordination, Kubernetes lifecycle handling, DNS, log preservation, and signal-driven cleanup without an atomic claim boundary.

Review effort: Balanced
Findings: None

Resolved since last review (2)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 22, 2026 16:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The orphan sweep can delete unrelated ownerless Pods in the namespace.

Review effort: Balanced
Findings: None

Previously missed (1)

In code that hasn't changed since last review

Medium severity Restrict startup sweep to PCv2 worker pods

src/​FSLibrary/​StellarOrphanSweep.fs:122

The automatic startup sweep now selects every ownerless pod older than the cutoff, not only the PCv2 worker pods this block is intended to reap. That can delete unrelated standalone/debug pods in the shared namespace. Restrict this list to the worker labels introduced by the chart (or an equally exact PCv2 naming predicate) in addition to checking owner references.

@Jonathan-Eid

Copy link
Copy Markdown
Contributor Author

Redis state is now out of the mission driver code. The Mission now reads from the job_monitor status. The job_monitor now uses redis to mark the workers as retireable.

@Jonathan-Eid
Jonathan-Eid requested a review from jayz22 September 22, 2026 17:09
Copilot AI review requested due to automatic review settings September 22, 2026 18:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The expanded destructive cleanup scope is not reflected in user-facing documentation, and one new diagnostic identifies the wrong configuration input.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

Comment thread src/FSLibrary/StellarOrphanSweep.fs
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 22, 2026 18:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The distributed retirement protocol and destructive orphan cleanup warrant final human validation.

Review effort: Balanced
Findings: None

Resolved since last review (1)

@anupsdf
anupsdf requested a review from sisuresh October 2, 2026 19:04

@sisuresh sisuresh left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Three findings inline. The mission hang in job_monitor.py is the important one. The other two are small.


Generated by Claude Code

candidates = sorted(w for w in (f"{WORKER_PREFIX}-{i}" for i in range(WORKER_COUNT))
if w not in busy and w not in retiring)
outstanding = queue_remain_count + queue_in_progress_count
keep = max(outstanding, MIN_UNMARKED_WORKERS)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the unmarked reserve (as few as MIN_UNMARKED_WORKERS) is lost to eviction or node failure, the mission hangs instead of failing. Bare pods aren't recreated, Phase 3 requeues their jobs so in_progress empties, and the driver's stall check only fires when in_progress > 0. That leaves queued jobs no pod can claim, and the monitor is still reachable, so no timeout trips. Before this PR, the StatefulSet would have recreated the pods. Suggest failing in the driver when queue_remain_count > 0 and no live unmarked worker remains.


Generated by Claude Code


while true; do
# Stop claiming once the job monitor marks us, so the driver can remove us without interrupting a range.
if [ "$(redis-cli -h "$REDIS_HOST" -p "$REDIS_PORT" SISMEMBER "retiring" "$POD_NAME")" = "1" ]; then

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This fails open. Any redis-cli error yields something other than "1", so a marked worker can still claim a range, and the driver may then delete it mid-job based on a stale status. Suggest claiming only on an explicit 0 and sleeping otherwise. Minor: retiring is hardcoded here, while the monitor reads it from RETIRING.


Generated by Claude Code

apiRateLimit
"Pod"
(fun () ->
podItems

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: podItems is listed before the helm uninstall step, so every worker pod of an uninstalled release is deleted a second time here, with a rate-limit sleep per call and a warning on each 404. Re-listing pods inside this lambda avoids that.


Generated by Claude Code

@sisuresh sisuresh left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One more finding inline, on the scope of the new pod sweep.


Generated by Claude Code

"Pod"
(fun () ->
podItems
|> Seq.filter (fun p -> isNull p.Metadata.OwnerReferences || p.Metadata.OwnerReferences.Count = 0)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This deletes every ownerless pod in the namespace that is older than 2 days, not just PCv2 workers. The sweep runs at every mission start, so any kubectl run or debug pod, or another tool's bare pod, in a shared namespace gets killed, even while running. That includes default when --namespace isn't passed. Suggest applying the same parallel-catchup-* / -stellar-core name filter the helm step uses, or matching on the worker app label.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants