Summary
The per-team sandbox index (sandbox:storage:{teamID}:index, a Redis SET) accumulates stale entries that are never cleaned up when sandbox keys disappear outside the normal Remove() path.
Root cause
ExpiredItems() sweeps orphaned ZSET members (entries in sandbox:storage:global:expiration whose corresponding sandbox key is gone), but only removes the ZSET member — it does not remove the sandboxID from the team index SET:
// items.go — current behaviour
if raw == nil {
staleMembers = append(staleMembers, ref.member) // ZREM the ZSET member ✓
orphanCount++
continue // team index entry is NOT cleaned ✗
}
No other code path cleans the team index for absent sandbox keys:
TeamItems / fetchSandboxBatch silently skip nil MGET results with a continue
- The healer (
healExpirationIndex) only fills missing ZSET members, never prunes the team index
Remove() / removeSandboxScript cleans both atomically, but only runs on the normal removal path
The team index SET therefore grows without bound whenever sandbox keys disappear externally — most commonly via Redis maxmemory key eviction or a Redis restart with incomplete persistence.
Observed impact
Teams with high sandbox churn (frequent create / pause / resume cycles) combined with Redis key eviction can accumulate thousands of stale SET members, inflating Redis memory and slowing TeamItems (SMEMBERS + MGET fan-out is proportional to SET size, including stale entries).
Trigger conditions
- Redis
maxmemory policy that can evict individual sandbox keys (allkeys-lru, allkeys-random, etc.)
- Redis restart with RDB/AOF disabled or incomplete
- Any path that deletes a sandbox key without calling
sandboxStore.Remove()
Frequent autoPause / autoResume amplifies the issue by increasing per-team sandbox churn and Redis memory pressure, raising the probability of key eviction.
Fix
When ExpiredItems finds an orphaned ZSET member (MGET returned nil), also SREM the sandboxID from the team index. MGET already confirmed the sandbox key is absent, so we cannot unindex a live sandbox. A concurrent Add that races the SREM will SADD the sandboxID back immediately, so the worst outcome is a brief gap in TeamItems results for one eviction cycle.
Summary
The per-team sandbox index (
sandbox:storage:{teamID}:index, a Redis SET) accumulates stale entries that are never cleaned up when sandbox keys disappear outside the normalRemove()path.Root cause
ExpiredItems()sweeps orphaned ZSET members (entries insandbox:storage:global:expirationwhose corresponding sandbox key is gone), but only removes the ZSET member — it does not remove the sandboxID from the team index SET:No other code path cleans the team index for absent sandbox keys:
TeamItems/fetchSandboxBatchsilently skipnilMGET results with acontinuehealExpirationIndex) only fills missing ZSET members, never prunes the team indexRemove()/removeSandboxScriptcleans both atomically, but only runs on the normal removal pathThe team index SET therefore grows without bound whenever sandbox keys disappear externally — most commonly via Redis
maxmemorykey eviction or a Redis restart with incomplete persistence.Observed impact
Teams with high sandbox churn (frequent create / pause / resume cycles) combined with Redis key eviction can accumulate thousands of stale SET members, inflating Redis memory and slowing
TeamItems(SMEMBERS + MGET fan-out is proportional to SET size, including stale entries).Trigger conditions
maxmemorypolicy that can evict individual sandbox keys (allkeys-lru,allkeys-random, etc.)sandboxStore.Remove()Frequent
autoPause/autoResumeamplifies the issue by increasing per-team sandbox churn and Redis memory pressure, raising the probability of key eviction.Fix
When
ExpiredItemsfinds an orphaned ZSET member (MGET returned nil), alsoSREMthe sandboxID from the team index. MGET already confirmed the sandbox key is absent, so we cannot unindex a live sandbox. A concurrentAddthat races the SREM willSADDthe sandboxID back immediately, so the worst outcome is a brief gap inTeamItemsresults for one eviction cycle.