Repository navigation
fix(worker): serve the healthcheck from the worker instead of resolving pg in the container - #192
Conversation
…ng `pg` in the container
The healthcheck added in v2.8.0 was broken on arrival. Caught by reading
`docker inspect .State.Health.Log` after the deploy:
Error: Cannot find module 'pg'
Require stack: - /app/apps/worker/[eval]
Under pnpm's isolated node_modules, `pg` is not resolvable from
/app/apps/worker even though pg-boss depends on it — so the probe failed
every interval and the container sat in `health: starting` heading for
`unhealthy`. The worker itself was fine throughout; only the probe was
wrong. A permanently-unhealthy container is worse than no healthcheck: it
trains operators to ignore health status.
Fix: the worker now serves its own liveness endpoint on 127.0.0.1:9091
(WORKER_HEALTH_PORT), and the compose probe is a plain Node `fetch` —
which needs NOTHING from node_modules, removing the module-resolution
question entirely rather than working around it.
The endpoint answers only after `boss.getQueue()` round-trips to the
pgboss schema, so a green check means "this process can still reach its
queue", not merely "the event loop is alive" — which is the property the
healthcheck exists to assert.
Bound to 127.0.0.1 (not 0.0.0.0) so it is reachable from the container's
own probe and nothing else; `.unref()`ed so it never holds the process
open; closed on shutdown alongside the pg-boss drain.
Note on method: `getQueueSize` was my first choice and does not exist in
pg-boss v12 — `getQueue(name)` / `getQueues()` do. Checked against the
package's own dist/index.d.ts before writing it, which is the third time
this session that reading the installed types caught an API I had
remembered wrong. The two I did NOT check first each cost a red CI round.
Test plan:
- [x] compose parses; healthcheck test renders as expected.
- [x] tsc on apps/worker/src/index.ts: no syntax errors.
- [x] Probe logic exercised locally against a stub server: 200 → exit 0,
404 → exit 1.
- [x] `getQueue` verified present in pg-boss@12.18.2's shipped types.
- [ ] CI: typecheck/test/build.
- [ ] Post-deploy: `docker inspect` must report worker `healthy`, which is
the check that failed last time. Compose healthchecks are exercised
by no CI gate, so this one is only provable in prod.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDtsHQ5kq3HHxswr2o2s2Q
📝 WalkthroughWalkthroughThe worker now serves a local ChangesWorker liveness
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant DockerHealthcheck
participant WorkerHealthServer
participant PgBoss
DockerHealthcheck->>WorkerHealthServer: GET /health
WorkerHealthServer->>PgBoss: boss.getQueue("kea.extract")
PgBoss-->>WorkerHealthServer: Queue result or error
WorkerHealthServer-->>DockerHealthcheck: HTTP 200 or 503 JSON response
Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@deploy/docker-compose.yml`:
- Around line 251-255: Add WORKER_HEALTH_PORT to the worker service’s
environment block so the healthcheck command can read the configured custom port
inside the container. Reuse the existing Compose variable-substitution
convention and preserve the default port behavior when no value is configured.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: 691b2d16-8386-4738-bd04-65cc2c4ae0b5
📒 Files selected for processing (3)
apps/worker/src/index.tsdeploy/docker-compose.ymldocs/KNOWN_ISSUES.md
| healthcheck: | ||
| test: | ||
| - CMD-SHELL | ||
| - >- | ||
| node -e "const{Client}=require('pg');const c=new Client({connectionString:process.env.DATABASE_URL});c.connect().then(()=>c.query('select 1 from '+(process.env.PG_BOSS_SCHEMA||'pgboss')+'.version limit 1')).then(()=>{c.end();process.exit(0)}).catch(()=>process.exit(1))" | ||
| node -e "fetch('http://127.0.0.1:'+(process.env.WORKER_HEALTH_PORT||9091)+'/health').then(r=>process.exit(r.ok?0:1)).catch(()=>process.exit(1))" |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Forward WORKER_HEALTH_PORT to the worker container.
Line 255 reads WORKER_HEALTH_PORT, but the worker.environment block does not define it. Docker Compose does not forward arbitrary host environment variables. A configured custom port therefore has no effect.
Proposed fix
environment:
DATABASE_URL: ${DATABASE_URL}
+ WORKER_HEALTH_PORT: ${WORKER_HEALTH_PORT:-9091}
NODE_ENV: ${NODE_ENV:-production}🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@deploy/docker-compose.yml` around lines 251 - 255, Add WORKER_HEALTH_PORT to
the worker service’s environment block so the healthcheck command can read the
configured custom port inside the container. Reuse the existing Compose
variable-substitution convention and preserve the default port behavior when no
value is configured.
The healthcheck added in v2.8.0 was broken on arrival — caught by
docker inspect .State.Health.Logafter the deploy, not by CI (compose healthchecks are exercised by no gate).Under pnpm's isolated
node_modules,pgisn't resolvable from/app/apps/workereven though pg-boss depends on it. The probe failed every interval; the worker itself was fine. A permanently-unhealthy container is worse than no healthcheck — it trains operators to ignore health status.Fix
The worker serves its own liveness endpoint on
127.0.0.1:9091, probed with Node's built-infetch— which needs nothing fromnode_modules, removing the resolution question rather than working around it.The endpoint answers only after
boss.getQueue()round-trips to the pgboss schema, so green means "this process can still reach its queue", not merely "the event loop is alive".Method note
getQueueSizewas my first choice and doesn't exist in pg-boss v12. I checkeddist/index.d.tsbefore writing it — the third time this session that reading the installed types caught a misremembered API. The two I didn't check first each cost a red CI round.Test plan
tscon the worker: no syntax errorsgetQueueverified present in pg-boss@12.18.2's shipped typesdocker inspectreports workerhealthy— only provable in prod🤖 Generated with Claude Code
https://claude.ai/code/session_01GDtsHQ5kq3HHxswr2o2s2Q
Summary by CodeRabbit
New Features
Bug Fixes