Skip to content

fix(worker): serve the healthcheck from the worker instead of resolving pg in the container - #192

Merged
bejranonda merged 1 commit into
mainfrom
bugfix/worker-healthcheck-probe
Aug 3, 2026
Merged

bejranonda merged 1 commit into
mainfrom
bugfix/worker-healthcheck-probe

Conversation

@bejranonda

@bejranonda bejranonda commented Aug 3, 2026 •

Copy link
Copy Markdown
Owner

The healthcheck added in v2.8.0 was broken on arrival — caught by docker inspect .State.Health.Log after the deploy, not by CI (compose healthchecks are exercised by no gate).

Error: Cannot find module 'pg'
Require stack: - /app/apps/worker/[eval]

Under pnpm's isolated node_modules, pg isn't resolvable from /app/apps/worker even though pg-boss depends on it. The probe failed every interval; the worker itself was fine. A permanently-unhealthy container is worse than no healthcheck — it trains operators to ignore health status.

Fix

The worker serves its own liveness endpoint on 127.0.0.1:9091, probed with Node's built-in fetch — which needs nothing from node_modules, removing the resolution question rather than working around it.

The endpoint answers only after boss.getQueue() round-trips to the pgboss schema, so green means "this process can still reach its queue", not merely "the event loop is alive".

Method note

getQueueSize was my first choice and doesn't exist in pg-boss v12. I checked dist/index.d.ts before writing it — the third time this session that reading the installed types caught a misremembered API. The two I didn't check first each cost a red CI round.

Test plan

  • compose parses; healthcheck test renders correctly
  • tsc on the worker: no syntax errors
  • probe logic exercised against a stub server: 200 → exit 0, 404 → exit 1
  • getQueue verified present in pg-boss@12.18.2's shipped types
  • CI
  • Post-deploy docker inspect reports worker healthy — only provable in prod

🤖 Generated with Claude Code

https://claude.ai/code/session_01GDtsHQ5kq3HHxswr2o2s2Q

Summary by CodeRabbit

  • New Features

    • Added a worker health endpoint that reports whether the service and job queue are available.
    • Returns clear success or failure status codes for monitoring systems.
  • Bug Fixes

    • Updated container health monitoring to use the worker’s health endpoint, improving reliability.
    • Corrected the documented healthcheck behavior and removed the faulty database-based probe.

…ng `pg` in the container

The healthcheck added in v2.8.0 was broken on arrival. Caught by reading
`docker inspect .State.Health.Log` after the deploy:

  Error: Cannot find module 'pg'
  Require stack: - /app/apps/worker/[eval]

Under pnpm's isolated node_modules, `pg` is not resolvable from
/app/apps/worker even though pg-boss depends on it — so the probe failed
every interval and the container sat in `health: starting` heading for
`unhealthy`. The worker itself was fine throughout; only the probe was
wrong. A permanently-unhealthy container is worse than no healthcheck: it
trains operators to ignore health status.

Fix: the worker now serves its own liveness endpoint on 127.0.0.1:9091
(WORKER_HEALTH_PORT), and the compose probe is a plain Node `fetch` —
which needs NOTHING from node_modules, removing the module-resolution
question entirely rather than working around it.

The endpoint answers only after `boss.getQueue()` round-trips to the
pgboss schema, so a green check means "this process can still reach its
queue", not merely "the event loop is alive" — which is the property the
healthcheck exists to assert.

Bound to 127.0.0.1 (not 0.0.0.0) so it is reachable from the container's
own probe and nothing else; `.unref()`ed so it never holds the process
open; closed on shutdown alongside the pg-boss drain.

Note on method: `getQueueSize` was my first choice and does not exist in
pg-boss v12 — `getQueue(name)` / `getQueues()` do. Checked against the
package's own dist/index.d.ts before writing it, which is the third time
this session that reading the installed types caught an API I had
remembered wrong. The two I did NOT check first each cost a red CI round.

Test plan:
- [x] compose parses; healthcheck test renders as expected.
- [x] tsc on apps/worker/src/index.ts: no syntax errors.
- [x] Probe logic exercised locally against a stub server: 200 → exit 0,
      404 → exit 1.
- [x] `getQueue` verified present in pg-boss@12.18.2's shipped types.
- [ ] CI: typecheck/test/build.
- [ ] Post-deploy: `docker inspect` must report worker `healthy`, which is
      the check that failed last time. Compose healthchecks are exercised
      by no CI gate, so this one is only provable in prod.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDtsHQ5kq3HHxswr2o2s2Q
@coderabbitai

coderabbitai Bot commented Aug 3, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The worker now serves a local /health endpoint on WORKER_HEALTH_PORT. The endpoint checks pg-boss connectivity. Docker uses this endpoint for healthchecks, and shutdown closes the server before draining jobs.

Changes

Worker liveness

Layer / File(s) Summary
Worker health endpoint
apps/worker/src/index.ts
The worker serves /health on localhost. The endpoint checks the kea.extract queue and returns 200, 503, or 404 responses. Shutdown closes the server before pg-boss drains jobs.
Container healthcheck integration
deploy/docker-compose.yml, docs/KNOWN_ISSUES.md
The Docker healthcheck uses Node fetch against the worker endpoint. The known-issues entry documents the corrected healthcheck flow.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant DockerHealthcheck
  participant WorkerHealthServer
  participant PgBoss
  DockerHealthcheck->>WorkerHealthServer: GET /health
  WorkerHealthServer->>PgBoss: boss.getQueue("kea.extract")
  PgBoss-->>WorkerHealthServer: Queue result or error
  WorkerHealthServer-->>DockerHealthcheck: HTTP 200 or 503 JSON response
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes replacing the container-side pg resolution probe with a worker-served healthcheck.
Description check ✅ Passed The description explains the issue and fix, provides a specific test plan, and clearly marks CI and post-deploy checks as pending.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch bugfix/worker-healthcheck-probe

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@deploy/docker-compose.yml`:
- Around line 251-255: Add WORKER_HEALTH_PORT to the worker service’s
environment block so the healthcheck command can read the configured custom port
inside the container. Reuse the existing Compose variable-substitution
convention and preserve the default port behavior when no value is configured.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 691b2d16-8386-4738-bd04-65cc2c4ae0b5

📥 Commits

Reviewing files that changed from the base of the PR and between e663a25 and 65a6294.

📒 Files selected for processing (3)
  • apps/worker/src/index.ts
  • deploy/docker-compose.yml
  • docs/KNOWN_ISSUES.md

Comment thread deploy/docker-compose.yml
Comment on lines 251 to +255
healthcheck:
test:
- CMD-SHELL
- >-
node -e "const{Client}=require('pg');const c=new Client({connectionString:process.env.DATABASE_URL});c.connect().then(()=>c.query('select 1 from '+(process.env.PG_BOSS_SCHEMA||'pgboss')+'.version limit 1')).then(()=>{c.end();process.exit(0)}).catch(()=>process.exit(1))"
node -e "fetch('http://127.0.0.1:'+(process.env.WORKER_HEALTH_PORT||9091)+'/health').then(r=>process.exit(r.ok?0:1)).catch(()=>process.exit(1))"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Forward WORKER_HEALTH_PORT to the worker container.

Line 255 reads WORKER_HEALTH_PORT, but the worker.environment block does not define it. Docker Compose does not forward arbitrary host environment variables. A configured custom port therefore has no effect.

Proposed fix
     environment:
       DATABASE_URL: ${DATABASE_URL}
+      WORKER_HEALTH_PORT: ${WORKER_HEALTH_PORT:-9091}
       NODE_ENV: ${NODE_ENV:-production}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@deploy/docker-compose.yml` around lines 251 - 255, Add WORKER_HEALTH_PORT to
the worker service’s environment block so the healthcheck command can read the
configured custom port inside the container. Reuse the existing Compose
variable-substitution convention and preserve the default port behavior when no
value is configured.

@bejranonda
bejranonda merged commit eed12d7 into main Aug 3, 2026
7 checks passed
@bejranonda
bejranonda deleted the bugfix/worker-healthcheck-probe branch August 3, 2026 06:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant