fix: Replace DB connection crash with passive reconnect polling - #3155
Conversation
Two related changes to make startup resilient to a brief DB outage: 1. database/manager.py: reduce initial background-reconnect delay from 30s to 5s. Previously, inside the 120s startup wait only 2 reconnect attempts fit (t=30, t=75). With 5s first delay we get several early attempts where recovery is most likely. 2. main.py: after the 120s window, don't SystemExit(1). Enter a passive wait loop (poll every DB_RECONNECT_POLL_INTERVAL, default 30s) until the background task reconnects. Avoids CrashLoop + Sentry storms on transient outages (e.g. the fatal events on 2026-04-03). https://claude.ai/code/session_01Dgaz4FwhYhLtERMozH2UZF
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
🧯 Dangerous deletes guard reportPolicy: see .cursorrules — dangerous deletions are blocked unless wrapped safely. Summary:
Flagged findings (file:line:snippet): Excluded matches (by path pattern) |
⏱️ Performance report(No performance test durations collected. Mark tests with |
📖 Documentation PreviewThe documentation has been built successfully!
To view locally:
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Follow-up to df08735. Cursor-bot flagged that main.py's new passive wait loop could hang forever: after max_bg_attempts (10) in _schedule_background_reconnect, no more attempts were scheduled, so is_connected would stay False and the main loop would busy-wait doing nothing useful. Fix: remove the hard cap on background reconnect attempts. Exponential backoff is already capped at 300s, so worst-case load is one attempt per 5 minutes — cheap to keep trying. We emit a one-time "escalating" event when the historical max is crossed, so observability is preserved. https://claude.ai/code/session_01Dgaz4FwhYhLtERMozH2UZF
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 158c732. Configure here.
Cursor-bot flagged that the new DB_RECONNECT_POLL_INTERVAL (introduced in df08735) wasn't registered in services/config_inspector_service.py or docs/environment-variables.rst. Also adds the previously-undocumented DB_RECONNECT_WAIT_BEFORE_POLL so both tunables are visible to operators using the config inspector. https://claude.ai/code/session_01Dgaz4FwhYhLtERMozH2UZF

✨ תיאור קצר
שינוי התנהגות כשל התחברות למסד הנתונים: במקום לעצור את התהליך (SystemExit) לאחר timeout, התהליך יישאר פעיל ויחכה בפולינג פסיבי עד שהחיבור יתחדש. זה מונע CrashLoop ו-Sentry storms בעת הפסקות זמניות של מסד הנתונים.
📦 שינויים עיקריים
פירוט:
main.py: החלפת
SystemExit(1)בלולאת פולינג פסיבית שמחכה לחיבור מחדשDB_RECONNECT_POLL_INTERVAL(ברירת מחדל: 30 שניות)criticalל-warningכדי לא להעלות אזעקות מיותרותdatabase/manager.py: הקטנת
delayבפונקציית_schedule_background_reconnectמ-30 שניות ל-5 שניות🧪 בדיקות
📝 סוג שינוי
✅ צ'קליסט
docs/environment-variables.rstוגםservices/config_inspector_service.pyDB_RECONNECT_POLL_INTERVAL– יש לעדכן את התיעוד🧩 השפעות/סיכונים
🧯 סיכון / החזרה
https://claude.ai/code/session_01Dgaz4FwhYhLtERMozH2UZF