chore(docker): add Update-Codeman.sh for scripted major-update rebuilds - #465
Conversation
docker/README.md and docs/docker-self-update.md both already point operators at "stop the stack, rebuild, restart" for anything the in-app updater refuses to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a new required .env key) — but that was a manual, hand-typed procedure with no script of its own, unlike every other start/update path this deployment has. docker/Update-Codeman.sh scripts it: `docker compose down`, then an unconditional `docker compose build --no-cache` (a major update should be certain of what actually ships, not reuse whatever layers happened to be cached), then hands off to the existing Start-Codeman.sh for the same careful PUID/PGID, override-file and fingerprint handling every other start already goes through — rather than reimplementing any of that by hand and risking it drifting out of step. An optional --volumes/-v flag also removes the codeman-node-modules/ codeman-dist named volumes, the scripted form of the "Resetting the build artefacts" procedure docs/docker-self-update.md already documents by hand. Safe: those two are the only named volumes this stack declares; application data and case workspaces are host bind mounts, never touched by `docker compose down` either way. Docs updated: a "Major updates" section in docker/README.md, and a pointer from docs/docker-self-update.md's existing "Resetting the build artefacts" troubleshooting entry. Tests: extended test/docker-entrypoint.test.ts (the existing home for Start-Codeman.sh's own static checks) with a bash -n parse check, the down-before-build-before-handoff ordering, the --volumes flag's effect, unrecognised-argument handling, and byte-for-byte agreement with Start-Codeman.sh's own override-file resolution logic (so `down` here and `up` there can never target different Compose files). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
…an.sh Start-Codeman.sh, its sibling and the script it hands off to, is itself committed non-executable (100644) upstream — it's documented and invoked as `bash docker/Start-Codeman.sh`, never `./docker/Start-Codeman.sh`. The "is executable" test I'd added for Update-Codeman.sh asserted the opposite convention, which the file correctly does not follow. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
|
Thanks for this. Scripting the major-update path has been an obvious gap for a while, and the write-up in the script header plus the override-file byte-identity test are exactly the kind of care this deployment needs. Three things before merge. Blocker: the handoff fails on every checkout. Build before taking the stack down. The default path can throw the rebuild away. Smaller things, fine to fold into the same push or leave to me at merge:
Everything else checks out here: typecheck, lint, format, frontend syntax and the full |
… handoff, the build/down ordering, and default-clear the build volumes Three blockers, all fixed and verified by actually running the script (not just string-matching it): 1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every checkout, since Start-Codeman.sh is committed non-executable (100644) — the same fact my own second commit on this branch established. Fixed to `exec bash "$script_dir/Start-Codeman.sh"`. 2. `down` ran before `build --no-cache`, so Codeman and every session it was running were offline for the entire rebuild, and a build failure left the stack down with nothing to bring it back — the exact ordering mistake Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists to prevent. Reordered to build, then down, then hand off. 3. The default path could throw the rebuild away: codeman-node-modules/ codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh only clears them when it detects the checkout's HEAD or package-lock.json moved, and neither condition is true for the Dockerfile-only change this script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an image whose fresh node_modules/dist then sat unused behind the old volumes. Made clearing them the default; `--keep-volumes` opts out (replaces the old `--volumes`/`-v` flag, which is no longer needed since clearing is now the default). Smaller items from the same review, also fixed: - The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's owner first, via the identical owner_of() helper Start-Codeman.sh uses (parity-tested) — without it, the build used Compose's default 1000:1000 regardless of the real appdata owner (99:100 on the Unraid layout docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd build during the handoff would then rebuild those layers anyway, so the --no-cache image never actually shipped. - docker/README.md's "rebuilds ... only when it detects ... moved" wrongly described BOTH the rebuild and the volume-clearing as conditional; Start-Codeman.sh rebuilds on every start, only the volume-clearing is conditional. Corrected, and reworded around the new default. - --help/-h now prints usage and exits 0 instead of falling into the unrecognised-argument branch. - "the ONLY named volumes this stack declares" now says docker-compose.yaml specifically, since a docker-compose.override.yml could add more. New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(), --help handling, and — the one that actually catches blocker #1, which five source-string-matching tests did not — a real end-to-end smoke test: a synthetic deployment, a stub `docker` on PATH logging every invocation, the real script executed via a real subprocess. Confirms the real command sequence (build --no-cache, then down --volumes or plain down, then evidence the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails honestly at Start-Codeman.sh's own later check rather than with EACCES. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
|
@Ark0N Thanks for reproducing it — all fixed in 5b4878d, and I verified the fix the same way you found the bug: ran the real script against a synthetic deployment with a stub
Also fixed all four smaller items: PUID/PGID now derives from Per your note about the five tests being pure string matches: added a real end-to-end smoke test — synthetic deployment, stub typecheck/format clean, bash -n and a real bash:3.2 container both parse it, and the full 🤖 Generated with Claude Code |
…pdate-Codeman.sh Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME Compose project as any other checkout on the host. It has to live here too, not just in Start-Codeman.sh: this script's own --no-cache build and its `down`/`down --volumes` both run BEFORE the handoff at the bottom of the file, so Start-Codeman.sh's copy of the guard would only fire after this script's own destructive calls already ran — and its default `down --volumes` is more destructive than Start-Codeman.sh's own targeted refresh, clearing every named volume the resolved project has. Also fixes a real bug the same guard shipped with: under `set -o pipefail`, `grep -v` legitimately exits 1 when nothing survives the filter (the ordinary, no-collision case), and without `|| true` on the pipeline that non-zero status propagates through the command substitution and `set -e` aborts the WHOLE script at the guard — every time, collision or not. Caught only by actually executing the guard end-to-end against a stub `docker` (the existing smoke-test harness), never by a static text/regex check on the source; the stub's `config --format json` response was also fixed to pretty-print like real Compose does, since a compact one-liner silently resolved project_name to empty and exercised neither script's guard the way production output does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose label after a plain `down`, instead of `down --volumes` (which also takes any volume an override file declares while the message named two). `down --volumes` remains only as a warned fallback when the project name cannot be resolved. - Report a failing first `docker compose config --format json` call with a clear error instead of exiting silently under `set -e`. - Filter empty label lines in the collision guard so an unlabelled container cannot hide a real collision; name the moved-checkout exit in its error. - Comments no longer cite a guard or incident in Start-Codeman.sh that does not exist; the README states the real gap (a Node base-image bump leaves codeman-node-modules stale because the lockfile did not move). - docs: Update-Codeman.sh in the docker-self-update.md short-version table and a mention in docker-compose.md; "Major updates" moved under "Updating" in docker/README.md. - test: smoke test covers the new sequence, the config failure and the empty-line case; quiet stdio; @fileoverview names the fourth concern. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Merged, thanks @opticon454! This ships in 1.32.1. The build-before-down ordering and the I applied the rest at merge rather than doing another round:
Two of the fixes were to your own reasoning, so you should know which bits did not hold. The comments cited a collision guard and an incident in |
Summary
docker/README.mdanddocs/docker-self-update.mdboth already tell operators to stop the stack, rebuild, and restart for anything the in-app updater itself refuses to apply — a changedserver.Dockerfile, a changeddocker-compose.yaml, or a new required.envkey — but that was a manual, hand-typed procedure with no script of its own, unlike every other start/update path this deployment has.docker/Update-Codeman.sh:docker compose down, then an unconditionaldocker compose build --no-cache(a major update should be certain of what actually ships, not reuse whatever layers happened to be cached), then hands off to the existingStart-Codeman.shfor the same careful PUID/PGID, override-file, and fingerprint handling every other start already goes through — rather than reimplementing any of that and risking drift.--volumes/-vflag also removes thecodeman-node-modules/codeman-distnamed volumes — the scripted form of the "Resetting the build artefacts" proceduredocs/docker-self-update.mdalready documents by hand. Safe: those two are the only named volumes this stack declares; application data and case workspaces are host bind mounts, never touched bydocker compose downeither way.docker/README.md, and a pointer fromdocs/docker-self-update.md's existing "Resetting the build artefacts" entry.Test plan
bash -n docker/Update-Codeman.sh(also verified against a realbash:3.2container, matching this repo's own bash-3.2-compatibility bar for its other install scripts)test/docker-entrypoint.test.ts(the existing home forStart-Codeman.sh's own static checks): parse check, down-before-build-before-handoff ordering, the--volumesflag's effect (and that the default path stays a plaindown), unrecognised-argument handling, and byte-for-byte agreement withStart-Codeman.sh's own override-file resolution logic (sodownhere andupthere can never target different Compose files)npm run typecheck— cleannpx eslint --config config/eslint.config.js "src/**/*.ts"— clean (nosrc/changes, ran for completeness)npx prettier --checkon the changed.tsfile — clean (docker/README.md/docs/docker-self-update.mdare markdown, outside this repo's ownnpm run formatglob — hand-formatted to match each file's existing style instead of running a bareprettier --write, which reflowed unrelated pre-existing content the first time I tried it)Start-Codeman.sh/cap_adddescribe blocks, reproduced identically on a cleanupstream/mastercheckout — a CRLF-on-Windows artifact from this sandbox'score.autocrlf=true, unrelated to this change and invisible on the Linux CI runner