flake.nix's rainix-sol-artifacts wraps the deploy in until do_deploy; do sleep 5 — ten attempts against the same endpoint before failing anyway.
That is structurally identical to Foundry's --fork-retries, and useless for the same reason #288 gives: retrying one URL cannot recover a deterministic failure. An exhausted plan quota, a dead host, or an empty --rpc-url fails identically on attempt 10 as on attempt 1. The only thing the loop buys is ~50 seconds of latency per failed run.
Visible in the current main failure, where all ten attempts produce the same line:
++ forge script script/Deploy.sol:Deploy -vvvvv --slow --rpc-url '' --etherscan-api-key *** --resume
Error: Internal transport error: Connection refused (os error 111)
Deploy failed, retrying in 5 seconds... (attempt 7)
...
Deploy failed after 10 attempts, aborting.
The fix is the one #289 already applied one level up
Classify the error once and either fail fast or move to another candidate. #289 built exactly that for fork reads — a candidate pool plus a typed rejection reason per endpoint — and the deploy task is the same problem with the same shape.
Retrying is only ever right for a transient failure, and nothing here distinguishes transient from deterministic. Either make that distinction or drop the loop; ten identical attempts is the worst of both.
Relationship to the other follow-up
This overlaps the sepolia work, since sepolia is the network this task deploys to and an empty --rpc-url is the failure the loop is currently spinning on. They can be done independently — dropping the useless retry does not require deciding sepolia's naming — but doing sepolia first makes this one's effect visible.
flake.nix'srainix-sol-artifactswraps the deploy inuntil do_deploy; do sleep 5— ten attempts against the same endpoint before failing anyway.That is structurally identical to Foundry's
--fork-retries, and useless for the same reason #288 gives: retrying one URL cannot recover a deterministic failure. An exhausted plan quota, a dead host, or an empty--rpc-urlfails identically on attempt 10 as on attempt 1. The only thing the loop buys is ~50 seconds of latency per failed run.Visible in the current
mainfailure, where all ten attempts produce the same line:The fix is the one #289 already applied one level up
Classify the error once and either fail fast or move to another candidate. #289 built exactly that for fork reads — a candidate pool plus a typed rejection reason per endpoint — and the deploy task is the same problem with the same shape.
Retrying is only ever right for a transient failure, and nothing here distinguishes transient from deterministic. Either make that distinction or drop the loop; ten identical attempts is the worst of both.
Relationship to the other follow-up
This overlaps the sepolia work, since sepolia is the network this task deploys to and an empty
--rpc-urlis the failure the loop is currently spinning on. They can be done independently — dropping the useless retry does not require deciding sepolia's naming — but doing sepolia first makes this one's effect visible.