Skip to content

fix(nightly): stop retrying an input that already failed - #93

Open
saphid wants to merge 1 commit into
mainfrom
fix/fork-nightly-skip-failed-inputs
Open

saphid wants to merge 1 commit into
mainfrom
fix/fork-nightly-skip-failed-inputs

Conversation

@saphid

@saphid saphid commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Fork Nightly runs every five minutes and on each upstream Nightly dispatch, which GitHub delivers in pairs. The Orchestrator v2 stream calls the same reusable workflow on its own five-minute schedule. When the patch stack conflicts with upstream, no release is published, so resolve keeps asking for a build and every run retries the identical failing input. Over the last 30 days that produced 391 failed Fork Nightly runs and 111 failed v2 runs, one failure email each.

A failed build now saves a small actions/cache marker keyed by release channel and stack fingerprint (record_failure). prepare looks that key up before building. If it exists, scheduled and dispatch runs log a notice and skip the build. The run is green and no email goes out.

  • A new upstream Nightly tag or v2 commit, or any manifest change, changes the fingerprint and builds normally. That first failure still emails once.
  • Manual workflow_dispatch runs, including v2's, skip the lookup, so they always retry. force_rebuild behaves as before.
  • Markers expire through normal cache eviction (7 days without access), so a stale stack gets retried about once a week.

DOWNSTREAM_NIGHTLY.md documents the behavior.

This does not repair the stacks themselves. Fork Nightly's first patch (d58c3273, natural-language thread search) conflicts with v0.0.43-nightly.20260929.2428, and several later entries also need refreshing. The v2 stack conflicts from 75be8c11 onward. See the separate v2 manifest PR for its 422 failure.

Verification: node --test .github/scripts/downstream-nightly.test.mjs (22 pass), actionlint clean, vp fmt --check clean. There is no end-to-end run yet. The first scheduled run after merge should fail once and record the marker, and later runs of the same input should skip.

Independent review: T3 delegated task on Codex gpt-6-sol (high) found no actionable findings. It confirmed that failed-job outputs propagate, that the cache path and scope match across callers, that manual dispatch under workflow_call bypasses the lookup, and the failure() semantics.

Work by Claude Opus 5.5 (1M context) in Claude Code, running in T3 Code.

🤖 Generated with Claude Code


Update: to stop the failure emails immediately, both Fork Nightly and Fork Nightly Orchestrator v2 are now disabled on saphid/t3code (gh workflow disable). After merging #93 (and #94), re-enable them with:

gh workflow enable 347086910 -R saphid/t3code   # Fork Nightly
gh workflow enable 352233980 -R saphid/t3code   # Fork Nightly Orchestrator v2

Fork Nightly and its Orchestrator v2 caller run every five minutes and on
each upstream dispatch. A conflicting patch stack never produces a release,
so every run rebuilt the same failing input: about 500 failure emails a
month. A failed build now saves a cache marker keyed by channel and stack
fingerprint, and later scheduled or dispatch runs skip that input with a
notice. New upstream sources, manifest changes, and manual runs still build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown

Thread transfer impact

⚠️ The latest CI run did not produce a thread transfer result for 48be308.

This comment will update automatically after the next completed run.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant