Skip to content

fix(core): raise idle timeouts so long turns survive - #237

Open
j35dev wants to merge 1 commit into
mainfrom
fix/harness-idle-timeouts
Open

j35dev wants to merge 1 commit into
mainfrom
fix/harness-idle-timeouts

Conversation

@j35dev

@j35dev j35dev commented Sep 17, 2026 •

Copy link
Copy Markdown
Owner

Task

Harness idle timeouts kill legitimate long work: Ari Core streams die after 120s of silence (endpoint sent nothing for 120s), and ACP prompts stall after 300s while a child is running pnpm verify / tests / builds with no protocol traffic.

What changed

  • Ari Core SSE/NDJSON idle deadline: 2 minutes → 20 minutes.
  • ACP prompt stall watchdog (Claude/Codex/Grok/etc.): 5 minutes → 20 minutes. ARI_ACP_PROMPT_STALL_MS still overrides; 0 still disables.
  • Ari Core bash default/max: 120s/600s → 600s/1800s, so verify-class commands finish without the model having to pass an explicit timeout.

Wedge detection remains: a truly silent endpoint or ACP child still fails legibly instead of hanging forever.

How verified

VITEST_MAX_WORKERS=2 pnpm verify green (typecheck + lint + test). Desktop: 154 passed, 3 skipped (1404 tests). Focused: http-retry.test.ts, tools.test.ts, connection.test.ts.


Devin Review

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

Devin Review

*/
export function acpPromptStallMs(raw = process.env['ARI_ACP_PROMPT_STALL_MS']): number {
if (raw === undefined || raw.trim() === '') return 300_000
if (raw === undefined || raw.trim() === '') return 1_200_000

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Idle sessions reject fresh prompts

After 20 minutes without ACP traffic, prompt can reject a fresh turn on its first watchdog tick. #lastInboundAt still timestamps the previous turn instead of the new request.

Learn more

The stall limit describes silence during one prompt, but the watchdog uses a connection-wide timestamp. A connection can remain open between turns. Starting a prompt does not reset that timestamp, so the first interval callback can see more than 20 minutes of silence even though the request was just sent. The default interval checks every two seconds.

Example: A user finishes one turn at 10:00 and sends another at 10:25. The new agent receives the request, but Ari can reject it around 10:25:02 instead of allowing 20 minutes for a response.

Recommended fix: Give each stalled request a fresh baseline when #request starts. Preserve subsequent updates from inbound traffic, preferably with request-local liveness state so unrelated concurrent requests cannot distort the deadline.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +411 to +412
const BASH_DEFAULT_TIMEOUT_SECONDS = 600
const BASH_MAX_TIMEOUT_SECONDS = 1_800

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Interrupted Bash commands keep running

After an Ari Core turn is interrupted, executeBash keeps its child alive for the new 600-second default. ToolContext carries no AbortSignal, so cancelled commands can keep changing workspace files in the background.

Learn more

The engine calls the adapter's interrupt method immediately, and Ari Core aborts only its model requests and parked UI requests. Tool execution awaits executeBash without passing that signal. Node's execFile therefore retains the shell until it exits or its independent timeout fires. This PR raises that default lifetime from two to ten minutes and the maximum from ten to thirty minutes.

Example: A model starts bash with command: "sleep 300; echo changed > result.txt". The user interrupts after five seconds. The turn becomes interrupted, but the shell writes result.txt five minutes later.

Recommended fix: Add an AbortSignal to ToolContext, pass the turn signal from runAgentLoop, and supply it to execFile. Treat abort termination as an interrupted tool execution and ensure the entire spawned process tree is stopped on every platform.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants