Skip to content

fix(responses): raise failed results with their error code for retry routing - #56

Merged
zhanghanduo merged 1 commit into
ApodexAI:mainfrom
Lucifus613:upstream/responses-error-code-routing
Oct 1, 2026
Merged

zhanghanduo merged 1 commit into
ApodexAI:mainfrom
Lucifus613:upstream/responses-error-code-routing

Conversation

@Lucifus613

Copy link
Copy Markdown
Contributor

Summary

OpenAIResponsesClient mishandles a Responses result whose status is failed:

  • Non-streaming: _parse_responses_output parses a failed result as if it were a normal answer. The partial output is returned to the caller, including any function_call items, so the agent loop can execute tool calls from a response that the server itself reported as failed.
  • Streaming: a response.failed event matches no branch and is silently ignored.
  • Retry routing: even after a failure is raised, a plain LLMError carries no status code. The shared retry classifier therefore cannot tell a rate limit from a permanent 400-class error. Rate limits get generic backoff, and errors such as invalid_prompt are retried pointlessly.

This PR raises on failed results and keeps the provider's error code on the error. The code is mapped to the equivalent HTTP status so that the existing call_llm ladder routes it:

Code Status Routing
rate_limit_exceeded 429 rate-limit backoff
server_error 500 transient backoff
vector_store_timeout 504 transient backoff
other SDK codes (e.g. invalid_prompt, bio_policy, invalid_image*) 400 stops immediately (non_transient)
unknown or missing none unchanged generic retry

The table's key set is exactly the ResponseError.code literal in openai 3.7.0. The existing text-based rules still take precedence: safety refusals advance the fallback chain, and upstream-timeout wording stays transient. For example, image_content_policy_violation advances the chain through the existing content_policy_violation rule.

Changes

  • agent_core/providers/openai_responses.py:
    • A failed result raises _ResponsesError(LLMError) with code and status_code. Its message is "<code>: <provider message>", falling back to Responses request failed.
    • The streaming response.failed event raises the same error.
  • tests/test_responses_error_routing.py (new):
    • Every code, as both a dict and an SDK-shaped object, keeps its code and status.
    • Each failure routes through the real call_llm: reason, call count and backoff band.
    • A streamed response.failed event raises the same error as the parsed result.
  • changes/responses-failed-error-routing.fix.md: changelog fragment.

Test plan

Run on this branch, which is main plus this commit:

  • python -m pytest tests -q: 1715 passed (openai 3.7.0, anthropic 1.3.0).
  • ruff check on the changed files: clean.
  • The new test file copied onto unmodified main: 99 failed. With this commit: 99 passed.
  • Live OpenAI call returning a failed result: not run (paid).

pyright agent_core reports 6 errors, all in agent_core/runtime/loop/agent_loop.py. Unmodified main reports the same 6, and this PR does not touch that file.

🤖 Generated with Claude Code

…routing

A failed Responses result was parsed like a successful one, so its partial
output (including tool calls) was returned as an answer, and a streamed
response.failed event was ignored. Raise instead, keeping the provider code
on the error and mapping it to the equivalent status (rate_limit_exceeded
429, server_error 500, vector_store_timeout 504, other SDK codes 400;
unknown codes keep generic retry) so the shared retry classifier routes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@zhanghanduo zhanghanduo left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fix. Very helpful. Cannot reproduce LLM errors easily. Should be correct.

@zhanghanduo
zhanghanduo merged commit 7b8fae0 into ApodexAI:main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants