Repository navigation
fix(execute): propagate agent 4xx in async lane instead of blanket 502 - #945
Merged
santoshkumarradha merged 1 commit intoAug 23, 2026
Conversation
Agent-Field#862) The async-completion branch in handleSync hardcoded HTTP 502 for all failed executions, making client-input rejections (e.g. 422) indistinguishable from upstream outages. The sync lane correctly propagated the agent's 4xx via writeExecutionError, but the async lane — which reads the execution from DB after the agent calls back — had no way to recover the original HTTP status. Root cause: the callError.statusCode is lost once the execution is persisted; the async lane only sees the stored record. Fix (two-pronged): Control plane: - Add error_status_code field to executionStatusUpdateRequest so the SDK can forward the HTTP status in its failure callback. - When error_status_code is 4xx, encode it in StatusReason as "agent_client_error:<code>" for persistence. - Add httpStatusForFailedExecution() helper that resolves the HTTP status from StatusReason (encoded client errors, known categories) and ErrorMessage ("agent error (NNN):" pattern fallback). - Replace hardcoded 502 in the async-completion branch with the helper. Python SDK: - In _execute_async_with_callback failure path, propagate error_status_code from exception.status_code / exception.code when the value is a valid HTTP status (400-599). Backward-compatible: SDKs that don't send error_status_code continue to get the existing behavior (502 default). SDKs that do send it get correct 4xx propagation immediately.
7vignesh
force-pushed
the
fix/862-async-502-for-client-errors
branch
from
August 22, 2026 21:25
5e04137 to
4b44d05
Compare
Contributor
Performance
✓ No regressions detected |
Contributor
📊 Coverage gateThresholds from
✅ Gate passedNo surface regressed past the allowed threshold and the aggregate stayed above the floor. |
Contributor
📐 Patch coverage gateThreshold: 80% on lines this PR touches vs
✅ Patch gate passedEvery surface whose lines were touched by this PR has patch coverage at or above the threshold. |
santoshkumarradha
approved these changes
Aug 23, 2026
santoshkumarradha
left a comment
Member
There was a problem hiding this comment.
Validated the control-plane and Python surfaces locally. The status propagation fix and regression coverage look good to me.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #862. The async-completion branch in
handleSynchardcoded HTTP 502 for all failed executions, making client-input rejections (e.g. 422 from input validation) indistinguishable from upstream outages. A reasoner validating its own input would returninvalid_input: ...but the caller saw 502 Bad Gateway — sending operators to check pods instead of the request.Root cause: The
callError.statusCodeis lost once the execution is persisted; the async lane (which reads the execution from DB after the agent calls back) had no way to recover the original HTTP status.Fix (two-pronged):
Control plane
error_status_codefield toexecutionStatusUpdateRequestso the SDK can forward the HTTP status in its failure callbackerror_status_codeis 4xx, encode it inStatusReasonasagent_client_error:<code>for persistencehttpStatusForFailedExecution()helper that resolves the HTTP status from:StatusReason— encoded client errors and known server-side categories (agent_timeout→504,target_not_found→404, etc.)ErrorMessage— parses"agent error (NNN):"pattern as fallback for existing executionsPython SDK
_execute_async_with_callbackfailure path, propagateerror_status_codefromexception.status_code/exception.codewhen the value is a valid HTTP status (400–599)Backward-compatible: SDKs that don't send
error_status_codeget existing behavior (502 default). SDKs that do send it get correct 4xx propagation immediately.Type of change
Test plan
cd control-plane && go test ./internal/handlers/ -run TestHttpStatusForFailedExecution -v(12 cases: encoded 4xx, category mapping, error message parsing, defaults)cd control-plane && go test ./internal/handlers/ -run TestUpdateExecutionStatusHandler_ErrorStatusCode -vcd control-plane && go test ./internal/handlers/ -run TestExecuteHandler_AgentError -v(existing 500→502 preserved)cd control-plane && go test ./internal/handlers/ -timeout 120s(full suite)cd control-plane && go vet ./internal/handlers/...cd control-plane && golangci-lint run ./internal/handlers/(no new warnings)cd sdk/python && python -m pytest tests/test_usage_transport.py(async callback path)python -c "import agentfield.agent"(syntax OK)Test coverage
coverage-baseline.jsonin this PR only if the removal caused a legitimate regression and I called it out in the summary above.Checklist
Related issues / PRs
Fixes #862
Note:
TestCallAgent_ErrorResponseis a pre-existing flaky test (timing-sensitiveelapsed > 0assertion) — confirmed by running it with-count=5on both main and this branch.