fix(responses): report truncation and normalize the translated output shape - #955
Conversation
… shape Truncated turns now return status "incomplete" with incomplete_details on chat-translated and Anthropic providers (stream and non-stream), and native OpenAI's incomplete_details is no longer dropped. Empty answers keep the required output_text "text" member, and the Responses usage object stays in the OpenAI shape; provider extras remain in usage records and cost calculation.
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
|
Warning Review limit reachedNext included review available in 15 minutes. View limit detailsLimit details: You’ve used all 4 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (17)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
Three
/v1/responsesoutput-shape gaps, all visible on translated (non-native-OpenAI) providers.Truncation is reported. A turn stopped by
max_output_tokens(or a content filter) now returnsstatus: "incomplete"withincomplete_details.reason, and the message item carriesstatus: "incomplete"— chat-translated providers (gemini, groq, deepseek, …) and Anthropic, in both stream and non-stream.core.ResponsesResponsegained the typedincomplete_detailsfield, so native OpenAI's value is no longer dropped either (it previously returnedstatus: "incomplete"with no details).Empty answers keep
text.output_textparts always serializetext, which the OpenAI schema requires; an empty assistant answer used to emit{"type":"output_text","annotations":[]}. The streamed path already did this.Usage is normalized. The Responses usage object now carries only the OpenAI shape (
input_tokens/output_tokens/total_tokensplus*_tokens_details), instead of leaking provider members such asthoughts_token_count,completion_reasoning_tokensorcompletion_time. Nothing with an OpenAI-shaped home is lost: reasoning tokens stay underoutput_tokens_details, cached tokens underinput_tokens_details(Anthropic's Responses usage now fills both). Provider extras remain inRawUsage, which feeds usage records and cost calculation, and the Anthropic streamed usage payload keeps its cache counts because stream costs are read back from it.Provider notes: gemini 2.5 flash can return a completely empty stream (no chunk, no
finish_reason, no usage) when the whole budget goes to thinking — nothing to map there; that stream still endscompleted.Tested: unit tests for each path (
internal/core,internal/providers,internal/providers/anthropic),go test ./...(pre-existing failures only: dashboard assets, a midnight-boundary version-cookie test),make lintclean, API docs regenerated. Live gateway run against anthropic/gemini/groq/openai withmax_output_tokens=16, stream and non-stream: all four now reportincomplete+max_output_tokens, usage normalized, empty answer serializes"text": "".