Repository navigation
fix(providers): honour tool_choice: auto for openai_compatible models - #708
Conversation
Some OpenAI-compatible endpoints ignore both response_format and a forced tool_choice and answer in prose, so every semantic analyzer fails structured-response validation and the scan degrades to static analysis. A tool_choice: auto registry entry now builds ChatOpenAI with tool_choice disabled and selects the function_calling method, and bind_structured_output gives such a model the same prompt instruction and fail-closed retry as Bedrock models restricted to toolChoice auto. Bundle spark-x2.5 (iFlytek Astron Token Plan) with tool_choice: auto. Fixes NVIDIA#707 Signed-off-by: FenjuFu <fufenjupku@gmail.com>
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation is scoped, preserves precedence rules, and includes targeted regression coverage.
Review effort: Balanced
Findings: None
What changed in this PR
Enables reliable structured output for OpenAI-compatible endpoints that require automatic tool selection.
Changes:
- Honors registry-based
tool_choice: auto. - Adds prompt-and-retry handling for prose responses.
- Registers
spark-x2.5and adds regression tests and documentation.
| File | Description |
|---|---|
README.md |
Documents OpenAI-compatible automatic tool selection. |
src/skillspector/llm_utils.py |
Detects disabled forced tool choice and requires tool calls. |
src/skillspector/providers/chat_models.py |
Passes disabled parameters to ChatOpenAI. |
src/skillspector/providers/openai_compatible/model_registry.yaml |
Registers spark-x2.5. |
src/skillspector/providers/openai_compatible/provider.py |
Applies registry tool-choice behavior. |
tests/unit/test_llm_utils.py |
Tests prose rejection and successful tool parsing. |
tests/unit/test_new_providers.py |
Tests registry and provider behavior. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
rng1995
left a comment
There was a problem hiding this comment.
[SkillSpector Review]
Hi @FenjuFu, thank you for the careful diagnosis in #707 and for a fix that reuses the existing Bedrock tool_choice: auto path instead of adding a new one!
Value and readiness: This solves the stated problem. A model with tool_choice: auto in the openai_compatible registry now gets a ChatOpenAI with disabled_params={"tool_choice": None} and the function_calling method. Prose answers become retryable StructuredOutputParseErrors, so an endpoint like spark-x2.5 no longer drops semantic analysis to static-only. Models without that entry, and the openai/nv_build/ollama providers that share create_openai_compatible_chat_model, are unchanged. Retry exhaustion still reports llm_structured_response_invalid, so completeness is not overclaimed. It is ready for final maintainer review. The notes below are optional.
Material findings
- [Non-blocking]
src/skillspector/llm_utils.py:413: the condition rewrite also changes Bedrock behavior, which the PR body describes as unchanged. Before, any explicit method skipped_require_tool_call(if kwargs or ...). Now an explicitfunction_callingon an auto-only Bedrock model is wrapped. That applies toSKILLSPECTOR_STRUCTURED_OUTPUT_METHOD=function_callingor a registrystructured_output: function_calling. The tool-call instruction is then appended, and a prose answer raises a retryableStructuredOutputParseErrorinstead of reachingparse_responseasNone. I think this is an improvement. Please mention it in the description, and consider pinning it with a test next totest_auto_only_tool_choice_asks_for_the_call_and_rejects_prose. - [Non-blocking]
src/skillspector/llm_utils.py:413:kwargs.get("method", "function_calling")assumes LangChain's default method isfunction_calling. That holds forChatBedrockConverse. It does not hold forChatOpenAI, whosewith_structured_outputdefaults tojson_schema(langchain-openai 1.3.3). Today this cannot fire, becauseOpenAICompatibleProvideralways returns an explicit method whenever it setsdisabled_params. Any future caller that binds adisabled_paramsChatOpenAIwith no method hint would get a "call the tool" instruction on aresponse_formatrequest. Wrapping thedisabled_paramsbranch only whenmethod == "function_calling"is explicit would remove that trap. - [Non-blocking]
tests/unit/test_llm_utils.py:859: a short companion case forSKILLSPECTOR_STRUCTURED_OUTPUT_METHOD=json_schema(or a declaredstructured_output: json_schema) with atool_choice: autoChatOpenAIwould show that the binding is left unwrapped and thatresponse_formatis sent.test_declared_structured_output_method_winscovers only the provider hint, not the binder.
PIC tradeoffs:
- Bundled model entry. The PR adds a vendor-specific
spark-x2.5entry to the bundledopenai_compatibleregistry. This follows the existing Groq/Together/DeepSeek entries. It also matters becauseSKILLSPECTOR_MODEL_REGISTRYreplaces the bundled file rather than merging with it, so users would otherwise have to copy the whole registry. The cost is a maintained entry that CI cannot exercise live. Itscontext_length: 262144and missingmax_output_tokensrest on the author's probe. The output budget then falls back to the percentage-of-context default inmodel_info.get_max_output_tokens. - Overlap with #691. #691 (chrisknvidia) edits the same
bind_structured_outputcondition to add a TP4preferred_method. The two PRs will textually conflict, so whichever lands second needs a rebase. They compose semantically: a preferredfunction_callingon an auto-only model would be wrapped under this PR's rule.
Verification and gaps: I read the full diff at the PR head and traced:
- Request path. The path runs
OpenAICompatibleProvider.create_chat_model→create_openai_compatible_chat_model→ChatOpenAI(disabled_params=...)→with_structured_output(method="function_calling"). In langchain-openai 1.1.10 (the floor), 1.3.3 (locked) and 1.6.6,bind_kwargs = self._filter_disabled_params(...)dropstool_choicewhile keepingtools. The newChatOpenAI-level test asserts exactly that. - Retry path.
_require_tool_callerrors reach_StructuredResponseValidationErrorand the bounded retry loop inllm_analyzer_base._invoke_batch_with_retries. - Precedence.
SKILLSPECTOR_STRUCTURED_OUTPUT_METHODcomes first, then a declaredstructured_output, thentool_choice: auto. - Other providers. Bedrock's default method is
function_callingin langchain-aws 1.6.1, so the empty-kwargs case is unchanged.
The live scan numbers in the PR body are the author's; I did not reproduce them. Tests were not executed locally per review policy. CI: all 6 checks are green, DCO is signed off, and there are no merge conflicts with main. The Copilot review raised no findings.
Decision: Approved (reviewed head cc49361ee74572439a62610ed27bec248aa6d71b)
…ed ChatOpenAI ChatOpenAI defaults to json_schema, so only wrap its tool_choice-disabled form in the tool-call prompt and retry when function_calling is selected explicitly. Bedrock auto-only models keep their default tool binding. Pin the binder for an explicit method on an auto-only model and for a non-tool method on a tool_choice-disabled ChatOpenAI. Signed-off-by: FenjuFu <fufenjupku@gmail.com>
rng1995
left a comment
There was a problem hiding this comment.
[SkillSpector Review]
Hi @FenjuFu, thank you for tightening this after the last review!
Value and readiness: ed66750 resolves the hidden-default note. _binds_unforced_tool_call now wraps only when the bound method actually binds a tool. For Bedrock auto-only models this is None or function_calling, as before. For a ChatOpenAI with tool_choice disabled it is only an explicit function_calling, so the default json_schema path stays unwrapped. The openai_compatible fix the PR is for still works: the provider test without an env override still gets the tool call asked for. Ready for final maintainer review.
Previous findings:
- The condition assumed the LangChain default method is
function_calling, which is not true forChatOpenAI: Resolved (llm_utils.py:420-433). - No test that a
json_schemaoverride on an auto-only model stays unwrapped: Resolved.test_auto_only_tool_choice_with_an_explicit_methodcovers both methods, andtest_chat_openai_with_tool_choice_disabled_keeps_a_non_tool_methodchecks thatresponse_formatis sent with notools. - Bedrock side effect of an explicitly requested
function_calling: now covered by the parametrized auto-only test. The behavior is deliberate and documented in the docstring.
Material findings: None.
PIC tradeoffs: Unchanged: the vendor-specific spark-x2.5 registry entry is bundled here, and #691 edits nearby lines, so whichever PR lands second needs a rebase.
Verification and gaps: I read the full cc49361..ed66750 delta and traced both branches of the new predicate. All 6 CI checks pass on ed66750. Tests were not run locally, per policy.
Decision: Approved (reviewed head ed667501c8e7cc8e599068d0ca2119dfa1159f48)

Fixes #707.
Some OpenAI-compatible endpoints ignore both
response_formatand a forcedtool_choiceand answer in prose. iFlytek's Astron Token Plan (spark-x2.5) is one. Behindopenai_compatible, every semantic analyzer then fails structured-response validation and the scan degrades to static analysis only. The Bedrock provider already handles this class of model throughtool_choice: autoand_require_tool_call. This PR letsopenai_compatibleuse the same path.Changes
OpenAICompatibleProviderreadstool_choicefrom its registry (bundled orSKILLSPECTOR_MODEL_REGISTRY). Fortool_choice: autoit buildsChatOpenAIwithdisabled_params={"tool_choice": None}. LangChain's_filter_disabled_paramsthen drops the forced choice, andstructured_output_methodreturnsfunction_calling. A declaredstructured_output:still wins, and so doesSKILLSPECTOR_STRUCTURED_OUTPUT_METHOD.create_openai_compatible_chat_modeltakes an optionaldisabled_params.bind_structured_outputtreats aChatOpenAIwithtool_choicedisabled like Bedrock'ssupports_tool_choice_values=("auto",), but only whenfunction_callingis selected explicitly, becauseChatOpenAIdefaults tojson_schema. It then adds the prompt instruction, and a prose answer raisesStructuredOutputParseError, which the analyzers retry.SKILLSPECTOR_STRUCTURED_OUTPUT_METHOD=function_callingon a Bedrock model restricted totoolChoiceautoused to skip_require_tool_call, so a prose answer parsed toNone. It is now wrapped too, so a prose answer becomes a retryable parse error.SKILLSPECTOR_STRUCTURED_OUTPUT_METHOD=json_schemastays unwrapped. (BedrockProviderhas nostructured_output_method, so a registrystructured_output:entry does not reach it; that is unchanged here.)spark-x2.5in theopenai_compatibleregistry:context_length: 262144(256K, as published for the hosted model) andtool_choice: auto. There is nomax_output_tokens, because the endpoint acceptedmax_completion_tokensup to 300000 in a probe.Apart from the Bedrock case above, other providers and models are unchanged. For a
ChatOpenAIwithoutdisabled_params, the binding is the same as before.Validation
tests/unit/test_new_providers.pycover: default models keep a forcedtool_choice;spark-x2.5getsdisabled_paramsandfunction_calling; a registry override can declaretool_choice: auto; and an explicitstructured_outputwins.tests/unit/test_llm_utils.pyuses a realChatOpenAIwith stubbed_generate. The request carries the tool withouttool_choice, the prompt asks for the call, a prose answer raisesStructuredOutputParseError, and the next tool-call answer parses. All 5 new tests fail onmainand pass with the change.function_calling(wrapped) andjson_schema(unwrapped) on an auto-only model, and atool_choice-disabledChatOpenAIwith no method orjson_schemasendingresponse_formatwithout tools or the prompt instruction.tests/unit/test_llm_utils.py,test_new_providers.py,test_providers.pyandtest_bedrock_provider.py: 291 passed, 9 skipped.ruff checkandruff format --checkpass.tests/fixtures/malicious_skillagainst the Token Plan endpoint withSKILLSPECTOR_MODEL=spark-x2.5:main(2226747)analysis_completeness.statuspartialcompletecoverage_percentOn
main,semantic_developer_intent,semantic_security_discoveryandsemantic_quality_policyend withllm_structured_response_invalid. On this branch the 4 retries are prose answers that the prompt-and-retry path recovers.