Skip to content

fix: prefer tool output for OpenAI Opus 5 TP4 - #691

Merged
rng1995 merged 9 commits into
NVIDIA:mainfrom
chrisknvidia:fix/christopherk/opus5-tp4-function-calling
Oct 8, 2026
Merged

rng1995 merged 9 commits into
NVIDIA:mainfrom
chrisknvidia:fix/christopherk/opus5-tp4-function-calling

Conversation

@chrisknvidia

@chrisknvidia chrisknvidia commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

TP4 can receive a {"json": ...} wrapper from an OpenAI-compatible Opus 5 gateway when LangChain uses json_schema, leaving the analysis incomplete after bounded retries. Prefer function_calling for TP4 on the reproduced route: effective provider openai, a native OpenAI chat client, exact model azure/anthropic/claude-opus-5, and reasoning/thinking controls unset. Initial and deadline-refreshed models use the same preference; explicit environment settings and provider hints retain precedence. The code documents excluded providers/aliases and the public workaround tracking record.

Selected native tool bindings now accept exactly one valid call to the expected schema tool. Retain the raw message because LangChain's parser can silently discard additional calls, including a contradictory assessment after a clean first assessment. Missing, unknown, additional, or malformed calls become structured-response errors in sync and async execution. They use existing bounded retries and remain degraded/incomplete if exhausted. Parsed schema types, usage callbacks, unforced-tool prompt behavior, and CLI adapter handling are preserved.

Validate analyzer preferences before override selection, and document the actual client default and full method precedence. Merged current main, preserving concurrent TP4 execution.

Verification

  • On final head: 588 focused tests passed, covering the shared binding, Bedrock, TP4, base analyzer, usage, and real-SDK HTTP suites. Ruff, formatting, diff checks, and wheel build passed. The earlier follow-up passed 990 focused tests; 9 optional internal-provider tests were skipped because nv_inference is absent.
  • Five fresh live CLI scans from the final installed wheel completed with 100% coverage on the requested Opus 5 label: default clean executable, default fenced-code mismatch, reasoning effort high, explicit function_calling/SARIF, and explicit json_schema/JSON. Clean exited 0; mismatch scans exited 1 for findings and retained TP4 ownership/code evidence.
  • 51 HTTP regressions cover sync/async, deadline refresh, titled schemas, refusals, malformed arguments, duplicate/contradictory/unknown tool calls, and recovery. Full graph tests assert the transmitted exact model and forced tool choice in addition to JSON/SARIF completeness and finding ownership.
  • 18 installed-wheel CLI cases passed with controlled SDK HTTP replies in JSON and SARIF. Checks include exits, four-attempt exhaustion/two-attempt recovery, ledger reasons, findings, SARIF schema, and usage accounting: 52 TP4 requests, 52 usage records, 1,040 tokens. Deadline, cancellation, concurrency-permit reuse, and mixed failure probes also passed.
  • Independent compatibility review passed 98 cases across OpenAI and Anthropic SDK HTTP interfaces and Bedrock Converse client stubs, including Pydantic/dictionary schema names, callbacks, and sync/async behavior. No remaining actionable code-review finding was identified.
  • All 106 Python source files in the tested wheel and installed package match the committed source. Wheel SHA256: fae8add71a53f88640af64fb97035ce09f0a67dbb4e9feae8ca8874623425b20.

Live evidence proves the requested label and selected gateway route; underlying deployment identity is not independently attested. Hosted refusal/malformed/multiple-call responses were not observed, and other providers were not exercised live. Controlled negative cases are separate from live provider evidence. Hosted CI on the updated head must finish. The specifically requested Security Review agent remains unavailable, so no specialized security-review pass is claimed.

Adoption

Ready for another review pass. All six original review threads are answered and resolved. Package release and downstream dependency adoption remain necessary before protected integration jobs can validate the delivered package.

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
@chrisknvidia
chrisknvidia force-pushed the fix/christopherk/opus5-tp4-function-calling branch from f89659e to 5710822 Compare September 30, 2026 00:44
@chrisknvidia
chrisknvidia marked this pull request as ready for review September 30, 2026 02:03

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SkillSpector Review]

@chrisknvidia, thank you for tracking down the Opus 5 {"json": ...} wrapper and keeping the workaround this tightly scoped. The PR description and the live-scan evidence made this easy to review.

The change looks correct and fails closed. An explicit SKILLSPECTOR_STRUCTURED_OUTPUT_METHOD and provider hints still take precedence. The new hook covers both the initial model and the deadline-refreshed model, and every failure path I tried leaves TP4 incomplete rather than silently clean. I left a few scope questions and some optional cleanups inline; none of them blocks the fix.

What I verified (head 5710822)

  • CI: all six checks are green. Locally, tests/test_mcp_tool_poisoning.py and tests/unit/test_llm_utils.py pass (139 passed), as do the LLM analyzer base unit tests (37 passed). None of the touched files has changed on main since the merge base.
  • Real ChatOpenAI with _generate patched, model azure/anthropic/claude-opus-5, provider openai:
    • TP4 now sends tool_choice={"type": "function", "function": {"name": "_TP4AnalysisResult"}} and no response_format.
    • A valid tool call: completed.
    • Out-of-range tool args (confidence: 7.0): 4 attempts, then degraded / llm_structured_response_invalid.
    • Prose with no tool call: 1 attempt, then failed / llm_batch_failed (NotImplementedError). More on this inline.
  • The new test_opus5_function_calling_refusal_stays_incomplete really reaches the patched sync _generate. It does not pass because of a network error to example.invalid.
  • ChatOpenAI.with_structured_output defaults to method="json_schema" in langchain-openai 1.3.3, which explains why this route was sending json_schema before.

Docs (README isn't in the diff, so noting it here)

The structured-output paragraph in README.md (around line 247) says LangChain's default "forces a tool call." That is no longer true for ChatOpenAI, whose default is json_schema. The paragraph also doesn't mention the new analyzer-level preference. One sentence would help operators who are debugging binding behavior: TP4 uses function_calling for provider openai with azure/anthropic/claude-opus-5, and SKILLSPECTOR_STRUCTURED_OUTPUT_METHOD still overrides it.

Comment thread src/skillspector/nodes/analyzers/mcp_tool_poisoning.py
Comment thread src/skillspector/nodes/analyzers/mcp_tool_poisoning.py Outdated
Comment thread src/skillspector/nodes/analyzers/mcp_tool_poisoning.py
Comment thread src/skillspector/llm_utils.py Outdated
Comment thread src/skillspector/llm_utils.py Outdated
Comment thread tests/test_mcp_tool_poisoning.py
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
@chrisknvidia

chrisknvidia commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Addressed all six original review threads, with answers and resolutions, and pushed the Review Guru follow-up in d9940d7.

The extra review found a P2: LangChain could discard a contradictory or malformed second tool call and accept the first clean assessment. Native tool bindings now validate the raw response before accepting exactly one expected assessment. Missing/ambiguous/malformed calls retry within the existing bound and stay degraded/incomplete on exhaustion. The graph tests now also assert the actual model and forced tool choice.

Final-head validation: 588 focused tests passed; 98 independent client-contract cases passed; 18 installed-wheel CLI cases passed for JSON/SARIF, including four-attempt exhaustion and transient recovery. Five fresh real-gateway scans from the installed wheel completed with 100% coverage, including high reasoning effort and both explicit methods. All 106 Python source files match the committed source, wheel, and installation. Ruff, formatting, diff checks, and wheel build passed.

The PR description now records the final scope, validation, and adoption requirements. Ready for another review pass; hosted CI is queued. Live checks cover the selected gateway/requested label, while malformed and ambiguous provider outputs were controlled HTTP cases. The requested specialized Security Review agent remains unavailable; no pass is claimed for it.

chrisknvidia and others added 2 commits October 5, 2026 13:54
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Resolve the README.md conflict by keeping main's NVIDIA Build
z-ai/glm-5.3 reasoning-effort paragraph and the PR's client-specific
structured-output wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SkillSpector Review]

Hi @chrisknvidia, thank you for working through every point from the last round, and for the real HTTP transport tests that pin down exactly what TP4 sends!

Value and readiness: Ready to merge. TP4 asks for tool output only on provider openai with the exact azure/anthropic/claude-opus-5 label, and only when no reasoning or thinking controls are set. Explicit and provider method choices still win. A missing or ambiguous tool call now gets bounded retries and ends degraded instead of failing on the first attempt.

Previous findings (review at 5710822)

  • Scope comment for the openai-only check: Resolved. The rationale sits beside _TP4_TOOL_OUTPUT_MODELS. openai_compatible, nv_inference, azure_openai and nv_build are covered by tests and bind {}.
  • Module constant and exact match: Resolved. Dated, suffixed, :latest, bare and uppercase labels are tested as excluded.
  • Forced tool_choice with reasoning: Resolved. Against a loopback HTTP server with a real ChatOpenAI:
    • no reasoning sends the forced _TP4AnalysisResult tool and no response_format;
    • SKILLSPECTOR_REASONING_EFFORT=high sends json_schema with reasoning_effort=high and no tool_choice;
    • an explicit function_calling override still forces the tool, as the README documents.
  • No tool call raised NotImplementedError outside the retry handling: Resolved for sync and async. A prose-only reply now makes four HTTP attempts and then ends degraded with llm_structured_response_invalid. Through the full graph, analysis_completeness.is_complete is False. A single valid tool call that also carries text still succeeds on the first attempt.
  • Validate preferred_method first: Resolved, and tested under each override.
  • Shared fixture: Resolved (patch_openai_chat_model).
  • README (body item): Resolved. It now gives the client-specific default, the precedence (env, then provider hint, then analyzer preference, then client default) and the TP4 rule.

Branch update: I merged current main into the branch (94d6c16, signed off). The only conflict was README.md: I kept main's NVIDIA Build z-ai/glm-5.3 reasoning paragraph and your corrected structured-output wording. Main's new default_reasoning_effort applies only to nv_build, and _tp4_reasoning_configured would detect it anyway.

Non-blocking notes

  • d9940d7 makes the "exactly one valid tool call" check apply to every function_calling or auto-tool binding, not just TP4. That includes Bedrock models registered with tool_choice: auto and openai_compatible spark-x2.5. Two tool calls used to be accepted silently, with the first winning. Now they are retried and the run ends degraded. That is stricter and fails safe, and the stubbed Bedrock tests pass, but it is worth a line in the release notes.
  • Any non-empty SKILLSPECTOR_REASONING_EFFORT, including none, turns off the TP4 tool preference. That is conservative, and a test codifies it.

Verification: On the merged tree:

  • the PR's test files: 226 passed;
  • tests/unit: 2,628 passed, 14 skipped;
  • the LLM, TP4, semantic and meta-analyzer suites: 954 passed;
  • ruff check and ruff format --check are clean.

All six CI checks pass on 94d6c16.


Decision: Ready to merge. A human maintainer needs to approve (reviewed head 94d6c1617d1990afb276867a0abd8a3b921e3fcf). This bot pushed commits to the branch, so it does not approve the PR itself.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants