Skip to content

feat(providers): adopt GPT-6.1 Sol and Claude Opus 5.5 defaults with compatible structured output - #806

Draft
rng1995 wants to merge 15 commits into
mainfrom
naren/gpt-6-1-sol-claude-5-5-defaults
Draft

rng1995 wants to merge 15 commits into
mainfrom
naren/gpt-6-1-sol-claude-5-5-defaults

Conversation

@rng1995

@rng1995 rng1995 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adopts GPT-6.1 Sol and Claude Opus 5.5 as the default models, and makes the Claude 5.5 models and GPT-6.1 Sol actually work on every route SkillSpector supports.

Provider Slot Before After
openai all slots gpt-5.4 gpt-6.1-sol
anthropic analyzers claude-opus-4-6 claude-opus-5-5
anthropic meta_analyzer claude-sonnet-4-6 claude-sonnet-5-5
bedrock all slots us.anthropic.claude-sonnet-4-6-20250915-v1:0 us.anthropic.claude-sonnet-4-6 (64K output cap)

Opened as a draft: the two default-flip commits wait for the acceptance run described below.

Customer impact

  • Scans on the openai and anthropic providers run on the current models, at lower per-token prices than today's defaults:
    • gpt-6.1-sol is $2 / $10 per MTok;
    • claude-opus-5-5 is $4 / $20, vs $5 / $25 for Opus 4.6.
  • Claude Opus/Sonnet 5.5 work on native Anthropic, Anthropic proxy, Bedrock, and through OpenAI-compatible gateways. Today all of them fail with HTTP 400.
  • The Bedrock default now names a model ID that exists in the AWS catalog.
  • Fail fast:
    • A temperature setting (SKILLSPECTOR_TEMPERATURE) for gpt-6.1-sol, claude-opus-5-5 or claude-sonnet-5-5 stops scan and baseline before analysis (exit 2) on every hosted provider, with a message that names the setting. The MCP server and the report show the same message in llm_error.
    • An effort setting (SKILLSPECTOR_REASONING_EFFORT) of none or minimal does the same for gpt-6.1-sol.
    • Users who had already picked these models got a 400 on every batch from these settings.
  • Breaking change: SKILLSPECTOR_TEMPERATURE worked with the old gpt-5.4 and claude-opus-4-6 defaults. With the new openai and anthropic defaults, a scan with it set now stops before analysis. Unset it, or pin the old model with SKILLSPECTOR_MODEL.

Root cause

  • Claude Opus 5.5 and Sonnet 5.5 reject forced tool calls (tool_choice: tool|any returns HTTP 400). LangChain's default structured output forces a tool call. OpenAI-compatible gateways also turn json_schema into a forced tool call for Claude models. Only claude-fable-5-1 and claude-mythos-5-1 were handled.
  • GPT-6.1 Sol supports reasoning effort low–max only, and rejects temperature at every effort.
  • The dated Bedrock ID us.anthropic.claude-sonnet-4-6-20250915-v1:0 is not in the AWS model catalog, and AWS caps Sonnet 4.6 output at 64K.

Fix

  • One model-name normalizer (model_name, with claude_model_name on top) reads bare, gateway-namespaced (<vendor>/anthropic/claude-opus-5-5, …/bedrock-claude-opus-5-5), dotted, dated, :latest, Bedrock profile and ARN forms. Forced-tool and sampling rejection and GPT-6.1 Sol detection use it on every route, and Bedrock's separate parser is removed.
  • Structured output per route:
    • anthropic and anthropic_proxy: native json_schema for 5.5.
    • bedrock: toolChoice: auto with a prompted call and bounded retry.
    • openai and openai_compatible with Claude 5.5 gateway IDs: the schema is bound as a tool with no tool_choice and no response_format, and prose answers are retried.
    • gpt-6.1-sol and claude-opus-5 keep today's path.
  • Registry entries:
    • claude-opus-5-5 and claude-sonnet-5-5 on anthropic and anthropic_proxy at 1M / 128K, json_schema.
    • Opt-in Bedrock 5.5 geo and global profiles at 1M / 128K, tool_choice: auto. These are never a default.
  • Control guard: resolve_sampling_parameters(model) runs reject_unsupported_controls, so every hosted provider checks controls before any request, azure_openai included. It runs after credentials are resolved, so the no-credentials path is unchanged. A Bedrock registry entry can declare sampling: rejected for an application-inference-profile ARN that serves Claude 5.5.
  • Scan preflight: scan and baseline build each slot model first and exit 2 with the guard's message. is_llm_available returns the same message, so the MCP server and the report show it in llm_error.
  • One forced-tool decision: forced_tool_choice_supported(model, registry_path) (registry entry first, then model name) is shared by the openai, openai_compatible and bedrock providers. The OpenAI-protocol builder takes forced_tool_choice and builds the disabled_params itself.
  • Docs: README.md, docs/DEVELOPMENT.md, .env.example.

Commits:

  • 93abc9f feat(providers): support Claude Opus and Sonnet 5.5 structured output on native routes
  • 0c9ab3b feat(providers): bind gateway Claude 5.5 IDs without a forced tool call
  • 8800d8c feat(providers): reject controls GPT-6.1 Sol refuses before any request
  • 2b461b7 fix(providers): repair the Bedrock default to us.anthropic.claude-sonnet-4-6 with a 64K cap
  • 6c69e2c feat(providers): default OpenAI to gpt-6.1-sol
  • 7b23977 feat(providers): default native Anthropic to Opus 5.5 analyzers and Sonnet 5.5 meta
  • b10b674 refactor(providers): read Claude model names through one normalizer on every route
  • a1e6ffc fix(providers): apply the control guard on every hosted provider
  • c66504b fix(cli): stop a scan before analysis when a model rejects a requested control
  • 89d342b test(semantic): name the OpenAI refusal test after the failed status it pins
  • cd2d771 docs: say an unsupported Sol reasoning effort stops the scan

Tests

  • make test-ci equivalent, offline, after the review fixes: 10,688 passed, 0 failed, 14 skipped, 4 xfailed. Two timing-bound tests in tests/test_batch_scan_security.py failed once under concurrent local load and passed on rerun. One preflight test added afterwards passes in the focused run. The 91% total coverage figure is from before the review fixes and was not re-measured. The run used Python 3.13 locally; CI uses 3.12.
  • New tests/unit/test_model_capabilities.py, 34 test functions (121 parametrized cases):
    • normalizer matrix, Sol :latest/:nitro/@… IDs included;
    • forced-tool and sampling tables;
    • request shapes for Sol and for gateway Claude 5.5;
    • control guards, Azure OpenAI included;
    • scan preflight: CLI exit 2 with the reason, llm_error, --no-llm.
  • Slot precedence: test_anthropic_meta_slot_default_yields_to_model_overrides in tests/unit/test_constants.py.
  • Bedrock: a sampling: rejected registry entry guards an application-inference-profile ARN (tests/unit/test_bedrock_provider.py).
  • New refusal regressions in test_semantic_security_discovery.py: a refused structured response, and an OpenAIRefusalError, both report failed or incomplete, never clean.
  • ruff check, ruff format --check and git diff --check are clean. Strict model validation passes for openai, anthropic, bedrock and anthropic_proxy.

Residual

Release note

  • The OpenAI default gpt-6.1-sol is also the OpenAI credential-fallback model.
  • Anthropic defaults are claude-opus-5-5 for analyzers and claude-sonnet-5-5 for meta.
  • Breaking: SKILLSPECTOR_TEMPERATURE worked with the old gpt-5.4 and claude-opus-4-6 defaults. With the new defaults, and with any Sol or Claude 5.5 model on every hosted provider, it stops scan and baseline before analysis with exit 2. Unset it, or pin the old model with SKILLSPECTOR_MODEL.
  • The Bedrock default is us.anthropic.claude-sonnet-4-6 with a 64K cap. Custom SKILLSPECTOR_MODEL_REGISTRY files must copy the new key.

Known limits

  • Claude 5.5 reached through the OpenAI credential fallback still requests json_schema. A follow-up will pass the effective provider into bind_structured_output.
  • A refused batch on the native Anthropic path fails that analyzer. This is fail-closed and pinned by tests; per-batch refusal recovery is a follow-up.
  • The Bedrock registry entries us.anthropic.claude-opus-4-6-20250915-v1:0 and us.anthropic.claude-opus-4-5-20250514-v1:0 predate this PR, and their dates do not match those releases. A follow-up will check their profile IDs and output caps against the AWS model cards, as this PR did for Sonnet 4.6.

Not run here (they need keys or paid runs):

  • make test-provider openai anthropic;
  • the acceptance run;
  • the gateway check;
  • the optional Bedrock check.
Plan: GPT-6.1 Sol and Claude Opus 5.5 defaults

Plan: GPT-6.1 Sol and Claude Opus 5.5 defaults

Default model changes

Provider Slot Current New Reasoning effort Why
openai analyzers + meta_analyzer gpt-5.4 gpt-6.1-sol Unset (API default medium); none/minimal rejected Current OpenAI model; strict JSON schema on Chat Completions without tools
anthropic analyzers claude-opus-4-6 claude-opus-5-5 Unset (medium); thinking always on Current Opus at 20% lower per-token price; native JSON schema avoids the forced-tool 400
anthropic meta_analyzer claude-sonnet-4-6 claude-sonnet-5-5 Unset (model default high) Keeps the cheaper meta pass; switch to Opus 5.5 if the acceptance run shows meta refusals
bedrock all us.anthropic.claude-sonnet-4-6-20250915-v1:0 (128K output) us.anthropic.claude-sonnet-4-6 (64K output) Not sent (unchanged) The dated ID is not in the AWS catalog; AWS caps output at 64K
anthropic_proxy all claude-sonnet-4-6 Unchanged Unchanged Gets 5.5 metadata and routing only
azure_openai, nv_build, gemini, ollama, openai_compatible all gpt-4o, z-ai/glm-5.3, gemini-3.8-flash, llama3.1:8b, llama-3.1-70b-versatile Unchanged Unchanged Separate provider changes

The PR also adds opt-in Bedrock metadata for the 5.5 models. It is never used as a default:

  • us./eu./global.anthropic.claude-sonnet-5-5 at 1M input / 128K output, tool_choice: auto (per the AWS model cards).
  • us./eu./au./jp./global.anthropic.claude-opus-5-5, same limits.

SKILLSPECTOR_MODEL still sets every slot.

GPT-6.1 Sol vs Claude Opus 5.5 for SkillSpector

Criterion GPT-6.1 Sol Claude Opus 5.5
Quality Ahead on some medium-effort vendor tasks Leads at max effort (Terminal-Bench Science 63.3 vs 57.0); no security-scanner benchmark for either
Refusals on security content Rated "Critical" for cyber; false-refusal rate unpublished Cyber, bio and reasoning-extraction classifiers (Sonnet 5.5: broader set); a refusal fails the analyzer
Structured output Strict json_schema on Chat Completions; no tools needed Native json_schema. Any forced tool call returns HTTP 400. Bedrock and OpenAI-compatible gateways need an unforced tool call.
Request controls Rejects temperature; effort low–max Rejects temperature; effort low–max; thinking cannot be disabled
Cost per MTok (in / out) $2 / $10 $4 / $20 (Opus 4.6: $5 / $25); per-scan cost depends on thinking tokens
Speed ~56 tok/s ~95 tok/s
Context / max output 1.05M / 128K 1M / 128K

Compatibility fixes

Route Model Today After this PR
anthropic, anthropic_proxy claude-opus-5-5, claude-sonnet-5-5 Forced tool call: 400; no budgets Native json_schema; 1M / 128K budgets
bedrock 5.5 geo/global profiles Forced toolChoice: 400; fallback budget toolChoice: auto with a prompted call and bounded retry; 1M / 128K
openai, openai_compatible Gateway IDs naming Claude 5.5 (e.g. <vendor>/anthropic/claude-opus-5-5) OpenAI-compatible gateways turn json_schema into a forced tool call: 400 Schema bound as a tool, no tool_choice or response_format; prose answers retried
Every hosted provider (OpenAI-protocol, azure_openai, anthropic, anthropic_proxy, bedrock) gpt-6.1-sol, Claude Opus/Sonnet 5.5 An explicit temperature (or Sol none/minimal effort) is sent and returns 400 on every batch The scan stops before analysis (exit 2) with a message naming the setting

Validation

Check What
Offline make lint format-check test-ci; strict model validation per changed provider
Live provider tests make test-provider openai anthropic, new and old defaults
Acceptance run (blocks the two default-flip commits) Fixtures ssd_clean, ssd3_nl_exfiltration, ssd1_semantic_injection, malicious_skill. Arms: Sol vs gpt-5.4; Opus 5.5 + Sonnet 5.5 meta at medium and high vs Opus 4.6 + Sonnet 4.6; an Opus 5.5 meta arm. Record completion, findings, refusals, latency, usage and cost.
Gateway An OpenAI-compatible gateway serving Opus 5.5, via openai, openai_compatible and anthropic with a base URL

Open decisions

  1. Keep one PR, or split before the two default-flip commits.
  2. Meta slot: Sonnet 5.5 (proposed), or Opus 5.5 for every slot.
  3. Gateway Claude detection: by model name (proposed), or registry entries only.

🤖 Generated with Claude Code

rng1995 and others added 7 commits October 8, 2026 02:01
… on native routes

Claude Opus 5.5 and Sonnet 5.5 answer a forced tool call with HTTP 400 and reject sampling controls. Add both to the shared forced-tool table and a new SAMPLING_REJECTED_MODELS table, so the anthropic and anthropic_proxy providers request the native JSON-schema response format and the bedrock provider binds the schema with toolChoice auto.

Add claude_model_name(), which reads the bare Claude name from bare, namespaced, prefixed, dotted, dated and Bedrock identifiers, and reject_unsupported_controls(), which fails before any request when SKILLSPECTOR_TEMPERATURE is set for a model that rejects it. anthropic, anthropic_proxy and bedrock call it after credentials resolve, so a missing key still returns None and the OpenAI fallback is unchanged.

Register both models for anthropic and anthropic_proxy (1,000,000 context, 128,000 output, json_schema) and the AWS-documented Bedrock profiles as explicit opt-in entries with toolChoice auto: Opus 5.5 on us/eu/au/jp/global and Sonnet 5.5 on us/eu/global. No default changes.

Bind the live OpenAI and Anthropic provider tests through bind_structured_output with room for reasoning tokens, and pin that a refused structured response fails the semantic analyzer instead of reading as a clean scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Gateways namespace or prefix Claude model names (azure/anthropic/claude-opus-5-5, aws/anthropic/bedrock-claude-opus-5-5), so the bare-name rule missed them and LangChain forced a tool call that Claude Opus and Sonnet 5.5 reject with HTTP 400. OpenAI-compatible gateways can also turn a json_schema response format for Claude into a forced tool call.

Read the Claude name with claude_model_name() on every route:
- anthropic and anthropic_proxy request the native JSON-schema format for gateway IDs too (registry entries still win).
- openai gains forced_tool_choice_supported() and structured_output_method(): for those Claude IDs it disables tool_choice and binds the schema as an unforced tool call that the prompt asks for, retrying a prose answer. Other models, including gpt-6.1-sol and earlier Claude IDs, keep LangChain's default.
- openai_compatible applies the same name rule when the registry declares no tool_choice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
GPT-6.1 Sol rejects an explicit temperature and the reasoning efforts none
and minimal with HTTP 400, and LangChain only strips temperature for the
gpt-5 family, so SkillSpector would send them and fail mid-scan.

reject_unsupported_controls() now also matches gpt-6.1-sol, bare or behind
a gateway namespace or prefix, with an optional snapshot suffix. Any
explicit SKILLSPECTOR_TEMPERATURE (1.0 included) and any
SKILLSPECTOR_REASONING_EFFORT other than low, medium, high, xhigh or max
raise ValueError with the setting to change. The shared OpenAI-protocol
builder calls it after credentials resolve, so openai, openai_compatible
and the other builder users get the Sol rules and the Claude 5.5
temperature rule; a missing key still returns None. Other models keep
passing both controls through.

SKILLSPECTOR_SEED is still forwarded to Sol and recorded as requested and
retained; the docs call it best-effort.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…net-4-6 with a 64K cap

The Bedrock default us.anthropic.claude-sonnet-4-6-20250915-v1:0 is not a
catalogued AWS ID, and its registry entry allowed 128K output tokens while
AWS caps Claude Sonnet 4.6 at 64K on Bedrock.

BEDROCK_DEFAULT_MODEL becomes the suffix-free cross-region profile ID
us.anthropic.claude-sonnet-4-6 that the AWS model card lists, and the
registry key moves with it, capped at 64,000 output tokens. A test pins
that no Claude 5.5 profile is the default; those stay explicit opt-in.
The dated Opus 4.6 and 4.5 keys are unchanged.

A custom SKILLSPECTOR_MODEL_REGISTRY file replaces the bundled registry,
so it has to carry the new key to keep the default's budgets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The openai provider defaults every model slot to gpt-6.1-sol instead of
gpt-5.4. The bundled registry already carries its 1,050,000-token context
and 128,000-token output budgets, and it binds structured output through a
strict, tool-free json_schema response format.

The OPENAI_API_KEY credential fallback uses the openai defaults, so it now
runs gpt-6.1-sol for any slot without a model override. An explicit
SKILLSPECTOR_TEMPERATURE, or a reasoning effort other than low, medium,
high, xhigh or max, now fails before the first request on these paths
instead of being sent. Callers who need the previous model set
SKILLSPECTOR_MODEL=gpt-5.4.

The README provider table and the OPENAI_API_KEY row are updated. The row
also records that Claude Opus or Sonnet 5.5 gateway IDs are not supported
through the fallback, which keeps the active provider's structured-output
method.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…onnet 5.5 meta

The anthropic provider defaults its analyzer slots to claude-opus-5-5
instead of claude-opus-4-6, and keeps a cheaper meta_analyzer pass with
claude-sonnet-5-5 instead of claude-sonnet-4-6. Both are registered with
1,000,000-token context and 128,000-token output budgets and bind
structured output through the native JSON-schema format, since they
reject a forced tool call.

Effort stays unset, so the API defaults apply: medium for Opus 5.5 and
high for Sonnet 5.5. SKILLSPECTOR_REASONING_EFFORT applies to every slot,
meta included. An explicit SKILLSPECTOR_TEMPERATURE now fails before the
first request with the default models. Slot precedence is unchanged:
SKILLSPECTOR_MODEL replaces the meta default too, and
SKILLSPECTOR_MODEL_META_ANALYZER wins over both. Callers who need the
previous models set SKILLSPECTOR_MODEL=claude-opus-4-6 and
SKILLSPECTOR_MODEL_META_ANALYZER=claude-sonnet-4-6. anthropic_proxy keeps
its default.

The README provider table and effort row are updated, and a
SKILLSPECTOR_MODEL_<SLOT> row documents the per-slot override. A test
pins the meta default against both overrides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…n every route

rejects_forced_tool_call() and rejects_sampling_controls() now normalize the
model ID themselves through claude_model_name(), so callers pass the raw ID
instead of wrapping it at each of the six call sites.

Bedrock's forced_tool_choice_supported() uses the same rule, which makes
claude_model_from_bedrock_id() dead; it is removed and its Bedrock ID cases
now pin claude_model_name(). Every Bedrock ID the old parser recognised
still resolves, and the registry tool_choice entry still wins first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review: 9 findings (0 P1, 3 P2, 6 P3); 7 posted inline.

Findings outside the diff or about the description:

  • PR_DESCRIPTION: [P2] SKILLSPECTOR_TEMPERATURE now breaks default-model users; the PR description says it already failed

Customer impact says these settings "caused a 400 on every batch" before this PR. That only holds for users who had already picked gpt-6.1-sol or Claude 5.5. On the old defaults, an explicit temperature worked:

  • langchain_openai validate_temperature removes any temperature other than 1 for gpt-5* models, so gpt-5.4 ran normally.
  • claude-opus-4-6 accepts temperature.

After this PR, a user on the openai or anthropic default with SKILLSPECTOR_TEMPERATURE set goes from working scans to failed analyzers on upgrade. DEVELOPMENT.md even uses 0 as its example value. The release note only says the setting "fails fast".

Suggested fix: List this as a breaking change under Customer impact and in the release note. For example: "SKILLSPECTOR_TEMPERATURE worked with the gpt-5.4 / claude-opus-4-6 defaults. With the new defaults, unset it, or pin the old model with SKILLSPECTOR_MODEL."

  • PR_DESCRIPTION: [P3] The test count for test_model_capabilities.py is wrong

The description says the new module has "67 tests" and lists "slot precedence" among them. At HEAD it has 25 test functions, and pytest --collect-only collects 104 cases. The slot-precedence test (test_anthropic_meta_slot_default_yields_to_model_overrides) lives in tests/unit/test_constants.py. The overall suite totals are fine.

Suggested fix: Change it to "25 test functions (104 parametrized cases)" and list slot precedence under test_constants.py.

Comment thread src/skillspector/providers/chat_models.py Outdated
Comment thread .env.example
Comment thread src/skillspector/providers/chat_models.py Outdated
Comment thread src/skillspector/providers/bedrock/model_registry.yaml Outdated
Comment thread src/skillspector/providers/bedrock/model_registry.yaml
Comment thread src/skillspector/providers/openai_compatible/provider.py Outdated
Comment thread tests/nodes/analyzers/test_semantic_security_discovery.py Outdated
rng1995 and others added 4 commits October 8, 2026 10:35
Resolve sampling parameters for a model: resolve_sampling_parameters now takes the model and runs reject_unsupported_controls itself, so azure_openai (whose registry lists gpt-6.1-sol) gets the temperature guard the other providers already had, and no caller can skip it. The guard raises UnsupportedControlError, a ValueError subclass.

GPT-6.1 Sol IDs are read through the same model_name normalizer as Claude IDs (last path segment, lowercased, cut at '@' or ':'), so gpt-6.1-sol:latest, openai/gpt-6.1-sol:nitro and gpt-6.1-sol@2026 are guarded too.

A registry entry can declare 'sampling: rejected'. Bedrock reads it, so an application-inference-profile ARN that serves Claude 5.5 can opt in to the temperature guard, as it can already opt out of forced tool calls with 'tool_choice: auto'.

The registry-then-name forced tool_choice decision moves into structured_output.forced_tool_choice_supported, and the openai, openai_compatible and bedrock providers delegate to it. create_openai_compatible_chat_model takes forced_tool_choice and builds the disabled_params dict that bind_structured_output looks for in one place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…d control

A SKILLSPECTOR_TEMPERATURE (or a GPT-6.1 Sol effort of none/minimal) that the configured model rejects used to surface only while each analyzer built its model: the scan logged one traceback per semantic analyzer, exited 2, and the report showed llm_available: true with a generic telemetry error.

unsupported_control_error builds each distinct slot model the way the analyzers do and returns the UnsupportedControlError text; it does nothing when neither control is set, and leaves other failures, such as missing credentials, to the availability path. The scan and baseline commands call it first and exit 2 with that message. is_llm_available also returns it, so the MCP server and the report carry the reason in llm_error.

The docs now say that the setting stops the scan on every hosted provider, that it worked with the previous gpt-5.4 and claude-opus-4-6 defaults, and how to declare 'sampling: rejected' for a Bedrock application-inference-profile ARN.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…it pins

test_openai_refusal_is_incomplete_not_clean asserts status == 'failed'; rename it to test_openai_refusal_is_failed_not_clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The README already says an unsupported SKILLSPECTOR_REASONING_EFFORT
for gpt-6.1-sol stops a scan before analysis with exit code 2. Use the
same wording in .env.example and docs/DEVELOPMENT.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
@rng1995

rng1995 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review follow-up for findings outside the diff:

  • A rejected temperature or effort does not stop the scan, and the report never shows why: Addressed in c66504b. scan and baseline now check every slot model before any analysis. They build each distinct slot model the way the analyzers do and exit 2 with the guard's message, for example "SKILLSPECTOR_TEMPERATURE is not supported by gpt-6.1-sol; unset it to use this model". is_llm_available returns the same message, so the MCP server and the report show it in llm_error. The guard now raises UnsupportedControlError (a ValueError subclass), so missing credentials still take the existing degraded path. New CliRunner and is_llm_available tests cover the default openai and anthropic models. cd2d771 updates the reasoning-effort wording in .env.example and docs/DEVELOPMENT.md to say the scan stops (exit 2), matching the README.
  • SKILLSPECTOR_TEMPERATURE now breaks default-model users; the PR description says it already failed: Addressed in c66504b. Customer impact and the release note now list this as a breaking change: SKILLSPECTOR_TEMPERATURE worked with the gpt-5.4 and claude-opus-4-6 defaults, so unset it or pin the old model with SKILLSPECTOR_MODEL. README, DEVELOPMENT.md and .env.example say the same.
  • azure_openai skips the temperature guard, but its registry lists gpt-6.1-sol: Addressed in a1e6ffc. resolve_sampling_parameters now takes the model and runs reject_unsupported_controls itself, so every caller gets the guard, azure_openai included, and no provider can skip it. New tests check that Azure with gpt-6.1-sol and SKILLSPECTOR_TEMPERATURE fails before any request, while gpt-4o keeps the temperature. The docs now say the guard covers every hosted provider.
  • The GPT-6.1 Sol matcher misses ':' and '@' suffixes that claude_model_name strips: Addressed in a1e6ffc. The last-segment and [@:] suffix stripping is now a shared model_name helper in structured_output.py. claude_model_name and the new is_gpt_6_1_sol both use it, and the local regex and re import are gone from chat_models.py. The Sol test now includes gpt-6.1-sol:latest, openai/gpt-6.1-sol:nitro and gpt-6.1-sol@2026.
  • Application-profile ARNs serving Claude 5.5 can opt out of forced tool calls, but no key enables the temperature guard: Addressed in a1e6ffc. Added an optional sampling: rejected registry key. rejects_sampling_controls reads it when given a registry path, and BedrockProvider passes its registry, so a declared application-inference-profile ARN that serves Claude 5.5 fails before any request. The registry header, provider docstring and README describe the key, and a new Bedrock test covers the ARN with and without the declaration.
  • The Opus 4.6 and 4.5 Bedrock entries keep the dated IDs this PR removes for Sonnet 4.6: Not changed. The Opus 4.6 and 4.5 entries predate this PR and are not defaults, and the AWS model cards could not be checked offline here. Rather than guess new IDs and caps, I listed the two entries under Known limits as a follow-up.
  • The test count for test_model_capabilities.py is wrong: Not changed. The description now gives the real counts, 34 test functions and 121 parametrized cases after the review fixes. It lists slot precedence under tests/unit/test_constants.py (test_anthropic_meta_slot_default_yields_to_model_overrides).
  • Registry-then-name forced_tool_choice logic is now copied in three providers: Addressed in a1e6ffc. Added structured_output.forced_tool_choice_supported(model, registry_path), which checks the registry entry first and then the model name. The openai, openai_compatible and bedrock providers delegate to it. create_openai_compatible_chat_model now takes forced_tool_choice: bool and builds {"tool_choice": None} itself, so the shape _binds_unforced_tool_call matches is defined in one place.
  • Test name says incomplete but the test asserts failed: Addressed in 89d342b. Renamed the test to test_openai_refusal_is_failed_not_clean to match the failed status it asserts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant