From 8ec2e1740e410a52ae203f85df35d02ebbff977f Mon Sep 17 00:00:00 2001 From: Tryanks Date: Wed, 23 Sep 2026 02:38:43 +0800 Subject: [PATCH 1/3] orchestrate: GPT-6 Sol replaces Astra as the bundled Codex executor The execution layer's bundled Codex profile is now `gpt-6-sol`; Astra stays a bundled collaboration peer and can still be added as an executor by hand. Sol's guidance chooses effort per task (low for small clear briefs, medium for routine work, higher as constraints grow) instead of Astra's low-only rule, and the workflow prompt and settings copy follow. Migration: an untouched bundled Codex executor row, whether the old Sol 5.6 default or either Astra default text, becomes the Sol 6 built-in unless the user already added Sol 6; rows with custom text, an endpoint profile, or changed switches stay as they are. --- assets/orchestrate/workflow.md | 8 +-- crates/core/src/settings.rs | 113 +++++++++++++++++++++----------- crates/runtime/src/app/tests.rs | 46 +++++++------ locales/en.yml | 2 +- locales/zh-CN.yml | 2 +- 5 files changed, 108 insertions(+), 63 deletions(-) diff --git a/assets/orchestrate/workflow.md b/assets/orchestrate/workflow.md index 0615faed..018bbf7d 100644 --- a/assets/orchestrate/workflow.md +++ b/assets/orchestrate/workflow.md @@ -12,10 +12,10 @@ If Orchestrate tool schemas are deferred, discover and load them before starting delegated execution. Read the current fleet, compare enabled execution profiles across all providers, and select a task-fit model, endpoint profile, and per-call effort using the configured strengths and caveats. Provider family gives no -preference. The bundled GPT-6 executor is dispatched at low effort only: never -pass it medium or above, since higher efforts cost more without better results. -Route UI-driving and eyes-on-screen verification to it first. Choose another -profile only when its description better fits the task. +preference. Take each call's effort from the profile's own guidance: the +bundled GPT-6 Sol executor runs small clear tasks at low and routine work at +medium, rising only as constraints or reasoning difficulty grow. Choose the +profile whose description best fits the task. ## Route the work diff --git a/crates/core/src/settings.rs b/crates/core/src/settings.rs index 1a5e5aa1..b9aadf8a 100644 --- a/crates/core/src/settings.rs +++ b/crates/core/src/settings.rs @@ -324,6 +324,7 @@ const OLD_DEFAULT_SOL_DEFINITION: &str = "Execution model for scoped implementat const OLD_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review. Consider it alongside Sol across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Baseline execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, and evidence gathering. Default to low effort for a clear brief; raise effort only when a specific piece demonstrably needs more depth. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; const DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use. It is exceptionally strong at driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows), so route eyes-on-screen verification and UI-driving work here first. Always dispatch it at low effort: low outperforms the former Sol executor at xhigh on quality and at a fraction of the token cost, so medium or higher is never justified for this profile and only wastes money; a task that seems to need more depth needs a better brief, not more effort. Keep unrelated improvements out of scope. Report the concrete result and relevant checks concisely."; +const DEFAULT_SOL_6_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use, including driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows). Use low for small mechanical tasks with a clear brief, medium for routine bounded work, high or xhigh as interacting constraints or reasoning difficulty grow, and max only for the hardest well-defined problems or when a lower effort has demonstrably stalled. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; const DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const DEFAULT_ASTRA_DEFINITION: &str = include_str!("../../../assets/orchestrate/astra.md"); const DEFAULT_FABLE_DEFINITION: &str = include_str!("../../../assets/orchestrate/fable-5-1.md"); @@ -351,7 +352,7 @@ pub fn orchestrate_efforts( .unwrap_or_default() } else { let fallback: &[&str] = match (provider, model) { - (ProviderKind::Codex, "gpt-5.6-sol" | "gpt-6-astra") => { + (ProviderKind::Codex, "gpt-5.6-sol" | "gpt-6-sol" | "gpt-6-astra") => { &["low", "medium", "high", "xhigh", "max", "ultra"] } ( @@ -427,11 +428,7 @@ impl Default for OrchestrateSettings { ), ], child_models: vec![ - builtin_model( - ProviderKind::Codex, - "gpt-6-astra", - DEFAULT_GPT_6_EXECUTION_DEFINITION, - ), + builtin_model(ProviderKind::Codex, "gpt-6-sol", DEFAULT_SOL_6_DEFINITION), builtin_model( ProviderKind::ClaudeCode, "claude-opus-5-5", @@ -471,21 +468,26 @@ impl LegacyOrchestrateModel { replace_untouched_sol: bool, replace_untouched_opus: bool, ) -> Option { - if !collaboration + // An untouched bundled Codex executor (Sol 5.6, then Astra) follows + // the bundle to GPT-6 Sol unless the user already added it; rows with + // custom guidance, an endpoint profile, or changed switches stay. + let untouched_codex_executor = !collaboration && self.effort.is_none() && self.entry.provider == ProviderKind::Codex - && self.entry.model == "gpt-5.6-sol" && self.entry.profile_id.is_none() - && self.entry.description == OLD_DEFAULT_SOL_DEFINITION && self.entry.enabled && !self.entry.fast - { + && match self.entry.model.as_str() { + "gpt-5.6-sol" => self.entry.description == OLD_DEFAULT_SOL_DEFINITION, + "gpt-6-astra" => { + self.entry.description == OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION + || self.entry.description == DEFAULT_GPT_6_EXECUTION_DEFINITION + } + _ => false, + }; + if untouched_codex_executor { return replace_untouched_sol.then(|| { - builtin_model( - ProviderKind::Codex, - "gpt-6-astra", - DEFAULT_GPT_6_EXECUTION_DEFINITION, - ) + builtin_model(ProviderKind::Codex, "gpt-6-sol", DEFAULT_SOL_6_DEFINITION) }); } if !collaboration @@ -593,7 +595,7 @@ impl From for OrchestrateSettings { .iter() .any(|entry| entry.entry.provider == provider && entry.entry.model == model) }; - let has_execution_astra = has_execution(ProviderKind::Codex, "gpt-6-astra"); + let has_execution_sol_6 = has_execution(ProviderKind::Codex, "gpt-6-sol"); let has_execution_opus_5_5 = has_execution(ProviderKind::ClaudeCode, "claude-opus-5-5"); ( decisions @@ -603,7 +605,7 @@ impl From for OrchestrateSettings { data.child_models .into_iter() .filter_map(|entry| { - entry.migrate(false, !has_execution_astra, !has_execution_opus_5_5) + entry.migrate(false, !has_execution_sol_6, !has_execution_opus_5_5) }) .collect(), ) @@ -674,6 +676,7 @@ impl OrchestrateSettings { pub fn builtin_child_definition(provider: ProviderKind, model: &str) -> Option<&'static str> { match (provider, model) { + (ProviderKind::Codex, "gpt-6-sol") => Some(DEFAULT_SOL_6_DEFINITION), (ProviderKind::Codex, "gpt-6-astra") => Some(DEFAULT_GPT_6_EXECUTION_DEFINITION), (ProviderKind::ClaudeCode, "claude-opus-5-5" | "claude-opus-5") => { Some(DEFAULT_OPUS_DEFINITION) @@ -1609,13 +1612,13 @@ mod tests { ["gpt-6-astra", "claude-fable-5-1"] ); assert_eq!(defaults.child_models.len(), 2); - assert_eq!(defaults.child_models[0].model, "gpt-6-astra"); + assert_eq!(defaults.child_models[0].model, "gpt-6-sol"); assert_eq!(defaults.child_models[0].provider, ProviderKind::Codex); assert!(!defaults.child_models[0].fast); assert!( defaults.child_models[0] .description - .contains("Always dispatch it at low effort") + .contains("Use low for small mechanical tasks") ); assert_ne!( defaults.child_models[0].description, @@ -1751,19 +1754,51 @@ mod tests { } #[test] - fn orchestrate_refreshes_untouched_previous_gpt_6_execution_text() { - let old_json = format!( - r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#, - serde_json::to_string(OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION).unwrap() - ); - let migrated: OrchestrateSettings = serde_json::from_str(&old_json).unwrap(); - assert_eq!( - migrated.child_models[0].description, - DEFAULT_GPT_6_EXECUTION_DEFINITION + fn orchestrate_moves_untouched_astra_executors_to_sol_6() { + for text in [ + OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION, + DEFAULT_GPT_6_EXECUTION_DEFINITION, + ] { + let old_json = format!( + r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#, + serde_json::to_string(text).unwrap() + ); + let migrated: OrchestrateSettings = serde_json::from_str(&old_json).unwrap(); + assert_eq!( + migrated.child_models, + [builtin_model( + ProviderKind::Codex, + "gpt-6-sol", + DEFAULT_SOL_6_DEFINITION + )] + ); + // Customised text, a profile, or a flipped switch keeps Astra. + for variant in [ + old_json.replace(text, "mine"), + old_json.replace(r#""enabled":true"#, r#""enabled":false"#), + old_json.replace(r#""fast":false"#, r#""fast":true"#), + old_json.replace( + r#""provider":"codex","#, + r#""provider":"codex","profile_id":"corp","#, + ), + ] { + let kept: OrchestrateSettings = serde_json::from_str(&variant).unwrap(); + assert_eq!(kept.child_models.len(), 1, "input: {variant}"); + assert_eq!( + kept.child_models[0].model, "gpt-6-astra", + "input: {variant}" + ); + } + } + // An untouched Astra row next to a user-added Sol 6 row is dropped + // rather than duplicating Sol 6. + let both = format!( + r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-sol","description":"Mine."}},{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#, + serde_json::to_string(DEFAULT_GPT_6_EXECUTION_DEFINITION).unwrap() ); - let customized = old_json.replace(OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION, "mine"); - let kept: OrchestrateSettings = serde_json::from_str(&customized).unwrap(); - assert_eq!(kept.child_models[0].description, "mine"); + let migrated: OrchestrateSettings = serde_json::from_str(&both).unwrap(); + assert_eq!(migrated.child_models.len(), 1); + assert_eq!(migrated.child_models[0].description, "Mine."); } #[test] @@ -1808,7 +1843,7 @@ mod tests { } #[test] - fn orchestrate_migration_does_not_duplicate_existing_execution_astra() { + fn orchestrate_migration_does_not_duplicate_existing_execution_sol_6() { let old_json = r#"{ "decision_models": [], "child_models": [ @@ -1821,7 +1856,7 @@ mod tests { }, { "provider": "codex", - "model": "gpt-6-astra", + "model": "gpt-6-sol", "profile_id": "custom-codex", "enabled": false, "fast": true, @@ -1832,7 +1867,7 @@ mod tests { let migrated: OrchestrateSettings = serde_json::from_str(old_json).unwrap(); assert_eq!(migrated.child_models.len(), 1); - assert_eq!(migrated.child_models[0].model, "gpt-6-astra"); + assert_eq!(migrated.child_models[0].model, "gpt-6-sol"); assert_eq!( migrated.child_models[0].profile_id.as_deref(), Some("custom-codex") @@ -1921,11 +1956,15 @@ mod tests { #[test] fn orchestrate_settings_patches_deduplicate_within_each_role() { let mut settings = Settings::default(); - let executor = settings.orchestrate.child_models[0].clone(); let peer = settings.orchestrate.decision_models[0].clone(); - assert_eq!(executor.provider, peer.provider); - assert_eq!(executor.model, peer.model); + // The same model serves both roles with role-specific guidance. + let executor = builtin_model( + peer.provider, + &peer.model, + DEFAULT_GPT_6_EXECUTION_DEFINITION, + ); assert_ne!(executor.description, peer.description); + settings.orchestrate.child_models[0] = executor.clone(); let mut duplicate = executor.clone(); duplicate.profile_id = Some("another-endpoint".into()); diff --git a/crates/runtime/src/app/tests.rs b/crates/runtime/src/app/tests.rs index f4c8f80c..7aad1266 100644 --- a/crates/runtime/src/app/tests.rs +++ b/crates/runtime/src/app/tests.rs @@ -1853,7 +1853,7 @@ fn orchestrate_guidance_and_current_configuration_are_composed() { ); assert!(first.contains("### Execution models — `dispatch`")); assert!(first.contains( - "#### `codex` / `gpt-6-astra` — available `effort`: `low`, `medium`, `high`, `xhigh`, `max`" + "#### `codex` / `gpt-6-sol` — available `effort`: `low`, `medium`, `high`, `xhigh`, `max`" )); assert!(first.ends_with("\n\nShip it")); settings.decision_models[0].enabled = false; @@ -1905,8 +1905,8 @@ fn dispatch_validates_against_live_efforts_instead_of_bundled_fallback() { let catalogs = HashMap::from([( ProviderKind::Codex, vec![ModelSpec { - id: "gpt-6-astra".into(), - display_name: "GPT-6 Astra".into(), + id: "gpt-6-sol".into(), + display_name: "GPT-6 Sol".into(), is_default: false, options: vec![OptionDescriptor::Select { id: "reasoningEffort".into(), @@ -1936,7 +1936,7 @@ fn dispatch_validates_against_live_efforts_instead_of_bundled_fallback() { .contains("unsupported effort max") ); let configuration = render_orchestrate_configuration(&settings, None, &catalogs); - assert!(configuration.contains("`gpt-6-astra` — available `effort`: `medium`, `high`, `deep`")); + assert!(configuration.contains("`gpt-6-sol` — available `effort`: `medium`, `high`, `deep`")); } #[test] @@ -1957,13 +1957,13 @@ fn loaded_catalog_marks_missing_orchestrate_model_unavailable() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("low"), None, &catalogs ) .unwrap_err(), - expected + expected.replace("gpt-6-astra", "gpt-6-sol") ); assert_eq!( resolve_orchestrate_collaboration( @@ -1979,9 +1979,15 @@ fn loaded_catalog_marks_missing_orchestrate_model_unavailable() { ); let configuration = render_orchestrate_configuration(&settings, None, &catalogs); assert!(configuration.contains("#### `codex` / `gpt-6-astra` — unavailable")); + assert!(configuration.contains("#### `codex` / `gpt-6-sol` — unavailable")); assert!(configuration.contains( "Unavailable: model `gpt-6-astra` is not present in the loaded `codex` catalog." )); + assert!( + configuration.contains( + "Unavailable: model `gpt-6-sol` is not present in the loaded `codex` catalog." + ) + ); } #[test] @@ -2014,7 +2020,7 @@ fn collaboration_and_execution_resolve_separate_profile_lists() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("low"), None, &HashMap::new() @@ -2045,14 +2051,14 @@ fn collaboration_and_execution_resolve_separate_profile_lists() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("low"), None, &HashMap::new() ) .unwrap() .1, - "gpt-6-astra" + "gpt-6-sol" ); settings.decision_models[0].enabled = true; settings.child_models[0].enabled = false; @@ -2402,7 +2408,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("low"), None, &HashMap::new() @@ -2410,7 +2416,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { .unwrap(), ( ProviderKind::Codex, - "gpt-6-astra".into(), + "gpt-6-sol".into(), Some("low".into()), false, None @@ -2421,7 +2427,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("medium"), Some("KIMI"), &HashMap::new() @@ -2429,7 +2435,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { .unwrap(), ( ProviderKind::Codex, - "gpt-6-astra".into(), + "gpt-6-sol".into(), Some("medium".into()), false, Some("kimi".into()), @@ -2438,7 +2444,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { let unknown_profile = resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("medium"), Some("missing"), &HashMap::new(), @@ -2469,7 +2475,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some(effort), None, &HashMap::new() @@ -2483,7 +2489,7 @@ fn orchestrate_dispatch_enforces_child_allow_list_and_defaults() { let wrong_effort = resolve_orchestrate_dispatch( &settings, "codex", - Some("gpt-6-astra"), + Some("gpt-6-sol"), Some("imaginary"), None, &HashMap::new(), @@ -6550,7 +6556,7 @@ fn orchestrate_dispatch_fast_override_beats_profile_setting() { .orchestrate .child_models .iter_mut() - .find(|child| child.model == "gpt-6-astra") + .find(|child| child.model == "gpt-6-sol") .unwrap(); max.fast = true; }); @@ -6562,7 +6568,7 @@ fn orchestrate_dispatch_fast_override_beats_profile_setting() { purpose: orchestrate_mcp::ThreadPurpose::Execution, parent_id: parent_id.clone(), provider: "codex".into(), - model: Some("gpt-6-astra".into()), + model: Some("gpt-6-sol".into()), effort: Some(effort.into()), profile: None, access: None, @@ -6617,7 +6623,7 @@ fn orchestrate_dispatch_resolves_cwd_before_reply() { purpose: orchestrate_mcp::ThreadPurpose::Execution, parent_id, provider: "codex".into(), - model: Some("gpt-6-astra".into()), + model: Some("gpt-6-sol".into()), effort: None, profile: None, access: None, @@ -6704,7 +6710,7 @@ fn orchestrate_worktree_dispatch_resolves_child_cwd_to_worktree() { purpose: orchestrate_mcp::ThreadPurpose::Execution, parent_id, provider: "codex".into(), - model: Some("gpt-6-astra".into()), + model: Some("gpt-6-sol".into()), effort: None, profile: None, access: None, diff --git a/locales/en.yml b/locales/en.yml index ee6cf7e7..862edc22 100644 --- a/locales/en.yml +++ b/locales/en.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "Collaboration unavailable until medium/high capabilities are known" children: title: "Execution models" - description: "Built-in executors: GPT-6 Astra and Opus 5.5. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Dispatch the GPT-6 executor at low only; it is the first choice for computer use, and higher efforts cost more without better results." + description: "Built-in executors: GPT-6 Sol and Opus 5.5. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Pick GPT-6 Sol's effort per task: low for clear small tasks, medium for routine work, higher only as constraints or reasoning difficulty grow." add: "Add execution model" empty: "No child-model profiles are configured. /orchestrate remains available, but dispatch calls will be rejected until a child model is added." none_enabled: "Every child-model profile is switched off. /orchestrate can still plan, but all dispatch calls will be rejected." diff --git a/locales/zh-CN.yml b/locales/zh-CN.yml index c174302d..50c851bc 100644 --- a/locales/zh-CN.yml +++ b/locales/zh-CN.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "尚无 medium/high 能力信息,暂不可协作" children: title: "执行模型" - description: "内置执行模型为 GPT-6 Astra 和 Opus 5.5。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。GPT-6 执行模型仅使用 low;它是电脑操作任务的首选,更高档位只会增加成本而不会提升效果。" + description: "内置执行模型为 GPT-6 Sol 和 Opus 5.5。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。按任务为 GPT-6 Sol 选择推理档位:明确的小任务用 low,常规工作用 medium,仅在约束交织或推理困难时再提高。" add: "添加执行模型" empty: "目前没有配置任何子模型。/orchestrate 仍然可用,但在添加子模型前,派发调用会被拒绝。" none_enabled: "所有子模型配置都已关闭。/orchestrate 仍可进行规划,但所有派发调用都会被拒绝。" From 2097d4e6f856cfcd35b34bd09a14eb0be562d33f Mon Sep 17 00:00:00 2001 From: Tryanks Date: Wed, 23 Sep 2026 02:48:57 +0800 Subject: [PATCH 2/3] orchestrate: describe GPT-6 Sol from its published guidance; Opus 5.5 is the second opinion Sol's profile now follows OpenAI's model guidance rather than Astra's low-only rule: built for complex coding and agentic workflows, about half the factual errors of GPT-5.6 Sol, Astra-level reliability at a fraction of the cost, Opus-class computer use at far lower cost per task, and medium as the recommended starting effort. It is named the primary executor so ordinary work goes there first. Opus 5.5's profile is reworded as the secondary executor: a Claude Code perspective, a cross-provider second implementation or review, or a fallback when Sol is unavailable or stalls. The previous Opus text is kept as a constant so untouched rows migrate to the new wording and customised ones are left alone. The workflow prompt and settings copy say the same. --- assets/orchestrate/workflow.md | 10 +++++----- crates/core/src/settings.rs | 31 ++++++++++++++++++++++++++++--- locales/en.yml | 2 +- locales/zh-CN.yml | 2 +- 4 files changed, 35 insertions(+), 10 deletions(-) diff --git a/assets/orchestrate/workflow.md b/assets/orchestrate/workflow.md index 018bbf7d..cc90250c 100644 --- a/assets/orchestrate/workflow.md +++ b/assets/orchestrate/workflow.md @@ -11,11 +11,11 @@ remain with you. If Orchestrate tool schemas are deferred, discover and load them before starting delegated execution. Read the current fleet, compare enabled execution profiles across all providers, and select a task-fit model, endpoint profile, and per-call -effort using the configured strengths and caveats. Provider family gives no -preference. Take each call's effort from the profile's own guidance: the -bundled GPT-6 Sol executor runs small clear tasks at low and routine work at -medium, rising only as constraints or reasoning difficulty grow. Choose the -profile whose description best fits the task. +effort using the configured strengths and caveats. The bundled GPT-6 Sol +executor is the default workhorse: route ordinary implementation, investigation, +and verification to it at medium effort, low for narrow edits, and higher only +when the work needs more planning or checking. Pick another profile only when +its description names a reason that fits the task, never by provider family. ## Route the work diff --git a/crates/core/src/settings.rs b/crates/core/src/settings.rs index b9aadf8a..53f35868 100644 --- a/crates/core/src/settings.rs +++ b/crates/core/src/settings.rs @@ -324,8 +324,9 @@ const OLD_DEFAULT_SOL_DEFINITION: &str = "Execution model for scoped implementat const OLD_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review. Consider it alongside Sol across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Baseline execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, and evidence gathering. Default to low effort for a clear brief; raise effort only when a specific piece demonstrably needs more depth. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; const DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use. It is exceptionally strong at driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows), so route eyes-on-screen verification and UI-driving work here first. Always dispatch it at low effort: low outperforms the former Sol executor at xhigh on quality and at a fraction of the token cost, so medium or higher is never justified for this profile and only wastes money; a task that seems to need more depth needs a better brief, not more effort. Keep unrelated improvements out of scope. Report the concrete result and relevant checks concisely."; -const DEFAULT_SOL_6_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use, including driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows). Use low for small mechanical tasks with a clear brief, medium for routine bounded work, high or xhigh as interacting constraints or reasoning difficulty grow, and max only for the hardest well-defined problems or when a lower effort has demonstrably stalled. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; -const DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; +const DEFAULT_SOL_6_DEFINITION: &str = "Primary execution model: dispatch ordinary implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use here first. GPT-6 Sol is built for complex coding and agentic workflows, with about half the factual errors of GPT-5.6 Sol, fewer misleading claims about its own work, and Astra-level reliability at a fraction of the cost; on computer-use tasks it matches Opus-class models at far lower cost per task. Start at medium, the recommended default for everyday and complex work; use low for narrow mechanical edits, and raise to high or xhigh only for work that needs more planning, analysis, or checking across many steps, sources, or tradeoffs; reserve max for the hardest well-defined problems. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; +const PREVIOUS_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; +const DEFAULT_OPUS_DEFINITION: &str = "Secondary execution model, not the default: Sol handles ordinary work. Choose Opus 5.5 when a task specifically benefits from a Claude Code perspective, when a second independent implementation or review across providers is wanted, or when Sol is unavailable or has demonstrably stalled on it. Strong at agentic coding, cross-file implementation, refactoring, debugging, and review, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const DEFAULT_ASTRA_DEFINITION: &str = include_str!("../../../assets/orchestrate/astra.md"); const DEFAULT_FABLE_DEFINITION: &str = include_str!("../../../assets/orchestrate/fable-5-1.md"); @@ -495,6 +496,7 @@ impl LegacyOrchestrateModel { && self.entry.model == "claude-opus-5" && self.entry.profile_id.is_none() && (self.entry.description == OLD_DEFAULT_OPUS_DEFINITION + || self.entry.description == PREVIOUS_DEFAULT_OPUS_DEFINITION || self.entry.description == DEFAULT_OPUS_DEFINITION) { // An untouched bundled Opus 5 row follows the release to Opus 5.5 @@ -512,6 +514,13 @@ impl LegacyOrchestrateModel { { self.entry.description = DEFAULT_GPT_6_EXECUTION_DEFINITION.into(); } + if !collaboration + && self.entry.provider == ProviderKind::ClaudeCode + && self.entry.model == "claude-opus-5-5" + && self.entry.description == PREVIOUS_DEFAULT_OPUS_DEFINITION + { + self.entry.description = DEFAULT_OPUS_DEFINITION.into(); + } let entry = &mut self.entry; let legacy = self.effort.is_some(); if legacy { @@ -1618,7 +1627,7 @@ mod tests { assert!( defaults.child_models[0] .description - .contains("Use low for small mechanical tasks") + .contains("Start at medium") ); assert_ne!( defaults.child_models[0].description, @@ -1753,6 +1762,22 @@ mod tests { ); } + #[test] + fn orchestrate_refreshes_previous_opus_5_5_default_text() { + let old_json = format!( + r#"{{"decision_models":[],"child_models":[{{"provider":"claude_code","model":"claude-opus-5-5","description":{}}}]}}"#, + serde_json::to_string(PREVIOUS_DEFAULT_OPUS_DEFINITION).unwrap() + ); + let migrated: OrchestrateSettings = serde_json::from_str(&old_json).unwrap(); + assert_eq!( + migrated.child_models[0].description, + DEFAULT_OPUS_DEFINITION + ); + let customized = old_json.replace(PREVIOUS_DEFAULT_OPUS_DEFINITION, "mine"); + let kept: OrchestrateSettings = serde_json::from_str(&customized).unwrap(); + assert_eq!(kept.child_models[0].description, "mine"); + } + #[test] fn orchestrate_moves_untouched_astra_executors_to_sol_6() { for text in [ diff --git a/locales/en.yml b/locales/en.yml index 862edc22..3ac9c771 100644 --- a/locales/en.yml +++ b/locales/en.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "Collaboration unavailable until medium/high capabilities are known" children: title: "Execution models" - description: "Built-in executors: GPT-6 Sol and Opus 5.5. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Pick GPT-6 Sol's effort per task: low for clear small tasks, medium for routine work, higher only as constraints or reasoning difficulty grow." + description: "Built-in executors: GPT-6 Sol, the default workhorse, and Opus 5.5 as a second opinion across providers. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Sol starts at medium; use low for narrow edits and raise effort only when the work needs more planning or checking." add: "Add execution model" empty: "No child-model profiles are configured. /orchestrate remains available, but dispatch calls will be rejected until a child model is added." none_enabled: "Every child-model profile is switched off. /orchestrate can still plan, but all dispatch calls will be rejected." diff --git a/locales/zh-CN.yml b/locales/zh-CN.yml index 50c851bc..3b2ec02a 100644 --- a/locales/zh-CN.yml +++ b/locales/zh-CN.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "尚无 medium/high 能力信息,暂不可协作" children: title: "执行模型" - description: "内置执行模型为 GPT-6 Sol 和 Opus 5.5。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。按任务为 GPT-6 Sol 选择推理档位:明确的小任务用 low,常规工作用 medium,仅在约束交织或推理困难时再提高。" + description: "内置执行模型为 GPT-6 Sol(主力执行者)和 Opus 5.5(跨提供方的第二意见)。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。Sol 默认从 medium 起步;窄范围修改用 low,仅在需要更多规划或核查时再提高。" add: "添加执行模型" empty: "目前没有配置任何子模型。/orchestrate 仍然可用,但在添加子模型前,派发调用会被拒绝。" none_enabled: "所有子模型配置都已关闭。/orchestrate 仍可进行规划,但所有派发调用都会被拒绝。" From c3f7e8e6dd37626a1a49dec694a12c587f94c8e0 Mon Sep 17 00:00:00 2001 From: Tryanks Date: Wed, 23 Sep 2026 03:02:44 +0800 Subject: [PATCH 3/3] orchestrate: position Sol as the workhorse and Opus for UI design, copy, and research Sol: close instruction following, reliable goal completion, computer use, lower cost. Opus 5.5: first choice for UI design, copywriting, creative and open-ended research work, and a cross-provider second implementation or review. Workflow prompt and settings copy say the same. --- assets/orchestrate/workflow.md | 10 ++++++---- crates/core/src/settings.rs | 4 ++-- locales/en.yml | 2 +- locales/zh-CN.yml | 2 +- 4 files changed, 10 insertions(+), 8 deletions(-) diff --git a/assets/orchestrate/workflow.md b/assets/orchestrate/workflow.md index cc90250c..8b487907 100644 --- a/assets/orchestrate/workflow.md +++ b/assets/orchestrate/workflow.md @@ -12,10 +12,12 @@ If Orchestrate tool schemas are deferred, discover and load them before starting delegated execution. Read the current fleet, compare enabled execution profiles across all providers, and select a task-fit model, endpoint profile, and per-call effort using the configured strengths and caveats. The bundled GPT-6 Sol -executor is the default workhorse: route ordinary implementation, investigation, -and verification to it at medium effort, low for narrow edits, and higher only -when the work needs more planning or checking. Pick another profile only when -its description names a reason that fits the task, never by provider family. +executor is the default workhorse: route ordinary implementation, verification, +and computer use to it at medium effort, low for narrow edits, and higher only +when the work needs more planning or checking. Route UI design, copywriting, +creative and open-ended research work to the bundled Opus executor. Pick any other profile +only when its description names a reason that fits the task, never by provider +family. ## Route the work diff --git a/crates/core/src/settings.rs b/crates/core/src/settings.rs index 53f35868..c79ecc7f 100644 --- a/crates/core/src/settings.rs +++ b/crates/core/src/settings.rs @@ -324,9 +324,9 @@ const OLD_DEFAULT_SOL_DEFINITION: &str = "Execution model for scoped implementat const OLD_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review. Consider it alongside Sol across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Baseline execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, and evidence gathering. Default to low effort for a clear brief; raise effort only when a specific piece demonstrably needs more depth. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; const DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use. It is exceptionally strong at driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows), so route eyes-on-screen verification and UI-driving work here first. Always dispatch it at low effort: low outperforms the former Sol executor at xhigh on quality and at a fraction of the token cost, so medium or higher is never justified for this profile and only wastes money; a task that seems to need more depth needs a better brief, not more effort. Keep unrelated improvements out of scope. Report the concrete result and relevant checks concisely."; -const DEFAULT_SOL_6_DEFINITION: &str = "Primary execution model: dispatch ordinary implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use here first. GPT-6 Sol is built for complex coding and agentic workflows, with about half the factual errors of GPT-5.6 Sol, fewer misleading claims about its own work, and Astra-level reliability at a fraction of the cost; on computer-use tasks it matches Opus-class models at far lower cost per task. Start at medium, the recommended default for everyday and complex work; use low for narrow mechanical edits, and raise to high or xhigh only for work that needs more planning, analysis, or checking across many steps, sources, or tradeoffs; reserve max for the hardest well-defined problems. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; +const DEFAULT_SOL_6_DEFINITION: &str = "Primary execution model: dispatch ordinary implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use here first. GPT-6 Sol follows a brief closely and reliably reaches the stated goal, at a lower cost than Opus; it is built for complex coding and agentic workflows, with about half the factual errors of GPT-5.6 Sol and fewer misleading claims about its own work, and it is strong at driving and reading real UIs, so eyes-on-screen verification and UI-driving work go here. Start at medium, the recommended default for everyday and complex work; use low for narrow mechanical edits, and raise to high or xhigh only for work that needs more planning, analysis, or checking across many steps, sources, or tradeoffs; reserve max for the hardest well-defined problems. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely."; const PREVIOUS_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; -const DEFAULT_OPUS_DEFINITION: &str = "Secondary execution model, not the default: Sol handles ordinary work. Choose Opus 5.5 when a task specifically benefits from a Claude Code perspective, when a second independent implementation or review across providers is wanted, or when Sol is unavailable or has demonstrably stalled on it. Strong at agentic coding, cross-file implementation, refactoring, debugging, and review, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; +const DEFAULT_OPUS_DEFINITION: &str = "Execution model for UI design and creative or investigative work: choose Opus 5.5 first for visual and interaction design, layout, copywriting and user-facing text, design exploration, open-ended research, and ideation where taste and breadth matter; it also gives a strong second implementation or review across providers. For ordinary implementation and verification prefer Sol, which follows briefs more closely at lower cost. Use medium for clear bounded work, high for substantial design or research, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely."; const DEFAULT_ASTRA_DEFINITION: &str = include_str!("../../../assets/orchestrate/astra.md"); const DEFAULT_FABLE_DEFINITION: &str = include_str!("../../../assets/orchestrate/fable-5-1.md"); diff --git a/locales/en.yml b/locales/en.yml index 3ac9c771..f005a9dd 100644 --- a/locales/en.yml +++ b/locales/en.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "Collaboration unavailable until medium/high capabilities are known" children: title: "Execution models" - description: "Built-in executors: GPT-6 Sol, the default workhorse, and Opus 5.5 as a second opinion across providers. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Sol starts at medium; use low for narrow edits and raise effort only when the work needs more planning or checking." + description: "Built-in executors: GPT-6 Sol, the default workhorse (close instruction following, reliable goal completion, computer use, lower cost), and Opus 5.5 for UI design, copywriting, creative and investigative work. Each provider/model appears once per list, including endpoint profiles, so collaboration and execution may use separate profiles for the same model. Sol starts at medium; use low for narrow edits and raise effort only when the work needs more planning or checking." add: "Add execution model" empty: "No child-model profiles are configured. /orchestrate remains available, but dispatch calls will be rejected until a child model is added." none_enabled: "Every child-model profile is switched off. /orchestrate can still plan, but all dispatch calls will be rejected." diff --git a/locales/zh-CN.yml b/locales/zh-CN.yml index 3b2ec02a..8b5a30bd 100644 --- a/locales/zh-CN.yml +++ b/locales/zh-CN.yml @@ -395,7 +395,7 @@ orchestrate: effort_unavailable: "尚无 medium/high 能力信息,暂不可协作" children: title: "执行模型" - description: "内置执行模型为 GPT-6 Sol(主力执行者)和 Opus 5.5(跨提供方的第二意见)。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。Sol 默认从 medium 起步;窄范围修改用 low,仅在需要更多规划或核查时再提高。" + description: "内置执行模型为 GPT-6 Sol(主力执行者:指令跟随强、可靠达成目标、擅长电脑操作、成本较低)和 Opus 5.5(UI 设计、文案、创意与调研类工作)。同一提供方的同一模型在每个列表中仅保留一条,因此同一模型可分别配置协作与执行角色。Sol 默认从 medium 起步;窄范围修改用 low,仅在需要更多规划或核查时再提高。" add: "添加执行模型" empty: "目前没有配置任何子模型。/orchestrate 仍然可用,但在添加子模型前,派发调用会被拒绝。" none_enabled: "所有子模型配置都已关闭。/orchestrate 仍可进行规划,但所有派发调用都会被拒绝。"