Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions assets/orchestrate/workflow.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,11 +11,13 @@ remain with you.
If Orchestrate tool schemas are deferred, discover and load them before starting
delegated execution. Read the current fleet, compare enabled execution profiles
across all providers, and select a task-fit model, endpoint profile, and per-call
effort using the configured strengths and caveats. Provider family gives no
preference. The bundled GPT-6 executor is dispatched at low effort only: never
pass it medium or above, since higher efforts cost more without better results.
Route UI-driving and eyes-on-screen verification to it first. Choose another
profile only when its description better fits the task.
effort using the configured strengths and caveats. The bundled GPT-6 Sol
executor is the default workhorse: route ordinary implementation, verification,
and computer use to it at medium effort, low for narrow edits, and higher only
when the work needs more planning or checking. Route UI design, copywriting,
creative and open-ended research work to the bundled Opus executor. Pick any other profile
only when its description names a reason that fits the task, never by provider
family.

## Route the work

Expand Down
126 changes: 95 additions & 31 deletions crates/core/src/settings.rs
Original file line number Diff line number Diff line change
Expand Up @@ -324,7 +324,9 @@ const OLD_DEFAULT_SOL_DEFINITION: &str = "Execution model for scoped implementat
const OLD_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review. Consider it alongside Sol across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely.";
const OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Baseline execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, and evidence gathering. Default to low effort for a clear brief; raise effort only when a specific piece demonstrably needs more depth. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely.";
const DEFAULT_GPT_6_EXECUTION_DEFINITION: &str = "Execution model for scoped implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use. It is exceptionally strong at driving and reading real UIs (find_roots → observe_ui → search_ui / inspect_ui / read_text, and act_ui / wait_for when the brief allows), so route eyes-on-screen verification and UI-driving work here first. Always dispatch it at low effort: low outperforms the former Sol executor at xhigh on quality and at a fraction of the token cost, so medium or higher is never justified for this profile and only wastes money; a task that seems to need more depth needs a better brief, not more effort. Keep unrelated improvements out of scope. Report the concrete result and relevant checks concisely.";
const DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely.";
const DEFAULT_SOL_6_DEFINITION: &str = "Primary execution model: dispatch ordinary implementation, debugging with a reproduction, migrations, code review, data analysis, evidence gathering, and computer use here first. GPT-6 Sol follows a brief closely and reliably reaches the stated goal, at a lower cost than Opus; it is built for complex coding and agentic workflows, with about half the factual errors of GPT-5.6 Sol and fewer misleading claims about its own work, and it is strong at driving and reading real UIs, so eyes-on-screen verification and UI-driving work go here. Start at medium, the recommended default for everyday and complex work; use low for narrow mechanical edits, and raise to high or xhigh only for work that needs more planning, analysis, or checking across many steps, sources, or tradeoffs; reserve max for the hardest well-defined problems. Keep unrelated improvements out of scope, match verification to the changed behavior, and report the concrete result and relevant checks concisely.";
const PREVIOUS_DEFAULT_OPUS_DEFINITION: &str = "Execution model for agentic coding, cross-file implementation, refactoring, debugging, and review across providers, including user-facing behavior and API or UI details. Use medium for clear bounded work, high for substantial implementation, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely.";
const DEFAULT_OPUS_DEFINITION: &str = "Execution model for UI design and creative or investigative work: choose Opus 5.5 first for visual and interaction design, layout, copywriting and user-facing text, design exploration, open-ended research, and ideation where taste and breadth matter; it also gives a strong second implementation or review across providers. For ordinary implementation and verification prefer Sol, which follows briefs more closely at lower cost. Use medium for clear bounded work, high for substantial design or research, and xhigh or max when difficult reasoning justifies the extra work; low can suit small mechanical tasks. Match verification to the changed behavior and avoid repetitive self-checking. Report evidence and unresolved limitations concisely.";
const DEFAULT_ASTRA_DEFINITION: &str = include_str!("../../../assets/orchestrate/astra.md");
const DEFAULT_FABLE_DEFINITION: &str = include_str!("../../../assets/orchestrate/fable-5-1.md");

Expand All @@ -351,7 +353,7 @@ pub fn orchestrate_efforts(
.unwrap_or_default()
} else {
let fallback: &[&str] = match (provider, model) {
(ProviderKind::Codex, "gpt-5.6-sol" | "gpt-6-astra") => {
(ProviderKind::Codex, "gpt-5.6-sol" | "gpt-6-sol" | "gpt-6-astra") => {
&["low", "medium", "high", "xhigh", "max", "ultra"]
}
(
Expand Down Expand Up @@ -427,11 +429,7 @@ impl Default for OrchestrateSettings {
),
],
child_models: vec![
builtin_model(
ProviderKind::Codex,
"gpt-6-astra",
DEFAULT_GPT_6_EXECUTION_DEFINITION,
),
builtin_model(ProviderKind::Codex, "gpt-6-sol", DEFAULT_SOL_6_DEFINITION),
builtin_model(
ProviderKind::ClaudeCode,
"claude-opus-5-5",
Expand Down Expand Up @@ -471,28 +469,34 @@ impl LegacyOrchestrateModel {
replace_untouched_sol: bool,
replace_untouched_opus: bool,
) -> Option<OrchestrateChildModel> {
if !collaboration
// An untouched bundled Codex executor (Sol 5.6, then Astra) follows
// the bundle to GPT-6 Sol unless the user already added it; rows with
// custom guidance, an endpoint profile, or changed switches stay.
let untouched_codex_executor = !collaboration
&& self.effort.is_none()
&& self.entry.provider == ProviderKind::Codex
&& self.entry.model == "gpt-5.6-sol"
&& self.entry.profile_id.is_none()
&& self.entry.description == OLD_DEFAULT_SOL_DEFINITION
&& self.entry.enabled
&& !self.entry.fast
{
&& match self.entry.model.as_str() {
"gpt-5.6-sol" => self.entry.description == OLD_DEFAULT_SOL_DEFINITION,
"gpt-6-astra" => {
self.entry.description == OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION
|| self.entry.description == DEFAULT_GPT_6_EXECUTION_DEFINITION
}
_ => false,
};
if untouched_codex_executor {
return replace_untouched_sol.then(|| {
builtin_model(
ProviderKind::Codex,
"gpt-6-astra",
DEFAULT_GPT_6_EXECUTION_DEFINITION,
)
builtin_model(ProviderKind::Codex, "gpt-6-sol", DEFAULT_SOL_6_DEFINITION)
});
}
if !collaboration
&& self.entry.provider == ProviderKind::ClaudeCode
&& self.entry.model == "claude-opus-5"
&& self.entry.profile_id.is_none()
&& (self.entry.description == OLD_DEFAULT_OPUS_DEFINITION
|| self.entry.description == PREVIOUS_DEFAULT_OPUS_DEFINITION
|| self.entry.description == DEFAULT_OPUS_DEFINITION)
{
// An untouched bundled Opus 5 row follows the release to Opus 5.5
Expand All @@ -510,6 +514,13 @@ impl LegacyOrchestrateModel {
{
self.entry.description = DEFAULT_GPT_6_EXECUTION_DEFINITION.into();
}
if !collaboration
&& self.entry.provider == ProviderKind::ClaudeCode
&& self.entry.model == "claude-opus-5-5"
&& self.entry.description == PREVIOUS_DEFAULT_OPUS_DEFINITION
{
self.entry.description = DEFAULT_OPUS_DEFINITION.into();
}
let entry = &mut self.entry;
let legacy = self.effort.is_some();
if legacy {
Expand Down Expand Up @@ -593,7 +604,7 @@ impl From<OrchestrateSettingsData> for OrchestrateSettings {
.iter()
.any(|entry| entry.entry.provider == provider && entry.entry.model == model)
};
let has_execution_astra = has_execution(ProviderKind::Codex, "gpt-6-astra");
let has_execution_sol_6 = has_execution(ProviderKind::Codex, "gpt-6-sol");
let has_execution_opus_5_5 = has_execution(ProviderKind::ClaudeCode, "claude-opus-5-5");
(
decisions
Expand All @@ -603,7 +614,7 @@ impl From<OrchestrateSettingsData> for OrchestrateSettings {
data.child_models
.into_iter()
.filter_map(|entry| {
entry.migrate(false, !has_execution_astra, !has_execution_opus_5_5)
entry.migrate(false, !has_execution_sol_6, !has_execution_opus_5_5)
})
.collect(),
)
Expand Down Expand Up @@ -674,6 +685,7 @@ impl OrchestrateSettings {

pub fn builtin_child_definition(provider: ProviderKind, model: &str) -> Option<&'static str> {
match (provider, model) {
(ProviderKind::Codex, "gpt-6-sol") => Some(DEFAULT_SOL_6_DEFINITION),
(ProviderKind::Codex, "gpt-6-astra") => Some(DEFAULT_GPT_6_EXECUTION_DEFINITION),
(ProviderKind::ClaudeCode, "claude-opus-5-5" | "claude-opus-5") => {
Some(DEFAULT_OPUS_DEFINITION)
Expand Down Expand Up @@ -1609,13 +1621,13 @@ mod tests {
["gpt-6-astra", "claude-fable-5-1"]
);
assert_eq!(defaults.child_models.len(), 2);
assert_eq!(defaults.child_models[0].model, "gpt-6-astra");
assert_eq!(defaults.child_models[0].model, "gpt-6-sol");
assert_eq!(defaults.child_models[0].provider, ProviderKind::Codex);
assert!(!defaults.child_models[0].fast);
assert!(
defaults.child_models[0]
.description
.contains("Always dispatch it at low effort")
.contains("Start at medium")
);
assert_ne!(
defaults.child_models[0].description,
Expand Down Expand Up @@ -1751,21 +1763,69 @@ mod tests {
}

#[test]
fn orchestrate_refreshes_untouched_previous_gpt_6_execution_text() {
fn orchestrate_refreshes_previous_opus_5_5_default_text() {
let old_json = format!(
r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#,
serde_json::to_string(OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION).unwrap()
r#"{{"decision_models":[],"child_models":[{{"provider":"claude_code","model":"claude-opus-5-5","description":{}}}]}}"#,
serde_json::to_string(PREVIOUS_DEFAULT_OPUS_DEFINITION).unwrap()
);
let migrated: OrchestrateSettings = serde_json::from_str(&old_json).unwrap();
assert_eq!(
migrated.child_models[0].description,
DEFAULT_GPT_6_EXECUTION_DEFINITION
DEFAULT_OPUS_DEFINITION
);
let customized = old_json.replace(OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION, "mine");
let customized = old_json.replace(PREVIOUS_DEFAULT_OPUS_DEFINITION, "mine");
let kept: OrchestrateSettings = serde_json::from_str(&customized).unwrap();
assert_eq!(kept.child_models[0].description, "mine");
}

#[test]
fn orchestrate_moves_untouched_astra_executors_to_sol_6() {
for text in [
OLD_DEFAULT_GPT_6_EXECUTION_DEFINITION,
DEFAULT_GPT_6_EXECUTION_DEFINITION,
] {
let old_json = format!(
r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#,
serde_json::to_string(text).unwrap()
);
let migrated: OrchestrateSettings = serde_json::from_str(&old_json).unwrap();
assert_eq!(
migrated.child_models,
[builtin_model(
ProviderKind::Codex,
"gpt-6-sol",
DEFAULT_SOL_6_DEFINITION
)]
);
// Customised text, a profile, or a flipped switch keeps Astra.
for variant in [
old_json.replace(text, "mine"),
old_json.replace(r#""enabled":true"#, r#""enabled":false"#),
old_json.replace(r#""fast":false"#, r#""fast":true"#),
old_json.replace(
r#""provider":"codex","#,
r#""provider":"codex","profile_id":"corp","#,
),
] {
let kept: OrchestrateSettings = serde_json::from_str(&variant).unwrap();
assert_eq!(kept.child_models.len(), 1, "input: {variant}");
assert_eq!(
kept.child_models[0].model, "gpt-6-astra",
"input: {variant}"
);
}
}
// An untouched Astra row next to a user-added Sol 6 row is dropped
// rather than duplicating Sol 6.
let both = format!(
r#"{{"decision_models":[],"child_models":[{{"provider":"codex","model":"gpt-6-sol","description":"Mine."}},{{"provider":"codex","model":"gpt-6-astra","description":{},"enabled":true,"fast":false}}]}}"#,
serde_json::to_string(DEFAULT_GPT_6_EXECUTION_DEFINITION).unwrap()
);
let migrated: OrchestrateSettings = serde_json::from_str(&both).unwrap();
assert_eq!(migrated.child_models.len(), 1);
assert_eq!(migrated.child_models[0].description, "Mine.");
}

#[test]
fn orchestrate_preserves_every_customized_sol_shape() {
let customized = [
Expand Down Expand Up @@ -1808,7 +1868,7 @@ mod tests {
}

#[test]
fn orchestrate_migration_does_not_duplicate_existing_execution_astra() {
fn orchestrate_migration_does_not_duplicate_existing_execution_sol_6() {
let old_json = r#"{
"decision_models": [],
"child_models": [
Expand All @@ -1821,7 +1881,7 @@ mod tests {
},
{
"provider": "codex",
"model": "gpt-6-astra",
"model": "gpt-6-sol",
"profile_id": "custom-codex",
"enabled": false,
"fast": true,
Expand All @@ -1832,7 +1892,7 @@ mod tests {
let migrated: OrchestrateSettings = serde_json::from_str(old_json).unwrap();

assert_eq!(migrated.child_models.len(), 1);
assert_eq!(migrated.child_models[0].model, "gpt-6-astra");
assert_eq!(migrated.child_models[0].model, "gpt-6-sol");
assert_eq!(
migrated.child_models[0].profile_id.as_deref(),
Some("custom-codex")
Expand Down Expand Up @@ -1921,11 +1981,15 @@ mod tests {
#[test]
fn orchestrate_settings_patches_deduplicate_within_each_role() {
let mut settings = Settings::default();
let executor = settings.orchestrate.child_models[0].clone();
let peer = settings.orchestrate.decision_models[0].clone();
assert_eq!(executor.provider, peer.provider);
assert_eq!(executor.model, peer.model);
// The same model serves both roles with role-specific guidance.
let executor = builtin_model(
peer.provider,
&peer.model,
DEFAULT_GPT_6_EXECUTION_DEFINITION,
);
assert_ne!(executor.description, peer.description);
settings.orchestrate.child_models[0] = executor.clone();

let mut duplicate = executor.clone();
duplicate.profile_id = Some("another-endpoint".into());
Expand Down
Loading