Skip to content

feat(playbooks): support Aurora review and publication loops - #164

Open
BastiHu wants to merge 1 commit into
codex/playbooks-08-authoringfrom
codex/playbooks-09-aurora-loops
Open

BastiHu wants to merge 1 commit into
codex/playbooks-08-authoringfrom
codex/playbooks-09-aurora-loops

Conversation

@BastiHu

@BastiHu BastiHu commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

The external Aurora Playbooks need review feedback and approval-controlled publication loops.

Extend YAML execution and publication adapters to collect PR feedback, revisit review and rework phases, and preserve explicit publication approval. Keep diagnostic persona execution read-only and document the external Playbook source.

External definitions are proposed separately in First-horizon/agent-skills#1: https://github.com/First-horizon/agent-skills/pull/1.

Stack 9/9. Depends on #163. Merge in order, retargeting to j5/main as predecessors land. Full-stack testing branch: codex/playbooks-09-aurora-loops.

Validation at the integrated stack tip: 179 server/shared tests and 104 web tests passed; server and web typechecks passed. Each regenerated runtime manifest was checked. The adapter slice independently passed 92 tests, including the parser regressions.

Review order:

  1. 01-persona-launch
  2. 02-engine
  3. 03-adapters
  4. 04-yaml-library
  5. 05-runtime-api
  6. 06-client-state
  7. 07-run-ui
  8. 08-authoring
  9. 09-aurora-loops

Model: GPT-5 · Harness: Codex

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 effective changed lines (test files excluded in mixed PRs). labels Sep 16, 2026
@bryantderosier
bryantderosier added this pull request to stack #165 September 16, 2026 13:12
@Jacksondr5

Copy link
Copy Markdown
Owner

Hey Bastian, here are some of my thoughts after reviewing the PR:

Control/Lifecycle

I think we should flip the order of control around. From what I can tell, in this stack, the playbook spawns agents and creates new subagents to execute tasks. That puts the playbook above the agent in terms of lifecycle and control.

I think it should be the opposite, an agent can run multiple playbooks sequentially through its lifespan. Or, it can be spawned specifically to run a playbook by itself. That means the playbook lifecycle is <= the agent lifecycle. I think having a playbook manage the lifecycle of agents is an interesting idea, but I think we should play with playbooks in the other form for a while before we expand their responsibility. A more controlling/complex feature tends to encode more opinions and be less flexible, so starting with the simpler one lets us evolve it as we discover more.

Code Steps & Transitions

This goes along with the simplicity argument I made above, but I think we should start out playbooks as a simpler tool that allows agents to execute complex workflows, but doesn't try to be a workflow engine.

I think, with where we are in the new cycle of agent capabilities that's come out with Astra and Fable, we are at a point where people are better served by just using the top models again (and not trying to bring in the older models that need more steering and coordination). Therefore, I think we all should focus on UX and tooling that helps the most capable agents when we're building J5.

From a UX standpoint for Playbooks, I think that means:

  • Showing the user clearly what step an agent is on at a glance
  • Making it really easy for them to spin up a playbook via slash commands or native language prompts (then the agent uses an MCP to trigger the playbook)
  • Turn a workflow that they already have with an agent into a playbook. IE, letting the agent create a playbook with them and edit it

From the agent tools perspective, I think that means:

  • Mitigating the context compaction problem by injecting prompts for the next step when the agent asks for it
  • Providing tools that give agents in a crew the ability to understand what step they're on and to progress the steps for the group in an easy but controlled manner. I think this is actually going to be pretty tricky to do because we have to figure out what the agent's behavior is here and create the right tools for it. For example, we might notice that some agents are really eager to progress to the next step before everyone in the crew is ready, and so maybe we have to do something like only allow the captain to progress the step. Or maybe we notice that some agents are really bad at remembering what step they're on, and so we need to inject more often into their context what step they're on. So, we'll want to iterate on this with live testing after merge.

Simplicity

This isn't just a normal coding best practice, I think given how fast T3 is moving and how they seem to be going into some of the areas we are building, keeping things simple is going to make it easier for us to integrate as they evolve. And we 100% want to be able to do this so we can spend more time on features and less time on the stuff they've already solved.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L 100-499 effective changed lines (test files excluded in mixed PRs). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants