Skip to content

feat: add Cursor and Grok as native providers (superseded by #584) - #583

Closed
Tryanks wants to merge 25 commits into
mainfrom
feat/native-cursor-grok
Closed

Tryanks wants to merge 25 commits into
mainfrom
feat/native-cursor-grok

Conversation

@Tryanks

@Tryanks Tryanks commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Superseded by #584. This branch was stacked on feat/provider-plugins (#580) and so carried #580's commit plus Grok plugin management. #584 is the same provider work rebased on main without the plugin dependency and is what merged. Kept closed for the record; the Grok plugin management part will return in a follow-up on top of #580.

Summary

Cursor and Grok as native providers (Principle 3). Both CLIs speak ACP with a vendor dialect, so the ACP machinery becomes one shared session with three dialects. See #584 for the final shape.

acp_session.rs            # connection, child lifetime, acks, permission settlement
  Dialect (crate-private)  # launch, auth, new/load/resume/fork, usage, steering, vendor methods
    acp.rs                 # registry ACP agents
    cursor.rs              # cursor-agent acp, cursor/*
    grok.rs                # grok agent stdio, _x.ai/*  (+ grok plugin management via #580's contract)

Evidence

Same as #584, plus Grok plugin tests on recorded CLI output. cargo nextest run --workspace --locked 905 passed, 6 skipped on the merge with main at the time.

Merge Danger

Not merged. See #584.

Tryanks added 25 commits October 3, 2026 11:37
Surface each native CLI's own plugin system as a Tcode capability (Principle 1):
Claude Code through `claude plugin … --json` and Codex through a temporary
`codex app-server`, one bounded process per operation. Tcode persists no plugin
state; the native files stay the truth and the host keeps an in-memory catalog
per profile with Fresh/Stale/Error state.

Contract (crates/agent): ProviderPluginEntry / PluginInstallation (one per
native scope) / DeclaredComponents / PluginAction, with Caps.plugin_management
declaring what each adapter implements. Actions are computed by the host for the
listing's project context; the UI renders only those and never branches on the
provider kind.

Trust: `-y` is never passed and `--dangerously-bypass-hook-trust` is never
used. A marketplace-declared install command surfaces as a PluginChallenge
carrying the CLI's own text and sha256; acceptance is bound on the host to that
pending operation and re-runs with `--accept-command`. Marketplace removal that
would uninstall plugins (Claude) is confirmed first; for Codex, whose
`marketplace/remove` would orphan installed plugins, Remove is offered only when
nothing is installed from it. The management process strips CLAUDECODE,
CLAUDE_CODE_ENTRYPOINT and CLAUDE_CODE_CHILD_SESSION, which make the CLI ignore
`--accept-command`.

Protocol v8: ProvidersStatus.plugins catalogs, eight plugin commands, plugin
operation toasts.

Command cache: `commands-<provider>.json` is now keyed by provider and the
profile's native home and invalidated after a plugin change, so one profile's
plugin commands no longer seed another home's menus. Draft seeding before the
provider starts is kept.

UI: Settings → Plugins, one section per enabled profile, with installed rows per
scope, declared-component details, Add plugin and marketplace dialogs, and the
challenge dialogs rendered from replicated state.

Verified live against claude 2.1.285 and codex-cli 0.159.3 in isolated homes
(fixtures recorded under crates/agent/tests/fixtures/plugins).
Move the protocol machinery out of acp.rs into acp_session: child and
connection lifetime, stderr/EOF handling, prompt delivery acknowledgement,
permission settlement, the fs/terminal client services, MCP and option
mapping, the standard session/update mapping and cancellation. What differs
per agent family is a crate-private Dialect: launch, session establishment
(auth, new/load/resume/fork, initial mode), session/update handling
including replay suppression, turn completion and usage, steering, live
approval-mode changes, and vendor requests/notifications routed by their
literal method names with raw params.

acp.rs keeps registry recipe resolution and the generic ACP policies as the
registry dialect, unchanged except for permission answers: a fixed decision
now selects only an offered allow_once/reject_once option and never falls
back to allow_always/reject_always, so nothing reaches the agent's
persistent allowlist unless the user explicitly picked that option, and an
explicit Option(id) must be one the agent offered. Pending permissions and
questions are also settled before session/close on shutdown.
A failed handshake (signed-out agent, broken initialize, early EOF) was
sent to the caller before the child process was killed. A caller that
exits on the error, such as the probe, left the agent running, reparented
to init; cursor-agent does not exit on stdin EOF, so it lingered. The
actor now kills and reaps the child first and reports the failure after.
ProviderKind::NATIVE is the one ordered collection of natively maintained
providers. The runtime's catalog/version/profile loops, the secret-name
index, the Settings provider cards and the composer's picker rail consume
it instead of keeping their own arrays; the Settings cards therefore follow
the same Claude-first order as the composer and the runtime. The provider
dialog copy becomes an exhaustive match, so a missing provider is a compile
error instead of a panic.
…sion

Both CLIs speak ACP as their only bidirectional protocol, so each is a
dialect of the shared ACP session rather than a new JSON-RPC engine.

Grok (`grok [--permission-mode M] agent [--model M] [flags] stdio`, home
GROK_HOME) authenticates with the non-interactive xai.api_key method,
creates sessions with session/new, resumes with session/resume, forks with
_x.ai/session/fork, drops updates flagged _meta.isReplay, steers through
_x.ai/interject (accepted when Grok echoes _x.ai/session/interjection),
answers _x.ai/ask_user_question through a blocking user-input request, and
lists its models from initialize's _meta.modelState.

Cursor (`cursor-agent [flags] acp`, home CURSOR_CONFIG_DIR) initializes
with its _meta client capabilities, never calls authenticate (it opens a
browser when signed out) and reports a missing login instead, and resumes
with session/load. Its sessions are not available yet: once a session would
be established, start fails with an explicit error until the cursor/*
requests are implemented and verified against a signed-in account.

The shared session gains what the dialects need: a blocking user-input
request that interrupt and shutdown cancel without holding the dispatch
loop, session state access for vendor request handlers, and a sessionless
query (spawn, initialize, ask, tear down) for model catalogs.

Identity and compatibility: wire and settings names are `cursor` and
`grok`; the two built-in profile ids are `native:cursor` and `native:grok`,
which no user profile slug can take, so an existing user profile named
"Cursor" or "Grok" keeps resolving to itself (Settings::builtin_profile_id
is the one owner, also used for thread colors and secrets). The registry's
`cursor` and `grok-build` agents are hidden: installed copies stay
resumable for their existing sessions and removable, but are no longer
offered for, or accepted as, the agent of a new session. PROTOCOL_VERSION
becomes 9 because ProviderKind has no unknown fallback on the wire.

Wire-option selections are persisted for every provider whose options come
from the wire, not only for registry ACP agents.
…vals

Three additive Dialect hooks, each defaulting to today's behaviour:

- client_services(): whether `initialize` offers tcode's fs/* and
  terminal/* services. An agent offered them routes its file and shell
  tools through tcode instead of its own implementation.
- owned_config_options(): session config options (by configId) that
  tcode already controls elsewhere, such as the composer's model picker.
  They stay in the option registry (the session's model is still read
  from it) but are left out of ProviderOptions.
- auto_approves(tool_call): whether a permission request is approved
  without asking, for an approval mode the agent does not apply itself.
  The shared session then selects the request's allow-once option only;
  without one the user is asked, so no automatic answer ever records a
  persistent allow/reject rule.

config_selection() reads a persisted wire selection (`acp:cfg:<id>`) so
a dialect can carry it into a new process at launch.
Verified against grok 1.0.46 with a scripted model backend (no xAI account
or live model).

- Grok runs its own file and shell tools. Offered tcode's client services
  it sends every command to terminal/create as one `/bin/bash -lc '…'`
  string, which cannot be spawned, and every read and edit to fs/*, which
  tcode's cwd containment rejected for Grok's canonical paths; so no Grok
  shell command or edit worked.
- Tool items: Grok's shell tool resends the whole output so far with
  every update, and the item shows exactly that, with the exit code only
  once the command has finished. The model's description of a command is
  no longer shown as its output, and Grok's typed rawOutput never becomes
  display text (a command printing nothing showed it as JSON). MCP calls
  through use_tool show the MCP result text.
- Usage: each model call's response_completed is a context observation
  (TokenUsage against the model's totalContextTokens); the session/prompt
  result's _meta is the turn's usage, reported once on TurnCompleted.
  turn_completed and prompt_complete notifications never complete a turn.
- Failures: a failed turn shows Grok's own message from the -32603 error
  data; retry_state retrying becomes a non-fatal error like Codex's
  willRetry; image_dropped notes become a warning; background_tasks maps
  to BackgroundTasksChanged (running tasks); auto_compact_started and
  auto_compact_completed map to ContextCompacted.
- Plan mode is native: Grok's default/plan modes (which session/set_mode
  takes and current_mode_update reports, but Grok does not list) back the
  Build/Plan toggle, including Grok entering plan mode itself. A continued
  session is told its mode, because Grok keeps plan mode across processes
  without reporting it. _x.ai/exit_plan_mode becomes a ProposedPlan and is
  answered `request_changes`, which keeps Grok planning: the decision is
  the user's, through the plan flow; approving would let Grok implement
  first, and holding the request would keep the turn open. An unanswered
  request cancelled the whole turn.
- Approval modes: `--permission-mode acceptEdits` still asks before every
  edit, so AutoAcceptEdits launches `default` and approves edit-kind
  requests once itself; ReadOnly and Supervised launch `default` (reads
  unprompted, mutations ask); FullAccess launches `bypassPermissions`,
  which survives leaving plan mode.
- Model and effort: the composer's model picker owns the model (the wire
  `model` option is no longer a second selector); reasoning effort stays
  a wire option, applied live, and is passed as `--reasoning-effort` when
  a session starts again.

The probe gains `--mcp <name> <url> <token>`, and `--effort` sets Grok's
wire option. The recorded fixture (tests/fixtures/grok) is replayed by a
stand-in Grok through the production session path.
`grok models` exits 0 and opens with the CLI's authentication state.
Grok 1.0.46 prints `You are not authenticated.` (reported Unauthenticated)
or `You are using XAI_API_KEY.` (Authenticated, "xAI API Key"; the key is
configured, not validated). Any other first line, such as a browser
sign-in whose wording was not observed, stays indeterminate.
Grok's native plugin system behind the provider plugin contract, through
its non-interactive CLI only, one bounded process per command (null stdin;
with an inherited non-terminal stdin `grok plugin` fails with ENXIO).
Verified against grok 1.0.46 in throwaway HOME/GROK_HOME, each operation
followed by a native read (fixtures and provenance in
tests/fixtures/plugins/grok).

- Listing: `plugin list --json --available` and `plugin marketplace list
  --json`, plus `plugin details` per installed plugin (its description and
  native text). Marketplace plugins are `name@marketplace`, as `install`
  takes them; the other commands take the plain name. Grok installs for the
  user only.
- Install: without `--trust` Grok explains what installing activates and
  prints the command that proceeds. That becomes the existing AcceptCommand
  challenge (Grok's command and text, bound by their SHA-256), and
  `--trust` is passed only on a re-run whose challenge still hashes to the
  accepted value; anything else is shown again.
- Update is offered for marketplace installs (it refreshes a git
  marketplace's checkout itself); a path install is a live link that
  `update` leaves as it is.
- Uninstall is not offered for a plugin of a multi-plugin repository:
  Grok then needs `--confirm` and removes them all, and the contract's
  uninstall carries no consent. `--confirm` and `--force` are never passed.
- Marketplaces: add and remove. Removal uninstalls the marketplace's
  plugins, which the listing reports so the host confirms first.
- Enable/Disable are left out: no `grok plugin` read reports enablement,
  and `grok inspect` (text and --json) reports a disabled plugin as
  enabled; a session confirmed the disabled plugin did not load.

ApplyNote gains ReloadOrNextSession: against a running session, an install
(skills, commands, hooks, an MCP server) and an update both applied after
`/reload-plugins`, and in sessions started afterwards, but not before.
`Setup::load` cleared its replay flag only once the awaiting task resumed,
and `block_task` lets the dispatch loop read the next message first. An
update the agent sends as soon as it has answered `session/load` was then
taken for replay and dropped: Cursor's slash commands after a resume
(`available_commands_update`, sent right after the response) were lost.

The flag is now cleared in an ordered response callback, which holds the
dispatch loop until it has run, so nothing after the response can be read
as replay. The Cursor client's load test sends the response and the next
update in one write and fails without this.
`configOptions` in `session/set_config_option` responses and
`config_option_update` is the agent's full set, but the option registry only
upserted, so an option the agent stopped offering stayed in the composer.
An agent whose options depend on the model (Cursor's model parameters: a
new model brings its own and drops the last one's) kept showing the
previous model's parameters, and setting one failed. A config list now
replaces the config options; the session modes are kept.
…t parse

`session/update` notifications went to a typed handler. One whose
`sessionUpdate` kind the schema does not know (Cursor's
`subagent_spawned` and `subagent_state_update`) failed to parse and the
SDK dropped it with a warning, so no dialect could see it. Notifications
are now received untyped: a `session/update` the schema parses goes to
`Dialect::session_update` as before, and anything else, a vendor update
kind included, to `Dialect::notification` with its literal method and raw
params, in arrival order.

`State::new` and `State::flush_text` become crate-visible so a dialect can
run the standard mapping for an agent's own sub-sessions.
Cursor sessions no longer stop at "not available yet". What follows was
read from the shipped source of cursor-agent 2026.10.01-e373342; only the
signed-out handshake has been observed from the binary. No signed-in
session has run.

- Catalog: `cursor/list_available_models`, whose model values are what
  `session/set_config_option` accepts. Signed out, it fails with the same
  sign-in guidance as a session instead of an empty catalog.
- Sessions: `session/new`, or `session/load` for a resume, whose failure
  is an error and never a fresh session; its replay is dropped.
  `authenticate` is never sent (signed out, it starts a browser login).
  The composer's model is applied with `session/set_config_option` after
  either, since a loaded conversation comes back on its last model; a
  model Cursor rejects fails the start. FullAccess launches with `--force`.
- Turns: a run Cursor reports as its failure text (`"\n\nError: …"` or an
  account action, then `end_turn`) completes as failed with the text kept;
  a tool is judged by its `rawOutput` (exit code, `error`, `rejected`,
  `permissionDenied`), since Cursor reports every finished tool
  `completed`, and a shell result shows its output as text.
- `cursor/ask_question` becomes a blocking question; answers go back as
  option ids, and a typed answer Cursor cannot take is reported, not
  dropped silently. `cursor/create_plan` becomes the proposed plan and is
  answered `accepted` without a plan URI, which has Cursor save the plan
  as its own CLI does without asking; the user builds or dismisses it
  through the proposed-plan flow. `cursor/update_todos` becomes the plan,
  merged by id when asked to. `cursor/task` and `cursor/generate_image` are
  acknowledged.
- Subagents (`_meta.subagents`): a task tool is a Subagent item;
  `subagent_spawned`, `subagent_state_update` and the subagent session's
  own updates become its status and its child items.

The tests drive `start` and `list_models` against a stand-in that relays
its stdio to the test, which answers from the recorded signed-out lines
and from source-derived messages labelled as such
(`tests/fixtures/cursor/PROVENANCE.md`). The probe gains `--binary`,
`--model` and a `read_only` approval mode.
`cursor-agent status --format json` reports the stored login. Signed in, the
provider is ready with the account's email. It says `unauthenticated` even
while CURSOR_API_KEY or CURSOR_AUTH_TOKEN signs the CLI in (observed with
2026.10.01-e373342), so it counts as signed out only when the profile's
environment and the inherited one set neither; otherwise the state stays
indeterminate. A key passed as a launch argument is not visible to the
probe.
Brings the Grok client and the shared ACP dialect hooks it added
(`client_services`, `owned_config_options`, `config_selection`,
`auto_approves`) under the Cursor client.

Conflicts, resolved keeping both sides:
- examples/probe.rs: the usage text lists both branches' flags
  (`--binary`, `--model` and `--mcp`).
- lib.rs: the ApprovalMode doc keeps Grok's new line and adds Cursor's.
- services/provider_probe.rs: both the Cursor (`cursor-agent status`) and
  the Grok (`grok models`) probes, with their parsers; only ACP stays
  without an auth probe.
Built on the shared ACP dialect hooks:

- Approvals (`auto_approves`): under ReadOnly tcode grants Cursor's read
  and search permission requests, and under AutoAcceptEdits its file-change
  requests, each with the request's allow-once option; everything else is
  asked, and nothing automatic ever picks Cursor's allow-always, which
  writes its permanent allowlist. A web fetch is asked about even under
  ReadOnly: it can carry data out. The kinds are Cursor's own
  (3351.index.js 2684-2753, 857-887).
- The model (`owned_config_options`): the composer's model picker owns
  Cursor's `model` option, so it is no longer surfaced as a second
  selector; `SessionStarted` still reports it. Its parameters stay
  provider options.
- Parameters (`config_selection`): a new process (resume, model change,
  approval-mode change) re-selects the model parameters the session chose
  if the model still offers them, since Cursor starts on the ones it last
  saved for the model, which another session may have changed.
- Client services (`client_services`): not offered. Cursor's ACP server
  reads `clientCapabilities` only for `_meta.parameterizedModelPicker` and
  subagents (1148, 1496, 3410) and runs its file and shell tools itself;
  declining keeps that so.
`--linger <seconds>` keeps the session after a turn completes until no turn
has run for that long, then shuts down and waits for the close.
`--follow-up <seconds> <prompt>` sends another turn that long after the
first completes, whether or not a turn is running then. Together they show
a turn an agent starts on its own and a turn sent during one.
When a background command finishes, Grok runs a prompt of its own
(`task-completed-<task id>`) reporting it. No `session/prompt` answers it,
so its events arrived with no turn around them. It now opens and completes
like any turn, following the contract the other providers already use for
turns the user did not send: Claude Code synthesizes TurnStarted for its
self-invoked turn and its `result` completes it; the runtime counts any
TurnStarted as a running turn, queues or steers sends meanwhile, and
dispatches the queue on TurnCompleted.

Shared session (acp_session):
- State::begin_agent_turn opens a turn the agent started; while another
  turn is still open it opens as soon as that one completes.
- State::end_turn completes a turn from the agent's own signal if it is
  still the open one. A sent turn's `session/prompt` result then completes
  nothing; finish_turn only completes the turn it names while that turn is
  open, and the turn id now travels with the result instead of a loop-local
  slot (a send overwriting that slot used to complete the wrong turn and
  then hit its `expect`).
- Interrupt acts on whatever turn is open, including an agent-started one.
- A turn the client sent while an agent-started turn ran takes its place,
  as a send during Claude's self-invoked turn does: one completion, and the
  agent turn's streamed text is closed when it ends.

Grok:
- `_x.ai/queue/changed` reports each running prompt; one the client did not
  queue opens an agent turn under Grok's prompt id.
- `turn_completed` ends the turn of the prompt it names, in order with
  everything else: when a task finishes during a sent turn, Grok starts its
  own prompt before the sent turn's `session/prompt` result arrives. The
  status comes from its stop reason (error text from `agent_result`), and
  the usage is the turn's total on top of the last model call's. The
  `session/prompt` result remains the fallback for a prompt without one.
- A drop to no background tasks is withheld while a sent turn is open, as
  Claude withholds it until its self-invoked turn, so the runtime does not
  take the process for idle before Grok's own turn.
- A backgrounded command's card stays completed: Grok streams its later
  output as running updates and never completes it again.

Verified against grok 1.0.46 with a scripted model backend through the
probe: a task finishing after a turn, during a turn, an interrupt during
Grok's turn, and a send during it. agent_started_turn.jsonl records the
during-a-turn case followed by another sent turn; its replay fails without
the agent turn, with a second completion, with the count not withheld, or
with the background card reopened.
Brings in the probe's `--linger` and `--follow-up` and the turns Grok
starts itself: a sent turn's outcome now carries its turn id, and
`finish_turn` completes a turn only while it is still the open one
(`State::complete_turn`, with `begin_agent_turn` and `end_turn` for
agent-started turns).

Conflict, resolved keeping both sides: examples/probe.rs, whose usage text
now lists `--binary`, `--model`, `--mcp`, `--linger` and `--follow-up`.
acp_session.rs merged cleanly beside this branch's untyped `session/update`
routing, ordered `session/load` response and full config lists; Grok's
recorded wire tests pass through them.
A comment is kept for a constraint the code cannot express. The npm
launch recipe is the registry's contract, not another client's; the
config-option ids are explained by what reads them rather than by the
protocol revision that introduced categories; the EOF note names the SDK
behaviour it works around without a stale version; and the approval
mapping and update mapping docs lose their history and test notes.
Cursor sessions run now, so the getting-started step lists `cursor-agent`
among the CLIs to have on PATH, and a bug report names any of the six native
agents rather than only Claude Code and Codex.
The UserInputQuestion id doc lists each provider's answer key; Cursor's
`cursor/ask_question` questions carry ids, which the Cursor client uses.
Cursor and Grok are native providers that speak ACP, so approval option
lists and `SetOption` routing are not the registry agent's alone; options
are set through `session/set_mode` or `session/set_config_option`, never
`session/set_model`.
The provider dialog pushed pi's "Native approvals" switch inside the
home-path block, so every provider with a home override (Claude, Codex,
Cursor, Grok) showed a switch whose help text describes pi and whose
setting nothing reads for them. Its only consumers, the runtime's launch
approval mode and the composer's approval-mode availability, read it behind
`downgrade_approval_without_native_approvals`; the switch is now gated on
that capability, and the home field stays under `home_path`.
@Tryanks
Tryanks marked this pull request as draft October 3, 2026 08:32
@Tryanks

Tryanks commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Superseded by the PR from feat/native-providers, which is based on main and independent of #580.

@Tryanks Tryanks closed this Oct 3, 2026
@Tryanks Tryanks changed the title feat: add Cursor and Grok as native providers feat: add Cursor and Grok as native providers (superseded by #584) Oct 6, 2026
@Tryanks
Tryanks deleted the feat/native-cursor-grok branch October 6, 2026 10:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant