Skip to content

test(expo): add a verify skill that drives the expo-native fixture on a local or borrowed device - #10087

Draft
mikepitre wants to merge 22 commits into
mike/expo-verify-hostfrom
mike/expo-verify-remote
Draft

mikepitre wants to merge 22 commits into
mike/expo-verify-hostfrom
mike/expo-verify-remote

Conversation

@mikepitre

@mikepitre mikepitre commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds verify-clerk-expo, a skill and CLI that an agent uses to prove a @clerk/expo change on an iOS simulator or Android emulator, with video, screenshots, and app state as evidence. It drives the expo-native fixture through the launch inputs from #10052, against a Clerk application that it creates for the worktree and deletes afterwards. The device runs on the developer's Mac. On a machine that cannot run one, such as a Linux machine or a cloud agent's sandbox, the CLI borrows a simulator or emulator on a CI runner and drives it the same way.

The skill is at .claude/skills/verify-clerk-expo/, the CLI is bin/control-clerk-expo inside it, and .cursor/skills/verify-clerk-expo is a symlink so Cursor reads the same files. packages/expo/AGENTS.md has the short version for agents, and the root AGENTS.md points to it. This pull request holds what #10053 held, and #10053 is closed.

Where to start reading: SKILL.md, then src/host.ts and src/fixture.ts, then .github/workflows/verify-remote.yml. About 3,700 of the 20,000 added lines are written for this repo. The rest are the lockfile and byte copies of the verification skills in clerk/clerk-ios (clerk/clerk-ios#629) and clerk/clerk-android (clerk/clerk-android#1046), where they are reviewed: src/core/, src/platform/ios/, src/platform/android/, specs/fixtures.ts, testing/, and every test file but test/host.test.ts, test/freshness.test.ts, and test/remote-host.test.ts. src/core/MANIFEST lists the hashes of the core files, and a unit test and control-clerk-expo doctor fail when one drifts.

The commits are in reading order, the local device first and the borrowed device after it:

  1. Six commits add the skill for a local device: the package and repo wiring, the lockfile, the copies, this repo's own code and specs (src/host.ts, src/fixture.ts, src/freshness.ts, specs/native.ts, ten spec files for six features under specs/golden/), the docs, and a CI job for the unit tests.
  2. Eight commits move it to e2e 0.18.0 and @e2e-dev/mobile 0.10.0 and take what came with that in the copies: a plain tap in host.tap, run --retries <n>, run --github-report, junit.xml beside each report, a host.fill that checks the text reached the field, unit tests that cannot write to the repository they run in, and an Android emulator with animations and system error dialogs off.
  3. Three commits add the borrowed device: the copies (src/core/remote/ and its tests), then this repo's own code (.github/workflows/verify-remote.yml, src/platform/session-device.ts, the second build product in src/fixture.ts, the two remote backends in src/host.ts), then the docs.
  4. Two commits pin the iOS runtime of the simulator a session creates and make the free GitHub-hosted runners the default.
  5. Two commits take the copies as clerk-ios trimmed them (settings in a JSON file beside a spec, one application that serves one run at a time, a smaller settings check, one Backend API host) and read a session's request on a free runner whatever the session's label.

What the CLI does:

  • up --platform ios|android builds the fixture as a Debug dev client, creates the worktree's Clerk application, takes a device the skill owns (a cloned simulator or a read-only emulator), installs the app, and starts tsdown --watch in packages/expo and expo start on a port fixed per device.
  • run <feature> runs that feature's specs and writes video.mp4, screenshots, and every verify.state the app reported under .claude/skills/verify-clerk-expo/.verify/runs/<run-id>/. The features are native-auth-view, user-button-and-profile, custom-flow-sign-in, custom-flow-sign-up, token-cache-persistence, and native-js-sync.
  • down releases the devices, stops Metro and the watch build, and deletes the application with every user in it. It keeps the evidence.
  • doctor only reads. It creates no file, device, or application.
  • A change under packages/expo/src reaches a local device on the next run with no native build. Before the specs start, run waits until the watch build has caught up and Metro serves a bundle built from the current dist, and fails with NOT_READY if it never does. A change to a native input (the list is nativeInputs in src/fixture.ts) rebuilds the dev client.
  • --backend auto is the default: a local device when the machine can run it, a runner when it cannot. --backend local|remote forces one.

How a borrowed device behaves:

  • The session builds the pushed commit, never the working tree. An uncommitted change to a build input, or a HEAD that GitHub does not have, fails with the fix git commit or git push.
  • A runner gets a Release build with expo-dev-client left out and the JS embedded, because it cannot reach Metro on the developer's machine. A JS edit reaches the app by commit, push, run, and the held session rebuilds.
  • Evidence lands in .verify/runs/<run-id>/ as for a local device.
  • doctor with the remote backend does git and REST reads only. doctor --live starts one probe run and one short session to prove the rest.

Things a reviewer may ask:

  • The skill uses no shared test instance. Each worktree gets one application in a workspace that holds nothing else (org_3KHungJxbvIscuSvy8oos5MHAli), configured from src/core/instances/base.json. A spec that needs other settings declares them in a <name>.settings.json file beside it. No golden spec does. The credential is a Platform API key from the environment or from a 1Password reference kept on the machine. A machine with none fails doctor with one fix line.
  • CI does not run these specs in this pull request. The tests in integration/tests/expo-native (ci(e2e): Replace maestro with e2e in the expo native integration tests #10032) stay the regression gate. The Verify Skill Tests job in ci.yml runs the skill's unit tests and typecheck on a Linux runner, with no device and no secret, when a pull request changes the skill or a path its tests read.
  • The skill installs e2e 0.18.0, @e2e-dev/mobile 0.10.0, and @e2e-dev/github 0.4.0 with npm ci from its own lockfile, and needs Node 24.8 or newer. It is outside the pnpm workspace, and the copied files are tested against those versions. Nothing is installed globally, on a developer's machine or on a runner.
  • Runners and cost. A session runs on macos-26 for iOS and ubuntu-24.04 for Android, GitHub-hosted labels that cost nothing for a public repository. They are slow. up was ready after 26 and 27 minutes on iOS in two sessions and after 11 to 14 minutes on Android in three, nearly all of it the build. --runner blacksmith-6vcpu-macos-26 or --runner blacksmith-8vcpu-ubuntu-2204 leases one of the two Blacksmith labels this repo already uses, which is ready in 5 to 7 minutes and is billed by the minute. A session stops itself after 15 minutes without a call from the CLI and always after 60, and down ends it.
  • The free Mac is slow to build. Both iOS session builds on macos-26 took 24 minutes, which was one minute inside the limit the CLI gave a session to be ready. The last commit raises that limit from 25 to 40 minutes. No session has run on the free Mac since that commit. In one of the two sessions the first run also lost its first test to a launch that took more than 30 seconds, and the second run passed.
  • On a GitHub-hosted label a session writes no Gradle cache and no pnpm cache for macOS, so it builds cold every time. The repository's Actions cache is at its limit, and a session per branch would push other entries out.
  • Who can start a session. The workflow runs on workflow_dispatch and on a push to verify-remote/**, so only someone with write access. It has no pull_request trigger. A machine that may not dispatch pushes a verify-remote/... branch whose commit message is the request, and the cleanup job deletes that branch, which is the one use of contents: write.
  • What reaches the runner and its public log. The driver creates the Clerk application, seeds users, and mints sign-in tickets. The runner gets a publishable key and a ticket as launch arguments, and never a secret key. The request carries only the SHA-256 of the session's bearer token. The job uploads no artifact, and every route on the tunnel requires the bearer.
  • The simulator of a session is an iPhone 17 Pro on iOS 26.5. The workflow names that runtime, so a new runner image cannot change the iOS version, and a runner without it fails with the list of runtimes it has. integration/tests/expo-native/boot-ios-simulators.sh, the script the e2e tests from ci(e2e): Replace maestro with e2e in the expo native integration tests #10032 use, still sets the simulator's keyboard preferences.
  • host.fill reads the focused field back after it types and types once more if the field is still empty. It cannot do that for a password field, which withholds its value, or on iOS for the fixture's own React Native fields, where an empty field reads as its placeholder. Those are typed once.
  • The auth specs use +clerk_test users. Four specs type a password or the test code 424242 and are tagged form-entry: custom-flow-sign-in/complete, custom-flow-sign-up/request-code, custom-flow-sign-up/complete, and native-auth-view/complete. They ran in CI before the move to e2e 0.18.0 (ci(repo): run the verify-clerk-expo device specs on pull requests #10090 has the runs). All four pass on iOS, and on Android the two that type an email code failed because the typed code did not reach the field. They have not run since, so the host.fill that reads the field back is unproven on them. Everything else runs with --skip form-entry and passes on local and borrowed devices, 7 tests on iOS and 6 on Android. Each sign-in flow is covered up to its code screen that way, and sign-up has no spec that runs without typing.
  • One spec is tagged known-bug and skipped unless --include known-bug is passed: the close button of an inline AuthView does not fire onDismiss.
  • Android specs find native views by text, because the pinned clerk-android release has no test tags. iOS specs use the clerk.* accessibility identifiers.
  • A worktree can hold an iOS and an Android device at once, on a runner as on a local device. They share one application, and it serves one run at a time. A run that starts while a run on the other platform is driving fails with DEVICE_BUSY and the fix to let that run finish. Before that was enforced, a ticket sign-in failed with resource_not_found in two of four tries with both platforms running specs at once, and never with one at a time. The cause is not known.
  • On Android the Release build names :app:createBundleReleaseJsAndAssets --rerun before assembleRelease. Gradle treats that task as up to date when the only change is inside a linked workspace package, and without the rerun a rebuilt app kept the old JS.
  • The fixture is built locally on macOS only. On Linux with KVM and an Android SDK, --backend auto picks the local emulator and the build then fails with UNSUPPORTED and the fix rerun with --backend remote.
  • .prettierignore skips the copied .ts files so the pre-commit hook does not reformat them.
  • Not tried: the idle stop and the cap of a session with this repository's workflow, which are the shared code's and were observed in clerk-ios and clerk-android; a rebuild on a held session after a change to a native input; and a run from inside a cloud agent's sandbox.

Checklist

  • pnpm test runs as expected.
  • pnpm build runs as expected.
  • (If applicable) JSDoc comments have been added or updated for any package exports
  • (If applicable) Documentation has been updated

Type of change

  • 🐛 Bug fix
  • 🌟 New feature
  • 🔨 Breaking change
  • 📖 Refactoring / dependency upgrade / documentation
  • other: test tooling

🤖 Generated with Claude Code

@changeset-bot

changeset-bot Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 8e4f27c

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
clerk-js-sandbox Ready Ready Preview Oct 6, 2026 7:36pm UTC
swingset Ready Ready Preview Oct 6, 2026 7:36pm UTC

Request Review

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

mikepitre and others added 6 commits October 6, 2026 02:41
The skill is outside the pnpm workspace and installs its pinned e2e with npm from its own lockfile.
.cursor/skills/verify-clerk-expo is a symlink to the skill, and .prettierignore keeps the pre-commit hook
off the files that are copied from other repositories.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
src/core, src/platform/ios, specs/fixtures.ts, testing, e2e.config.ts, and these tests are byte copies of
clerk/clerk-ios. src/platform/android and test/android.test.ts are byte copies of clerk/clerk-android.
src/core/MANIFEST pins the core.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… check, and the golden specs

src/host.ts builds the expo-native fixture as a Debug dev client, starts the watch build and one Metro per
lane, and waits until Metro serves current JS before specs start. specs/golden holds ten spec files for
six features.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…them

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mikepitre and others added 7 commits October 6, 2026 12:39
…dev/mobile 0.10.0

The new mobile engine brings agent-device 0.21.22. e2e 0.18.0 refuses Node 24.0 to 24.7, so `engines` and the
`node` check of `doctor` ask for 24.8.0 or newer on 24. The lockfile is regenerated with npm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
agent-device 0.21.22 no longer refuses a tap on a control of a native view, so `host.tap` is `locator.tap()`
with the assertion timeout and the helper that tapped the middle of a node is gone. `host.fill` still taps the
field and types through agent-device, because a native text field on iOS shows no text input until it has focus.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pt of a flaky test

`run --retries <n>` runs a failed test again, up to n more times. The default is 0. A test that passes on a
retry is printed and recorded as flaky with the error, failure page, and screenshot of its failed attempt, and it
does not fail the run. The e2e config no longer sets retries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gh @e2e-dev/github

After the run is sealed, the reports of every settings group are merged and handed to the reporter once, with
the platform as its key, so one job gives one job summary and one pull request comment per platform. The switch
is off by default. A GitHub token in the environment is treated as a secret of every run: output is redacted, a
file that holds it is marked tainted, and a tainted run is never reported.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After typing, `host.fill` reads the focused text input back. If it is still empty after two seconds, the helper
taps it and types once more, and it fails with "the text never reached the field" if it is empty again. It
confirms nothing when it cannot tell: a password field withholds its value, and on iOS an empty React Native
field of the fixture reads as its placeholder.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…repository

A test that runs `git init` under a git hook or a rebase exec inherited `GIT_DIR` and wrote to the real
checkout. `testing/git-env.ts` clears the variables that name a repository, every test file imports it first,
and `test/git-env.test.ts` fails when a test file does not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mikepitre
mikepitre force-pushed the mike/expo-verify-remote branch from 2d05137 to 43ca8cc Compare October 6, 2026 16:55
mikepitre and others added 6 commits October 6, 2026 13:35
…id lane emulator

A lane now boots with the window, transition, and animator animation scales at 0, so a spec sees each screen
change at once. It also sets `hide_error_dialogs`: on a slow CI runner the launcher can stop responding while the
emulator boots, and agent-device refuses every tap while that dialog covers the app. The local backend reads the
settings back and fails `up` when one did not take. A borrowed emulator boots its lane through the same code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…k-android

src/core/remote, the launcher, and these tests are byte copies of clerk/clerk-ios at its remote device
layer. src/platform/android is a byte copy of clerk/clerk-android at the same layer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ify skill

verify-remote.yml runs a session for either platform. A session builds the pushed commit as a standalone
Release app with the JS embedded, because a runner cannot reach Metro on the developer's machine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n creates

The session created its simulator on the runtime that matched the default Xcode of the runner image, so an image
update could change the iOS version under the specs without a change in this repository. The workflow now names
the runtime. On a runner without it, the step fails and lists the runtimes the runner has.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A remote session now starts on `macos-26` for iOS and `ubuntu-24.04` for Android, which GitHub hosts at no cost
for a public repository. `--runner <label>` and `VERIFY_REMOTE_RUNNER` still choose a Blacksmith label, which is
faster and billed by the minute.

The free Mac is slow, so the workflow waits longer: five minutes for the cloudflared download, ten for the
tunnel, and as long as the step allows for the first boot of the simulator. On a GitHub-hosted label the session
writes no Gradle cache and no pnpm cache for macOS, because the repository's Actions cache is already at its
limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mikepitre and others added 2 commits October 6, 2026 14:15
The shared files are byte copies of clerk/clerk-ios after eight cuts there. What changes for this skill:

- A spec declares other instance settings in a JSON file beside it, `<name>.settings.json`. An exported
  `instanceSettings` is refused, and so is any other JSON file under `specs/`. No golden spec declares any.
- A worktree has one Clerk application, which serves one run at a time. A run that starts while a run on the
  other platform is driving fails with `DEVICE_BUSY`.
- The settings check compares the 61 leaves a spec cannot run without and the leaves a settings file declares.
- The Backend API is called on api.clerk.com only. A 401 to an instance's own key fails with its cause.
- The switch for e2e's AI judge and the doctor check for a daemon of a deleted install are gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The short job that reads a borrowed-device request ran on a Blacksmith label when the session's own label was a
Blacksmith one. It now runs on `ubuntu-latest` for every session. `VERIFY_REMOTE_PLAN_RUNNER` still names
another label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mikepitre mikepitre changed the title test(expo): borrow a simulator or emulator on a CI runner for the verify skill test(expo): add a verify skill that drives the expo-native fixture on a local or borrowed device Oct 6, 2026
`up` ended a session that was not ready 25 minutes after it asked for
the build. The Expo iOS build on the free GitHub-hosted Mac took 1,432
and 1,441 s in the two sessions measured, about a minute short of that,
so a slightly slower runner would have failed `up` with NOT_READY and
thrown the session away.

A session that ends or fails its build is still reported at once, so
the longer limit only applies to one that is alive and still building.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

2 active deployments
Preview – swingset — 8e4f27cf Deployed Oct 6, 2026 by vercel[bot]
Preview – clerk-js-sandbox — 8e4f27cf Deployed Oct 6, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant