Skip to content

Make the UI test suite runnable again, and shardable across remote Macs - #4

Draft
yusuftor wants to merge 12 commits into
mainfrom
refactor/remote-test-execution
Draft

yusuftor wants to merge 12 commits into
mainfrom
refactor/remote-test-execution

Conversation

@yusuftor

Copy link
Copy Markdown
Contributor

The suite had not run in CI since it was written: the workflow only pushed a
tag for Xcode Cloud, and on a current Mac the project did not compile at all.
This gets it building, running and splitting into shards, so it can go wide on
Limrun from Linux runners.

Draft, because the stored images still have to be taken again — see What is
left
below.

What was broken

It did not build. Configuration_ObjC.h names four SuperwallKit types
while importing only Foundation. It used to compile because another header
pulled the module in first; Xcode 27 builds modules explicitly, so include
order no longer leaks those types in and all four schemes failed. Confirmed
on an untouched main.

Every test hung for five minutes and then failed. Three faults in a row:

  • deleteApp asked for each button through the view holding it
    (collectionViews.buttons["Remove App"]). The wording of those steps is
    unchanged, but the nesting is not, so the query matched nothing and the app
    was never removed.
  • With the app already installed, the StoreKit session was set up too late and
    the app started with no products — the exact case the helper's own error
    message warns about. The helper retried by restarting the same
    SKProductsRequest, which answers once and is then spent, so no callback
    ever came. It then gave up with assertionFailure, which traps in a debug
    build, so the app died rather than reporting.
  • launchApp discarded the result of waiting for the foreground, and the
    runner had no timeout of its own, so it waited on an app that was no longer
    running.

It tested the wrong code. The project tracks the SDK's develop, so a run
picked up whatever that branch held rather than the commit that asked for it.
The dispatch already carried a commit; nothing used it.

What this changes

  • One import, and the project builds again.
  • deleteApp asks Springboard for each button directly and waits for it, then
    waits for the icon to go.
  • A fresh products request per attempt; a failure reports instead of killing
    the app.
  • The launch is checked, and the runner has a timeout for an app that stops
    answering. The app keeps its own shorter one and still reports first.
  • scripts/pin-sdk.py --commit <sha> pins the SDK to the commit under test.
  • scripts/run-tests.py plan|run splits the suite into shards, on this Mac or
    on a remote one with --runner lim. Classes are dealt round-robin, since
    test numbers group by feature and contiguous slices hand one shard all the
    slow paywall tests.
  • The workflow plans the shards and runs ten at a time — Limrun caps instances
    per org and a shard holds two, a sandbox and a simulator. The Xcode Cloud tag
    job stays until this path has proven itself.
  • Snapshots compare on how close each pixel is, not only how many match
    exactly. Text rendering moves a lot of pixels by a little between machines
    and iOS minor versions, and Limrun does not promise a fixed minor version.

Verified

  • Both a Swift and an Objective-C scheme build.
  • Test 0 passes end to end: app removed, StoreKit session set up, app
    installed and launched, SDK configured, paywall rendered with the attribute
    the test sets, image matched on a second run. 70 seconds, down from 542.
  • Shards are complete and disjoint at 1, 3, 10, 179 and 200 shards.
  • pin-sdk.py round-trips the project file byte for byte.

Nothing has run on Limrun yet — that needs LIM_API_KEY.

What is left

  • Re-record the images. They were taken on an iPhone 14 Pro at 393x818
    points; the suite now targets 402x840, so every image assertion fails on
    size alone. scripts/run-tests.py run --record does it, but it is roughly
    700 cases and many hours, and recording makes whatever renders the new
    truth, so it wants a person watching. Test 0 is recorded here as proof the
    loop closes.
  • Test 126 crashes the app. A real fault, not a snapshot difference.
  • In a sample of eight tests the failures were four snapshot mismatches and
    that one crash, so the expected mode dominates.
  • A crashed app costs the runner its full timeout. Noticing the app has gone
    would cut that.

🤖 Generated with Claude Code

yusuftor and others added 12 commits September 14, 2026 21:22
The image assertions only set `precision`, the share of pixels that have to
match exactly. Anything that does differ has to be an exact match, which is
the wrong way round for running on a machine we don't own: text rendering and
anti-aliasing nudge a large share of pixels by a tiny amount between Macs and
between iOS minor versions, so screens that look identical fail.

Add a `perceptualPrecision` to each precision value and send it along with
the pixel share, so a small shift in a lot of pixels passes and a real layout
change still fails.

Also drop `snapshotsPathComponent`, which nothing reads.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The project tracks the SDK's `develop` branch, so a run picks up whatever that
branch holds when it resolves rather than the commit that asked for the run.
The dispatch from the SDK repo already carries a commit; nothing used it.

`scripts/pin-sdk.py --commit <sha>` rewrites the package requirement to that
revision, and `--branch develop` puts it back. Both are idempotent and leave
the project file byte-identical when nothing changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suite is 179 test classes under four schemes, run one after another
against a simulator that no longer ships. End to end that is several hours on
one machine, which is why it only ever ran by hand.

`scripts/run-tests.py plan` prints the job matrix and `run` executes one entry,
so a CI job can fan the shards out and rejoin on the exit code. Classes are
dealt out round-robin rather than in blocks, because test numbers group by
feature and contiguous slices hand one shard all the slow paywall tests.
`--runner lim` builds and runs on a remote Mac instead of this one.

runTests.sh keeps working as the run-everything entry point and now aims at a
simulator that exists. DESTINATION overrides it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The workflow pushed a tag for Xcode Cloud to pick up and did nothing else, so
the suite had no run of its own and the commit in the dispatch payload went
unused.

Plan the shards, then run them on Limrun from Linux runners, pinning the SDK
to the commit that asked for the run. Ten at a time: Limrun caps instances per
org and each shard holds two, an Xcode sandbox and a simulator.

The tag job stays for now so Xcode Cloud keeps running the suite the old way
until the Limrun path has proven itself on a real commit.

Needs a LIM_API_KEY secret.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suite does not build. Configuration_ObjC.h names SWKPaywallViewController,
SWKPaywallResult, SWKSuperwallDelegate and SWKSuperwallEventInfo but imports
only Foundation, and forward-declares the protocols alone. It used to compile
because another header pulled SuperwallKit in first; Xcode 27 builds modules
explicitly, so include order no longer leaks those types in and all four
schemes fail.

Import the module in the header that needs it, and keep the derived data a
local build writes out of git.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The README described a VNC Mac and said to run all four schemes by hand.
Describe the runner, the shards, and how the SDK version gets chosen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every test begins by fetching the two custom products, and waits on a
continuation that the products-request delegate resumes. When the first reply
comes back empty the helper retries by calling start() again on the same
SKProductsRequest. That object answers once and is then spent, so the retry
never produces a callback, the continuation is never resumed, and the test sits
there until the app's own five minute timeout fires.

Build a request per attempt and hold it until it replies.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The images were recorded on an iPhone 14 Pro running iOS 16.4, a simulator
that no longer exists, so they have to be taken again on the device the suite
now targets. Nothing in the project turned recording on.

`run --record` sets SnapshotTesting recording for the run, passed in through
the TEST_RUNNER_ prefix that xcodebuild hands to the test process. Recording
only makes sense locally, since a remote sandbox keeps the images it writes,
so pairing it with the remote runner is refused.

A recording run comes back red by design: every recorded assertion reports a
failure saying it wrote a new image.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A shard is its own process already, so xcodebuild's parallel testing only
adds simulator clones underneath it. The app and the test runner talk over a
port worked out from the clone number, which makes the handshake fragile for
no gain, and the project README has long told people to turn it off by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things kept a run alive with nothing left to wait for.

deleteApp asked for each button through the view that held it, as in
`collectionViews.buttons["Remove App"]`. The wording of those steps has not
changed, but the nesting has, so on current iOS the query matched nothing and
the app was never removed. The next test then set up its StoreKit session
against an app that was already installed, which is the one thing the helper's
own error message warns against, and it started with no products. Ask
Springboard for the buttons directly and wait for each one, then wait for the
icon to go.

The products helper gave up with assertionFailure, which traps in a debug
build, so the app died rather than reporting. Both call sites already resume
the caller straight afterwards, so the trap only cost us the message. Print it
and carry on.

launchApp threw away the result of waiting for the app to come to the
foreground, so a launch that never happened looked like a launch that did, and
the runner then waited on an app that was not there. Check it, and give the
runner a timeout of its own for the case where the app stops answering
mid-test. The app keeps its own shorter timeout and still reports first when
it can.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Proves the loop end to end: the app is removed, the StoreKit session is set
up, the app installs and launches, the SDK configures, the paywall renders
with the user attribute the test sets, and the image matches on a second run.

The other images are still from the old device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant