Skip to content

FUG-89 S3: per-case ownership worksheet (for review) - #242

Draft
fughilli wants to merge 1 commit into
mainfrom
s3-attribution-worksheet
Draft

fughilli wants to merge 1 commit into
mainfrom
s3-attribution-worksheet

Conversation

@fughilli

@fughilli fughilli commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Handoff (2026-10-05): the remaining work and the decisions this worksheet needs are tracked in #244. Open questions 1, 2 and 6 below are superseded: (1) is handled by rr migrate apply --trust-main-guard (rules_requirements v0.2.1); (2) the collection check is replaced by rr migrate verify against CI evidence; (6) is replaced by #244's drift check. The worksheet's inputs.evidence now holds the placeholders sw/testlogs and hw/testlogs.

What this is

Today a single test case can count toward several requirements at once: once for each requirement id it is tagged with, and once more for each requirement whose verified_by names its target. FUG-89 moves to per-case ownership. Each test case gets exactly one owner, or none, and a requirement is verified by the set of cases it owns. requirements/attribution.rrplan is the worksheet for that change. It is the output of rr migrate plan (rules_requirements v0.2.0) with every case decided, and it holds a proposed owner and a one-line reason for each of the 210 cases that count toward two or more requirements today. Nothing changes yet. The worksheet is data only: no test, tag or requirements.yaml is touched, and nothing reads the file until S4 runs rr migrate apply. Every verdict stays as it is until you confirm the owners below and the S4 codemod applies them.

Evidence used: main e709ffc's traceability-evidence (Test run 37236424393) and hitl-traceability-evidence (HITL tests run 37226247132, main 907da5a). Of 1294 cases, 554 count toward a requirement, and 210 of those count toward two or more. All 210 have a proposed owner: 21 groups, 14 per-case overrides, 0 open.

How to read the tables:

  • Counts toward today: every requirement the case is credited to now.
  • Proposed owner: the one requirement it would verify after S4.
  • ⚠ needs your call: a judgment call or a departure from the design's Part C. The alternative is given in the row.

No case under firmware/, solver/, shared/ or fx_compiler/ is contested, so those areas have no table.

web

Test (target / case group) Cases Counts toward today Proposed owner Reason
//web:clocksync_test / clocksync 4 PR-13, PR-29 PR-13 SNTP offset/RTT math, min-RTT pick, ServerClock used when connecting to a device: PR-13 "LAN connection". No readiness or reconnect behaviour (PR-29).
//web:improv_provision_test / provisionViaBle wire + error 2 PR-13, PR-29 PR-13 Drives the real provisionViaBle against a fake Improv peripheral: PR-13 "device provisioning".
//web:improv_provision_test / "survives Android's first-attempt GATT flake via retry" 1 PR-13, PR-29 PR-29 Succeeds after two failed GATT connects and writes the credentials once: PR-29 "reconnect-hardened provisioning".
//web:flashUsb_test / flashUsb 6 PR-15, PR-16, PR-32 PR-15 USB VID/PID → chip family identification: PR-15 "detect hardware appropriately". No platform or transport fallback is asserted.
//web:flashEnv_test / UA detectors, iOS reason, API-less desktop hint, insecure context 5 PR-16, PR-32 PR-32 Capability detection and clear fallback reasons: PR-32 "detect platform capability … clear fallback behavior". Part C put the two UA detectors under PR-16. ⚠
//web:flashEnv_test / capable desktop, two Android WebUSB cases 3 PR-16, PR-32 PR-16 Asserts these platforms can flash: PR-16 "flashing across supported client platforms". The Android cases are a WebUSB polyfill, which PR-32's "alternate transport strategies" also describes. ⚠
//web:costModel_test / parseFxb, walkEntry, histCycles, mathFeature, costFor, dynamic histogram 9 PR-18, PR-27 PR-27 Opcode walk and per-opcode costing: PR-27 "expand coverage of modeled operations, calibrate the cost model".
//web:costModel_test / estimateFrameTime, confidenceOf 2 PR-18, PR-27 PR-18 Budget fraction and the green/yellow/red fit colouring the user sees: PR-18 "performance visibility … within device budget". PR-27's text also says "expose budget visibility". ⚠
//web:pinhole_test / lookAtQuat, project 2 PR-11, PR-34 PR-34 The TS camera model matches the Python model through a Python-generated golden: PR-34 "cross-language conformance tests".
//web:pinhole_test / quatToRotMat 1 PR-11, PR-34 PR-11 A rotation-matrix property with no cross-language golden: reconstruction geometry.

pi/server

Test (target / case group) Cases Counts toward today Proposed owner Reason
//pi/server:server_test / test_handler (mapping session, live map, submit_map/topology, errors) 31 PR-11, PR-13 PR-11 The mapping-session handler contract: PR-11 "capture, reconstruction, upload … conflict handling".
//pi/server:server_test / test_handler::test_hello_returns_welcome, test_time_sync_pong_timestamps 2 PR-11, PR-13 PR-13 The hello→welcome handshake and the server half of clock sync open the browser's LAN session: PR-13 "LAN connection". This departs from Part C, which has them under PR-11. ⚠
//pi/server:server_test / test_handler::test_playback_off_is_universal… 1 PR-11, PR-13 none get/set_playback is effect playback control. It is neither mapping (PR-11) nor provisioning/connection (PR-13). ⚠
//pi/server:server_test / test_proto_wire 45 PR-11, PR-34 PR-34 Every client/server message round-trips the protobuf envelope and re-validates against the pydantic contract: PR-34 "schema … compatibility tests, validate payload dimensions/boundaries".
//pi/server:server_test / test_session 8 PR-11, PR-22 PR-11 SessionManager start/stop/persist and MapStore JSON/CSV save: PR-11 "capture … import/export". Part C's reconnect/resume cases do not exist, so PR-22 is left with no software evidence (see the open questions). ⚠
//pi/server:server_integration_test / test_full_capture_flow 1 PR-11, PR-13 PR-11 Over a real socket: hello → mapping → detections → result_ready → GET /maps. Its stated acceptance is "a recorded session persists and is reconstructable".

pi/reconstruction and pi/led_driver

Test (target / case group) Cases Counts toward today Proposed owner Reason
//pi/reconstruction:reconstruct_test / test_reconstruct 7 PR-11, PR-31 PR-11 Projection, triangulation/BA recovery, outlier rejection, warm start: PR-11 "reconstruction". No case is large or interrupted; those cases live in test_interrupted_capture (PR-31) and test_reconstruct_scale (PR-12).
//pi/reconstruction:rust_parity_test / test_rust_parity 2 PR-11, PR-34 PR-11 Asserts solver accuracy: the Rust and Python solvers each land within 5 mm of ground truth and within 3 mm of each other, and the CLI reports a benchmark score. Only one line compares wire keys. Part C proposed PR-34, and a Rust-vs-Python parity test can be read as "cross-language conformance". ⚠
//pi/led_driver:led_driver_test / test_graycode 20 PR-11, PR-34 PR-11 The capture code: Gray/SEC-DED codewords, colour plans, every LED id recoverable. Part C's golden-vector cases do not exist here; the cross-language checks are //web:gray_test and //web:fec_test (PR-34).

pi/hitl

Test (target / case group) Cases Counts toward today Proposed owner Reason
//pi/hitl/harness:e2e_netstack / hitl_e2e::flash_boot 1 PR-13, PR-21, PR-26 PR-13 The flashed app boots and advertises Improv over BLE: PR-13 "discovery". Nothing measures heap headroom (PR-21) or aborts an unsafe workload (PR-26). ⚠
//pi/hitl/harness:e2e_netstack / hitl_e2e::improv_provision 1 PR-13, PR-29 PR-13 Provisions WiFi over BLE Improv and gets a device URL back: PR-13 "device provisioning". No readiness check, retry or repeated provisioning (PR-29). ⚠
//pi/hitl/harness:e2e_netstack / hitl_e2e::websocket_checks 1 PR-13, PR-22, PR-35 PR-13 WebSocket connect, time sync, rename echo, cert page: PR-13 "rename … LAN connection". No transient disconnect (PR-22). PR-35 is open question 5. ⚠
//pi/hitl/tests:hitl_test / test_fx_bench (bundle schema, golden margin, stable_cycles, LED count) 18 PR-17, PR-27 PR-27 The device-measurement bundle fed to deviceProfile.ts and the golden frame-cycle margin: PR-27 "calibrate the cost model, validate against held-out workloads". PR-17's text also names "calibrated performance estimation inputs". ⚠
//pi/hitl/tests:hitl_test / test_fx_bench run-health (stamps_run_health + 3 run_health_failure) 4 PR-17, PR-27 none Harness run-health gating (aborted sweep, recovered reboot). This is HITL trustworthiness, which is PR-36's text, not PR-17 or PR-27. ⚠
//pi/hitl/tests:hitl_test / test_improv 6 PR-13, PR-29 none Tests the harness's own Improv client (pi/hitl/harness/improv.py). The product codecs have their own tests: //firmware/player_app:improv_codec_test, //web:improv_test. ⚠
//pi/hitl/tests:hitl_test / test_sync 5 PR-13, PR-29 none Tests the harness's SNTP math (sync.py). The app's clock sync is //web:clocksync_test. ⚠
//pi/hitl/tests:hitl_test / test_map_upload 7 PR-12, PR-31 none Tests the harness's window plan and synthetic fixtures. It never runs the web chunker or the firmware's reassembly, and nothing is interrupted. ⚠
//pi/hitl/tests:hitl_test / test_mapping_trigger 9 PR-12, PR-31 none Tests the harness's mapping-triggered verdict. ⚠
//pi/hitl/tests:hitl_test / test_video_bench 5 PR-10, PR-17 none Tests the harness's RESULT parser and fps verdict. The TouchDesigner encoder is //tools/touchdesigner/stream_bench:stream_bench_test (PR-10). ⚠

requirements (CI)

Test (target / case group) Cases Counts toward today Proposed owner Reason
//requirements:model_test / [target] 1 PR-23, PR-25 PR-25 Validates the requirements model and the coverage gate: PR-25 "CI shall enforce quality gates". It validates no HITL/CI operational flow (PR-23). With this change PR-23 loses its only verified_by target and has to rely on HITL evidence. ⚠

Open questions (each with the proposed default)

  1. S4 blocker: the static guards refuse every Python file. In the real tree, rules_requirements v0.2.0's static guards refuse all 14 Python files. The cause is pi/provisioning/nix/improv/improv_codec_test.py, which calls importlib.util.spec_from_file_location. That call only runs under if __name__ == "__main__", never at import time. On a copy without that file, the dry run rewrites 13 files and refuses none. Default: fix the guard upstream so that code reached only from a __main__ block does not count as import-time, then run S4. The fallback is to migrate those files by hand.
  2. The collection check has not run. The tag dry run passed with --no-collect-check only. pytest collection needs uvicorn, python.runfiles and the Bazel import roots, and none of these could be installed here. Default: S4 runs rr migrate apply with the collection check on, in an environment that has the py_test deps (CI), followed by a re-plan.
  3. Harness self-tests and fx_bench run-health are none. The five pi/hitl/tests harness modules and the 4 fx_bench run-health cases are proposed none. They would fit PR-23 (the harness validates operational flows) and PR-36 (HITL/CI flakiness). That would be a new claim, and the worksheet can only choose among the claims that exist. Default: none now. Add PR-23/PR-36 tags in a follow-up if you want them counted.
  4. PR-22 has no software evidence. test_session has no disconnect case. web/tests/captureDeviceDrop.test.ts and pi/server/tests/test_capture_recovery.py are about surviving device drops, but both are tagged PR-31. Default: accept PR-22 as UNVERIFIED and write a dedicated PR-22 test. The alternative is moving one of those two tests from PR-31 to PR-22.
  5. Rename echo: PR-13 or PR-35? The S5 design assigns build_info/board_caps to PR-35. The rename echo arguably fits PR-35's "refresh network-facing identity correctly after changes" better. Rename is also in PR-13's text. Default: PR-13 until CheckPlan (S5) splits websocket_checks into separate cases. Then pick per case.
  6. The HITL run was red overall. HITL tests run 37226247132 concluded failure, although the e2e_netstack phases passed. Default: re-plan with --merge against a green HITL run before S7.

How to change a decision

  • Edit the worksheet on this branch. In requirements/attribution.rrplan, set owner: on a group (it applies to all its cases), or add owner: and reason: to a single case to override the group. Valid values are a requirement id from that case's counts_toward, or none. To make an id outside that list the owner, change the tags or verified_by first, in a separate change.
  • Or comment on this PR. Name the test and case and give the owner you want, and it will be applied here.
  • Don't regenerate the file. rr migrate plan --merge (v0.2.0) keeps the owners but drops every reason: field, so edit the file in place.
  • Check after editing. rr migrate plan --model requirements --evidence <sw testlogs> <hw testlogs> --merge requirements/attribution.rrplan should report "0 open".

Once you approve, S4 applies the owners with rr migrate apply. The implied verified_by edits are: remove PR-23 → model_test, PR-29 → clocksync_test, and PR-16/PR-32 → flashUsb_test. costModel, flashEnv, improv_provision and pinhole are split within one target, which needs v0.3 case selectors.

Do not merge until the owners are confirmed.

🤖 Generated with Claude Code

@fughilli fughilli changed the title FUG-89 S3: attribution worksheet with proposed per-case owners FUG-89 S3: per-case ownership worksheet (for review) Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://fughilli.github.io/splanc/pr-preview/pr-242/

Built to branch gh-pages at 2026-10-05 19:07 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

`rr migrate plan` (rules_requirements v0.2.0) over main e709ffc's
traceability-evidence (Test run 37236424393) and hitl-traceability-evidence
(HITL tests run 37226247132) lists 210 test cases that count toward two or
more requirements, plus 7 targets that several requirements name.
requirements/attribution.rrplan records one proposed owner per case, each
with a reason taken from what the test asserts and the requirement's text.
The PR owners confirm or change them. The worksheet is a data file only: no
test, tag or requirements.yaml change, and no verdict change.

The .rrplan extension keeps it out of the model, since CI passes
`--model requirements` as a directory and every YAML file there is read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch had an error being deployed

1 failed deployment
HITL — ead7a22f Deployed Oct 5, 2026 by fughilli via hitl_tests #764
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants