Repository navigation
FUG-133: HITL gate — wss:443 handshake + cert-page GET under worst-case load - #121
Merged
Merged
Conversation
Contributor
|
…se load Add //pi/hitl/harness:loaded_tls, a manual+hitl on-hardware driver that reproduces the PR #114 field failure directly: the HTTPS cert-trust page timing out and wss:443 handshakes failing because mbedTLS could not allocate its ~28 KB session on a heap starved by a resident map + effect. :rename_wss exercises the wss re-issue path but only on a CLEAN device, so it never reaches the OOM-during-handshake that actually broke. The driver reserves/flashes/provisions like the sibling drivers, then: 1. LOADs the device to its worst case over wss — a full kMaxLeds map, an activated texture-sampling effect, and a resident texture keyframe — then drops that socket so its own TLS session frees while the map/effect/texture stay resident in device RAM. 2. GATEs: opens a FRESH wss:443 session (hello/welcome), holds it open so it occupies one of the device's two mbedTLS slots, and does a concurrent HTTPS GET / on the landing/cert page asserting 200 — forcing the second ~28 KB session on the now-loaded heap, the exact path that OOM'd. 3. Asserts the captured serial shows no esp_tls_create_server_session / mbedTLS-alloc failure across the window. --device-ws hits a reachable board directly (skips the serial leg). Pure logic — the load fixtures, the set_texture wire dict, the arena/frame cap guards, and the serial-OOM scanner — lives in loaded_tls_core and is unit-tested in //pi/hitl/tests (test_loaded_tls). No firmware change: the landing page + 2-session cap already exist; this is the regression gate that they hold under load. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…n model The gate was written (PR #121, 2026-08-20) against the vendor two-mbedTLS-slot wss server (main.cpp wss_start, now behind #if !defined(LM_NETSTACK) and deleted from the shipping image). The rigs run the netstack transport (netstack_transport.cpp), a SINGLE-connection TLS server that RST-sheds concurrent SYNs, so the old step-2 assertion — hold a wss session open and force a *concurrent* HTTPS GET / -> 200 ("the second ~28 KB session") — is incompatible with how the firmware now behaves and would false-FAIL on connection-refused. Rework: - Sequential single-connection model: LOAD the heap to worst case (full kMaxLeds map + activated texture effect + resident texture keyframe) over wss, drop that socket so its TLS session frees while map/effect/texture stay RAM-resident, then run N sequential rounds, each a FRESH wss:443 handshake+welcome on the loaded heap (the exact PR #114 alloc) followed by a SEQUENTIAL cert-page GET / and a recovery probe. - Gate on RECOVERY, reusing tls_churn_core's unit-tested verdict (baseline=SKIP when unreachable; recovery of the final loaded-heap round is the anti-wedge gate; a serial crash marker fails an otherwise-PASS run). A strict single-shot handshake+GET==200 assert red-lines the genuine transient shed-then-recover (the FUG-136 lesson), so it is not used. - Read HITL_BUNDLE_RUNFILE (the removed esp32c6_flashbundle default is gone); the target ships esp32c6_netstack_flashbundle + HITL_CHUNK_BYTES=1024 (a 4096-byte TLS record overflows the netstack record buffer -> mbedtls alloc failure, which was flaking the map-upload LOAD leg). - Drop the NVS-join reboot: netstack keeps creds in RAM, so a reboot strands the DUT; test the just-provisioned link directly (mirrors tls_churn --skip-nvs-reboot). - Use the shared hitl_ws jitter tolerances (OPEN/RPC 25s) instead of tight 8/15s. - --led-count default 512 (LM_MAX_LEDS); 768 exceeded the firmware cap and the LOAD was rejected. - scan_serial_for_oom now matches the netstack strings ([heap] alloc FAILED, TLS handshake err -0x7f00); the vendor esp_tls_create_server_session string is never printed. The scan is INFORMATIONAL, never a gate. Unit test updated to the netstack OOM strings and the 512 cap; //pi/hitl/tests:hitl_test passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
fughilli
force-pushed
the
agent/fug-133-hitl-assert-wss-443-handshake-ce
branch
from
September 28, 2026 20:42
8b2ffdf to
3e4afd1
Compare
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
FUG-133 — HITL gate: wss:443 handshake under worst-case heap load
An asserting HITL gate for the PR #114 failure class: on a heap already loaded with a
resident map + a texture-sampling effect + a streamed texture keyframe, mbedTLS could
not allocate its TLS session and fresh wss:443 handshakes / the HTTPS cert page failed
until reboot.
rename_wssandtls_churnexercise the wss re-issue / connection-churnpaths but only on a clean device, so neither reproduces the OOM-during-handshake on a
loaded heap. This gate does.
Rebased + reworked for the netstack transport
This PR was originally written (2026-08-20) against the vendor two-mbedTLS-slot wss
server (
main.cpp wss_start, now behind#if !defined(LM_NETSTACK)and deleted from theshipping image). The rigs now run the netstack transport (
netstack_transport.cpp),a single-connection TLS server that RST-sheds concurrent SYNs. The original step-2
assertion — hold a wss session open and force a concurrent HTTPS
GET /→ 200 ("thesecond ~28 KB session") — is incompatible with that model and would false-FAIL on a
refused/reset connection. The gate has been reworked to the single-connection reality:
LM_MAX_LEDSmap + activated texture effect + resident texture keyframe) over wss, drop that
socket so its TLS session frees while the map/effect/texture stay RAM-resident, then run
N sequential rounds — each a fresh wss:443 handshake + welcome on the loaded heap
(the exact Fix device connection reliability: heap headroom, JIT-on, and web (re)connect flow #114 alloc), followed by a sequential cert-page
GET /(not concurrent),then a recovery probe.
tls_churn_core's unit-tested verdict (baseline unreachable⇒ SKIP, never FAIL; recovery of the final loaded-heap round is the anti-wedge gate; a
serial crash marker fails an otherwise-PASS run). A strict single-shot "handshake +
GET==200" assert red-lines the genuine transient shed-then-recover, which is graceful
degradation on a heap-tight board, not the regression (the FUG-136 lesson). On unfixed
firmware this gate can legitimately go red on the lane — that is the wedge reproducing,
not a test defect.
video_stream/LOAD-leg flake — root cause + fixThe
hitl_testscheck was red because of the LOAD leg. Root causes, all fixed://firmware/player_app:esp32c6_flashbundle,which no longer exists — the vendor image was deleted; the shipping image is
esp32c6_netstack_flashbundle. A stale data dep fails the tag-discoveredbazel test→ the whole lane reds. Target now ships the netstack bundle and readsHITL_BUNDLE_RUNFILE.map_upload_core.CHUNK_BYTESdefaults to 4096, whichoverflows the netstack TLS record buffer → mbedTLS alloc failure mid-map-upload. The
target now sets
HITL_CHUNK_BYTES=1024(as the other netstack drivers do).so the old post-provision reboot lost them and the board never rejoined. The driver now
tests the just-provisioned link directly (mirrors
tls_churn --skip-nvs-reboot).-L→DUTjitter flake; the driver now uses the shared
hitl_wstolerances.--led-countdefault 768 > the firmware cap (512). A 768-LED map was rejected; thedefault is now 512 (
LM_MAX_LEDS).The serial OOM scan now matches the strings the netstack actually prints (
[heap] alloc FAILED,TLS handshake err -0x7f00); the vendoresp_tls_create_server_sessionstring isnever emitted. That scan is informational (a shed-line count for the reviewer), never
the gate.
Verification
bazel test //pi/hitl/tests:hitl_test— pure verdict + fixtures + serial-scan, green.bazel build --platforms=@embedded//platforms:esp32c6 //firmware/player_app:esp32c6_netstack_flashbundle— green.over Improv):
netstack), every fresh handshake on the loaded heap succeeds, and each sequential cert
GET /returns 200. (The Pi 3 rig is fenced from the gating lane for unrelated USB-buscontention, so this was validated on a Pi 5.)
The gate joins the HITL lane automatically (tag
hitl); it ismanualso it stays out ofbazel test //....--device-ws wss://<ip>/wshits a reachable board directly.