Skip to content

FUG-133: HITL gate — wss:443 handshake + cert-page GET under worst-case load - #121

Merged
fughilli merged 2 commits into
mainfrom
agent/fug-133-hitl-assert-wss-443-handshake-ce
Oct 4, 2026
Merged

fughilli merged 2 commits into
mainfrom
agent/fug-133-hitl-assert-wss-443-handshake-ce

Conversation

@issuefleet

@issuefleet issuefleet Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

FUG-133 — HITL gate: wss:443 handshake under worst-case heap load

An asserting HITL gate for the PR #114 failure class: on a heap already loaded with a
resident map + a texture-sampling effect + a streamed texture keyframe, mbedTLS could
not allocate its TLS session and fresh wss:443 handshakes / the HTTPS cert page failed
until reboot. rename_wss and tls_churn exercise the wss re-issue / connection-churn
paths but only on a clean device, so neither reproduces the OOM-during-handshake on a
loaded heap. This gate does.

Rebased + reworked for the netstack transport

This PR was originally written (2026-08-20) against the vendor two-mbedTLS-slot wss
server (main.cpp wss_start, now behind #if !defined(LM_NETSTACK) and deleted from the
shipping image). The rigs now run the netstack transport (netstack_transport.cpp),
a single-connection TLS server that RST-sheds concurrent SYNs. The original step-2
assertion — hold a wss session open and force a concurrent HTTPS GET / → 200 ("the
second ~28 KB session") — is incompatible with that model and would false-FAIL on a
refused/reset connection. The gate has been reworked to the single-connection reality:

  • Sequential single-connection model. LOAD the heap to worst case (full LM_MAX_LEDS
    map + activated texture effect + resident texture keyframe) over wss, drop that
    socket so its TLS session frees while the map/effect/texture stay RAM-resident, then run
    N sequential rounds — each a fresh wss:443 handshake + welcome on the loaded heap
    (the exact Fix device connection reliability: heap headroom, JIT-on, and web (re)connect flow #114 alloc), followed by a sequential cert-page GET / (not concurrent),
    then a recovery probe.
  • Gate on RECOVERY, reusing tls_churn_core's unit-tested verdict (baseline unreachable
    ⇒ SKIP, never FAIL; recovery of the final loaded-heap round is the anti-wedge gate; a
    serial crash marker fails an otherwise-PASS run). A strict single-shot "handshake +
    GET==200" assert red-lines the genuine transient shed-then-recover, which is graceful
    degradation on a heap-tight board, not the regression (the FUG-136 lesson). On unfixed
    firmware this gate can legitimately go red on the lane
    — that is the wedge reproducing,
    not a test defect.

video_stream/LOAD-leg flake — root cause + fix

The hitl_tests check was red because of the LOAD leg. Root causes, all fixed:

  1. Build error (primary red). The target referenced //firmware/player_app:esp32c6_flashbundle,
    which no longer exists — the vendor image was deleted; the shipping image is
    esp32c6_netstack_flashbundle. A stale data dep fails the tag-discovered
    bazel test → the whole lane reds. Target now ships the netstack bundle and reads
    HITL_BUNDLE_RUNFILE.
  2. Oversized TLS records. map_upload_core.CHUNK_BYTES defaults to 4096, which
    overflows the netstack TLS record buffer → mbedTLS alloc failure mid-map-upload. The
    target now sets HITL_CHUNK_BYTES=1024 (as the other netstack drivers do).
  3. NVS-join reboot stranded the DUT. The netstack keeps WiFi creds in RAM (no NVS),
    so the old post-provision reboot lost them and the board never rejoined. The driver now
    tests the just-provisioned link directly (mirrors tls_churn --skip-nvs-reboot).
  4. Tight timeouts. Hardcoded 8 s/15 s reintroduced the driver→tailnet→rig→ssh -L→DUT
    jitter flake; the driver now uses the shared hitl_ws tolerances.
  5. --led-count default 768 > the firmware cap (512). A 768-LED map was rejected; the
    default is now 512 (LM_MAX_LEDS).

The serial OOM scan now matches the strings the netstack actually prints ([heap] alloc FAILED, TLS handshake err -0x7f00); the vendor esp_tls_create_server_session string is
never emitted. That scan is informational (a shed-line count for the reviewer), never
the gate.

Verification

  • bazel test //pi/hitl/tests:hitl_test — pure verdict + fixtures + serial-scan, green.
  • bazel build --platforms=@embedded//platforms:esp32c6 //firmware/player_app:esp32c6_netstack_flashbundle — green.
  • On real hardware (esp32c6 on a Pi 5 rig, flashed with the netstack bundle + provisioned
    over Improv):
    [load] map: 30658 B -> 30 window(s)            # 512-LED map, sharded at 1024 B/record
    [load] resident: 512-LED map + effect '__fug133' + 40x40 texture (3200 B frame)
    [round 1] handshake OK under load … cert-page GET / -> 200 … wss RECOVERED 2.2s
    [round 2] handshake OK under load … cert-page GET / -> 200 … wss RECOVERED 2.3s
    [round 3] handshake OK under load … cert-page GET / -> 200 … wss RECOVERED 2.2s
    RESULT rounds=3 baseline=ok recovered=3/3 served_total=3 cert_ok=3/3 crashed=no verdict=PASS
    
    The LOAD leg that was flaking now completes cleanly (30 × ~1024 B windows over wss on the
    netstack), every fresh handshake on the loaded heap succeeds, and each sequential cert
    GET / returns 200. (The Pi 3 rig is fenced from the gating lane for unrelated USB-bus
    contention, so this was validated on a Pi 5.)

The gate joins the HITL lane automatically (tag hitl); it is manual so it stays out of
bazel test //.... --device-ws wss://<ip>/ws hits a reachable board directly.

@github-actions

github-actions Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-04 18:24 UTC

Claude Agent and others added 2 commits September 28, 2026 20:28
…se load

Add //pi/hitl/harness:loaded_tls, a manual+hitl on-hardware driver that
reproduces the PR #114 field failure directly: the HTTPS cert-trust page timing
out and wss:443 handshakes failing because mbedTLS could not allocate its ~28 KB
session on a heap starved by a resident map + effect. :rename_wss exercises the
wss re-issue path but only on a CLEAN device, so it never reaches the
OOM-during-handshake that actually broke.

The driver reserves/flashes/provisions like the sibling drivers, then:
  1. LOADs the device to its worst case over wss — a full kMaxLeds map, an
     activated texture-sampling effect, and a resident texture keyframe — then
     drops that socket so its own TLS session frees while the map/effect/texture
     stay resident in device RAM.
  2. GATEs: opens a FRESH wss:443 session (hello/welcome), holds it open so it
     occupies one of the device's two mbedTLS slots, and does a concurrent HTTPS
     GET / on the landing/cert page asserting 200 — forcing the second ~28 KB
     session on the now-loaded heap, the exact path that OOM'd.
  3. Asserts the captured serial shows no esp_tls_create_server_session /
     mbedTLS-alloc failure across the window.

--device-ws hits a reachable board directly (skips the serial leg). Pure logic —
the load fixtures, the set_texture wire dict, the arena/frame cap guards, and the
serial-OOM scanner — lives in loaded_tls_core and is unit-tested in
//pi/hitl/tests (test_loaded_tls). No firmware change: the landing page + 2-session
cap already exist; this is the regression gate that they hold under load.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…n model

The gate was written (PR #121, 2026-08-20) against the vendor two-mbedTLS-slot
wss server (main.cpp wss_start, now behind #if !defined(LM_NETSTACK) and deleted
from the shipping image). The rigs run the netstack transport
(netstack_transport.cpp), a SINGLE-connection TLS server that RST-sheds concurrent
SYNs, so the old step-2 assertion — hold a wss session open and force a *concurrent*
HTTPS GET / -> 200 ("the second ~28 KB session") — is incompatible with how the
firmware now behaves and would false-FAIL on connection-refused.

Rework:
- Sequential single-connection model: LOAD the heap to worst case (full kMaxLeds
  map + activated texture effect + resident texture keyframe) over wss, drop that
  socket so its TLS session frees while map/effect/texture stay RAM-resident, then
  run N sequential rounds, each a FRESH wss:443 handshake+welcome on the loaded
  heap (the exact PR #114 alloc) followed by a SEQUENTIAL cert-page GET / and a
  recovery probe.
- Gate on RECOVERY, reusing tls_churn_core's unit-tested verdict (baseline=SKIP
  when unreachable; recovery of the final loaded-heap round is the anti-wedge gate;
  a serial crash marker fails an otherwise-PASS run). A strict single-shot
  handshake+GET==200 assert red-lines the genuine transient shed-then-recover (the
  FUG-136 lesson), so it is not used.
- Read HITL_BUNDLE_RUNFILE (the removed esp32c6_flashbundle default is gone); the
  target ships esp32c6_netstack_flashbundle + HITL_CHUNK_BYTES=1024 (a 4096-byte
  TLS record overflows the netstack record buffer -> mbedtls alloc failure, which
  was flaking the map-upload LOAD leg).
- Drop the NVS-join reboot: netstack keeps creds in RAM, so a reboot strands the
  DUT; test the just-provisioned link directly (mirrors tls_churn --skip-nvs-reboot).
- Use the shared hitl_ws jitter tolerances (OPEN/RPC 25s) instead of tight 8/15s.
- --led-count default 512 (LM_MAX_LEDS); 768 exceeded the firmware cap and the LOAD
  was rejected.
- scan_serial_for_oom now matches the netstack strings ([heap] alloc FAILED,
  TLS handshake err -0x7f00); the vendor esp_tls_create_server_session string is
  never printed. The scan is INFORMATIONAL, never a gate.

Unit test updated to the netstack OOM strings and the 512 cap; //pi/hitl/tests:hitl_test
passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@fughilli
fughilli force-pushed the agent/fug-133-hitl-assert-wss-443-handshake-ce branch from 8b2ffdf to 3e4afd1 Compare September 28, 2026 20:42
@fughilli
fughilli merged commit e7c7a38 into main Oct 4, 2026
11 checks passed

This branch was successfully deployed

1 active deployment
HITL — 3e4afd18 Deployed Sep 28, 2026 by fughilli via hitl_tests #655
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants