Skip to content

docs: dot4-tile (4-lane parallel MAC) on silicon -- 3/3 bit-exact (full ladder + tile done) - #199

Merged
gHashTag merged 1 commit into
mainfrom
docs/gft-dot4-on-silicon
Aug 6, 2026
Merged

gHashTag merged 1 commit into
mainfrom
docs/gft-dot4-on-silicon

Conversation

@gHashTag

@gHashTag gHashTag commented Aug 6, 2026

Copy link
Copy Markdown
Owner

gft_dot4_ax7203 (gft_dot4_tile = 4× gft16_mul + gft_add tree) verified 3/3 on the AX7203: 4x(41,0)²=16 (0x5800), 9+4+8+4=25 (0x5920), 4x(41,256)²=36 (0x5A40) — parallel one-shot counterpart to the streaming gft_macc row.

Completes the on-silicon set: GF-T16 mul, dot2, streaming row, 4-lane tile, and the GF-T32 top rung — every GF-T compute primitive now runs bit-exact on a live FPGA from the same .t27 the A2A verifier uses.

Debug: first dot4 bitstream was silent; full-UART top sim proved the RTL correct, isolating the fault to a dead --timing-allow-fail route (seed 5) → re-PnR seed 8 fixed it.

…ll ladder + tile done)

gft_dot4_ax7203 (gft_dot4_tile = 4x gft16_mul + gft_add reduction tree) verified 3/3 on the
AX7203: 4x(41,0)^2=16 (0x5800), 9+4+8+4=25 (0x5920), 4x(41,256)^2=36 (0x5A40) -- the parallel
one-shot counterpart to the streaming gft_macc row, same golden results.

Completes the on-silicon set: GF-T16 multiply, dot2, streaming row, 4-lane tile, and the
GF-T32 top rung -- every GF-T compute primitive now runs bit-exact on a live FPGA from the
same .t27 the over-wire A2A verifier uses.

Debug note recorded: the first dot4 bitstream was silent; a full-UART top sim proved the RTL
correct, isolating the fault to a functionally-dead nextpnr --timing-allow-fail route (seed 5),
fixed by re-PnR on seed 8. README dot4 status -> 3/3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant