Skip to content

Publish the scale-field-width result, and give the priority away where it belongs - #658

Closed
gHashTag wants to merge 2 commits into
mainfrom
post/scale-field-width-v2
Closed

gHashTag wants to merge 2 commits into
mainfrom
post/scale-field-width-v2

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

The finding this line was built around — the shared scale in a block format does not need an eight-bit exponent — was published first by someone else. arXiv:2603.08713 §3.3: "for nearly all weight tensors and over 98% of activation tensors, a 4-bit exponent suffices to capture the scaling factor's dynamic range", on Llama-3.1-8B-Instruct and Qwen3-8B. The post says so in its second paragraph, not its last.

A correction made after reading the section rather than the abstract: that paper does not propose narrowing the exponent from E8. It keeps E8M0 at 1×16 and adds a macro-block E0M8 at 1×128, explicitly to avoid native E4M3 in hardware. The observation is theirs; the prescription is not in that paper. Getting this backwards would have handed away more than was owed.

Claimed as ours: b_min = ceil(log2 S(W,K)) with a sufficiency proof and bit-identity to E8M0 rather than an empirical threshold; the R < S < R+2 bound checked on all 3,780 tensor-K pairs without exception; and sufficiency separated from necessity with a counterexample — a truncated field can return the same dequantised weights, so per-tensor necessity is false.

Limits in the body, not in footnotes: four bits hold only for K ≥ 8 (GPT-2 needs five at K = 4), and activations break the constant — OPT needs five bits at layer inputs, 12 of 72 layers over four, driven by post-ReLU near-zero blocks dragging the lower bound down ~28 binades, not by outliers. "E8M0 is over-provisioned" survives; "four bits" does not survive as a single number.

Receipts: the MXFP4 paper, QLoRA double quantisation, Shared Microexponents, the OCP MX spec. Full Russian translation included.

…e it belongs

The finding this line of work was built around -- that the shared scale in a
block format does not need an eight-bit exponent -- was published first by
someone else. Chhugani et al., arXiv:2603.08713, section 3.3: "for nearly all
weight tensors and over 98% of activation tensors, a 4-bit exponent suffices to
capture the scaling factor's dynamic range", on Llama-3.1-8B-Instruct and
Qwen3-8B among others. The post says so in its second paragraph rather than its
last, because a reviewer who knows the field finds that paper in one search.

The section was read before citing it. One correction to how the prior art was
described internally: that paper does NOT propose narrowing the exponent from
E8. It keeps E8M0 at 1x16 and adds a macro-block E0M8 at 1x128, explicitly to
avoid the hardware cost of native E4M3. The observation is theirs; the
prescription that follows from it is not in that paper. Getting this backwards
would have handed away more than was owed.

What the post claims as its own is the theory: b_min = ceil(log2 S(W,K)) with a
sufficiency proof and bit-identity rather than an empirical threshold, the
R < S < R+2 bound checked on all 3,780 tensor-K pairs, and the separation of
sufficiency from necessity with a counterexample -- a truncated field can return
the same dequantised weights, so per-tensor necessity is false.

The limits are in the body, not in footnotes: four bits hold only for K >= 8
(GPT-2 needs five at K = 4), and activations break the constant outright -- OPT
needs five bits at layer inputs, twelve of seventy-two layers over four, driven
by post-ReLU near-zero blocks dragging the lower bound down ~28 binades, not by
outliers. "E8M0 is over-provisioned" survives; "four bits" does not survive as a
single number.

Prior art is cited as receipts: the MXFP4 paper, QLoRA double quantisation,
Shared Microexponents, and the OCP MX specification. Full Russian translation
included, since the ru overlay exists and this is the audience that asked for it.
Five runs on my own designs proved the harness runs. They could not prove
anything else, because a gallery that only ever checks its author is a showroom.
Three public Tiny Tapeout submissions, open licence, same checks, same day:

  pongsagon/tt_um_pongsagon_tinygpu_v2          34,223 cells   3,383 flops
  NikLeberg/tt_um_float_synth                      900 cells     183 flops
  divadnauj-GB/tt_um_divadnauj-GB_serv_soc_wb    8,301 cells   1,751 flops

All three passed all five checks, including the one my own chips failed: every
declared source file was present and the design elaborated from that list alone.
Two of mine did not — Phi was missing three files from its info.yaml and Gamma
one, and a Tiny Tapeout shuttle builds from exactly that list.

That contrast is the reason to publish this section rather than a nicer one. The
instrument is credible because it was pointed at its author first and found
something there.

Nothing here judges anyone's design. These are structural facts about public
code, reproducible with the commands printed on each card. The commit stamp is
suppressed for third-party rows: I did not check them out at a pinned SHA and
will not print one I cannot stand behind.
@gHashTag

Copy link
Copy Markdown
Owner Author

Закрываю без мержа: содержимое уже на main. Правки этой ветки были внесены на main напрямую позднее, поэтому мерж не добавит нового, а вернёт старые версии затронутых файлов — это регресс, не улучшение.

@gHashTag gHashTag closed this Aug 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant