Repository navigation
Conversation
…e it belongs The finding this line of work was built around -- that the shared scale in a block format does not need an eight-bit exponent -- was published first by someone else. Chhugani et al., arXiv:2603.08713, section 3.3: "for nearly all weight tensors and over 98% of activation tensors, a 4-bit exponent suffices to capture the scaling factor's dynamic range", on Llama-3.1-8B-Instruct and Qwen3-8B among others. The post says so in its second paragraph rather than its last, because a reviewer who knows the field finds that paper in one search. The section was read before citing it. One correction to how the prior art was described internally: that paper does NOT propose narrowing the exponent from E8. It keeps E8M0 at 1x16 and adds a macro-block E0M8 at 1x128, explicitly to avoid the hardware cost of native E4M3. The observation is theirs; the prescription that follows from it is not in that paper. Getting this backwards would have handed away more than was owed. What the post claims as its own is the theory: b_min = ceil(log2 S(W,K)) with a sufficiency proof and bit-identity rather than an empirical threshold, the R < S < R+2 bound checked on all 3,780 tensor-K pairs, and the separation of sufficiency from necessity with a counterexample -- a truncated field can return the same dequantised weights, so per-tensor necessity is false. The limits are in the body, not in footnotes: four bits hold only for K >= 8 (GPT-2 needs five at K = 4), and activations break the constant outright -- OPT needs five bits at layer inputs, twelve of seventy-two layers over four, driven by post-ReLU near-zero blocks dragging the lower bound down ~28 binades, not by outliers. "E8M0 is over-provisioned" survives; "four bits" does not survive as a single number. Prior art is cited as receipts: the MXFP4 paper, QLoRA double quantisation, Shared Microexponents, and the OCP MX specification. Full Russian translation included, since the ru overlay exists and this is the audience that asked for it.
Five runs on my own designs proved the harness runs. They could not prove anything else, because a gallery that only ever checks its author is a showroom. Three public Tiny Tapeout submissions, open licence, same checks, same day: pongsagon/tt_um_pongsagon_tinygpu_v2 34,223 cells 3,383 flops NikLeberg/tt_um_float_synth 900 cells 183 flops divadnauj-GB/tt_um_divadnauj-GB_serv_soc_wb 8,301 cells 1,751 flops All three passed all five checks, including the one my own chips failed: every declared source file was present and the design elaborated from that list alone. Two of mine did not — Phi was missing three files from its info.yaml and Gamma one, and a Tiny Tapeout shuttle builds from exactly that list. That contrast is the reason to publish this section rather than a nicer one. The instrument is credible because it was pointed at its author first and found something there. Nothing here judges anyone's design. These are structural facts about public code, reproducible with the commands printed on each card. The commit stamp is suppressed for third-party rows: I did not check them out at a pinned SHA and will not print one I cannot stand behind.
Owner
Author
|
Закрываю без мержа: содержимое уже на main. Правки этой ветки были внесены на main напрямую позднее, поэтому мерж не добавит нового, а вернёт старые версии затронутых файлов — это регресс, не улучшение. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The finding this line was built around — the shared scale in a block format does not need an eight-bit exponent — was published first by someone else. arXiv:2603.08713 §3.3: "for nearly all weight tensors and over 98% of activation tensors, a 4-bit exponent suffices to capture the scaling factor's dynamic range", on Llama-3.1-8B-Instruct and Qwen3-8B. The post says so in its second paragraph, not its last.
A correction made after reading the section rather than the abstract: that paper does not propose narrowing the exponent from E8. It keeps E8M0 at 1×16 and adds a macro-block E0M8 at 1×128, explicitly to avoid native E4M3 in hardware. The observation is theirs; the prescription is not in that paper. Getting this backwards would have handed away more than was owed.
Claimed as ours:
b_min = ceil(log2 S(W,K))with a sufficiency proof and bit-identity to E8M0 rather than an empirical threshold; theR < S < R+2bound checked on all 3,780 tensor-K pairs without exception; and sufficiency separated from necessity with a counterexample — a truncated field can return the same dequantised weights, so per-tensor necessity is false.Limits in the body, not in footnotes: four bits hold only for K ≥ 8 (GPT-2 needs five at K = 4), and activations break the constant — OPT needs five bits at layer inputs, 12 of 72 layers over four, driven by post-ReLU near-zero blocks dragging the lower bound down ~28 binades, not by outliers. "E8M0 is over-provisioned" survives; "four bits" does not survive as a single number.
Receipts: the MXFP4 paper, QLoRA double quantisation, Shared Microexponents, the OCP MX spec. Full Russian translation included.