feat(vsa): Phase 3 — synthesis top, benchmark 135x, release v0.2, arXiv draft - #13
Merged
Merged
Conversation
Phase 3 of issue #7: synthesis + demo preparation Verilog: - vsa_matmul_top.v: synthesizable top module for QMTECH XC7A100T Autoregressive ternary inference loop: Embed -> VSA MatMul -> Argmax -> UART TX MMCM 50->81.25 MHz, LED heartbeat, 10-byte binary UART frame - vsa_matmul_top.xdc: pin constraints (clk U22, UART K20/L20, LED T23) Rust: - SynthConfig::vsa_matmul() — one-command VSA synthesis config - CLI: synth-vsa, bench commands - Benchmark: 9,255 tok/s CPU vs 1,250,000 tok/s FPGA estimate = 135x speedup Docs: - release/RELEASE-v0.2.md — release manifest - docs/arxiv-trinity-stack-draft.md — arXiv paper draft Tests: 21/21 cargo test, 0 clippy warnings
This was referenced May 6, 2026
Open
This was referenced May 14, 2026
gHashTag
pushed a commit
that referenced
this pull request
Aug 10, 2026
…tiplier BitNet stores ternary weights plus a real per-layer scale alpha = mean|W|, and multiplying by that alpha puts the multiplier back at the layer boundary. Snapping the scale to a grid removes it. The phi grid is denser than powers of two by log(2)/log(phi) = 1.440 at the same cost class, so the prediction made before measuring was that its excess error over the unreachable exact alpha would be about half. Measured over 210 layers: exact 0.476781, phi^k 0.488424 (+2.4420%), 2^k 0.499943 (+4.8579%). Ratio of excesses 0.501 against a predicted 0.500. phi wins 163 of 210 layers -- not all, since a layer whose optimum lands near a power of two is better served by the coarser grid. Together with dot_exact this closes the multiplier out of the entire layer: weights, accumulation, and now the scale. Defect #13 recorded rather than reported as a result. The first attempt asked this through perplexity, where post-hoc ternarisation destroys a model not trained for it: every arm including BitNet's exact alpha landed at ppl ~2.2e7 against a baseline of 14.49. Read naively that says phi is refuted by 2x. It says nothing -- the tell was that the control arm was destroyed too, and a comparison whose control fails is not a comparison. Refs #516
gHashTag
added a commit
that referenced
this pull request
Aug 10, 2026
* Measure the actual frontier: four levers against MXFP4, and what each is worth
The goal is a format that leads the world, so the measurement has to happen where
the leader stands. That is not where this project's map is. MX puts the exponent
outside the number -- one per block of 32 -- and MXFP4 runs natively on Blackwell
and MI355X. It is a fifth axis, and our own classifier had already said so: five
of six MX formats landed in "range < 6, unclassifiable", which I read as a limit
of the instrument rather than as the instrument telling me these formats are not
in its domain.
Four levers, block 32, NRMSE, three workloads:
finer shared scale than e8m0 1.05-1.09x, and 0.82x on outliers
other minifloat shapes 0.27-1.11x
uniform int4 1.05-1.12x
Lloyd-Max levels on the data 1.51-1.74x, but 0.82x on activations
dropping the sign on ReLU data 2.35x
The control matters more than the largest number: the same unsigned format on
weights collapses to 0.22x, as it must. Without that, 2.35x would mean nothing.
What the numbers say is uncomfortable. At four bits the format is close to
saturated -- MXFP4 sits within 1.5x of the MSE-optimal quantiser -- and every
remaining gain is workload-specific: the level set that wins on weights loses on
activations. The largest lever found, 2.35x, is not a new format at all. It is
dropping the sign bit on one-sided data, which is standard quantisation practice.
MX carries a sign because it is a general format, not because that is optimal for
post-ReLU tensors.
So the measurement confirms this project's own corollary from the other side:
formats cannot be ranked without naming a workload. The universal 4-bit format is
saturated. What is left is not a better format but a rule for choosing one per
tensor, and whether that composes end to end is the next experiment rather than a
claim.
One wrong result was caught and is recorded: the first version had int4 beating
MXFP4 by 3.5x, which is nonsense. I had normalised by amax, putting every value in
[0,1] and denying the element format its entire upper range. The committed script
fits Lloyd-Max in the element's own normalised domain and asserts all three
findings, including the control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Seven measurements at the frontier, and the honest verdict they add up to
The goal is a format that leads the world, so the measurement has to happen where
the leader stands: MX, block-scaled, hardware-native on Blackwell and MI355X.
Seven levers measured. Finer shared scale: 1.05-1.09x, and 0.82x on outliers.
Other minifloat shapes: at most 1.12x. Dropping the sign on one-sided data: 2.35x,
with the control confirming it collapses to 0.22x on weights as it must. Block
size from 8 to 128: a straight exchange, 1.18x of accuracy for 1.18x of bits, so
the standard's 32 sits on a flat part of the curve. A per-tensor selector: 1.73x
across a mixed network, with fitted levels transferring to held-out data at 98-99%
and never below 1.26x cross-distribution.
Two findings are structural rather than numerical.
The levers do not compose. On every workload exactly one wins and the combination
is always worse than the better single lever: weights 0.22x / 1.50x / 0.21x,
activations 2.38x / 0.78x / 1.91x, Laplace 0.31x / 1.74x / 0.31x. The cause is
that both fix the same thing -- the mismatch between the level set and the
distribution's support -- so fixing it twice can only hurt. A selector is
therefore a classifier, not an accumulator, and its ceiling is the maximum over
levers rather than their product.
And the sign lever exceeds any choice inside the signed family. On post-ReLU
activations the best signed minifloat among e1m2 / e2m1 / e3m0 is e2m1 itself at
1.00x, while unsigned uint4 gives 2.35x.
That matters because the selector idea is already published and active --
BlockDialect picks a per-block format from a formatbook, MixFP4 switches between
E2M1 and E1M2 per block, dMX learns a per-layer assignment. Experiment 6
rediscovered known work, and the file says so. What those methods appear not to
touch is the sign axis, which is where the largest measured lever lives. That is
recorded as a claim requiring verification by reading the papers, not as a result.
The verdict is that no claim to a leading format follows from these numbers. At
four bits the universal format is saturated, and the largest remaining lever is
standard quantisation practice rather than a new format. Two of my own wrong
results are recorded in the file: int4 beating MXFP4 by 3.5x, caused by
normalising away the element format's upper range, and a "levers are almost
multiplicative" verdict computed against a lever that was losing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Read BlockDialect, measure the sign lever on real activations, and close the question
Two things needed checking before the sign lever could be called a gap: whether
published selectors already use it, and whether it survives on activations that
real networks produce. The first says no, the second says the lever does not
survive.
BlockDialect, section 3.1 and Figure 4, read rather than inferred: sixteen
dialects, each a set of magnitudes {7.5, 5.5, 3, 2, 1.5, 1, 0.5, 0}, stored as
"1-bit sign and 3-bit index". Every dialect is signed. The formatbook varies
magnitudes, not signedness, so the sign axis is genuinely untouched.
That stops being an opportunity at experiment 8. Measured against real activation
functions, unsigned gives 2.35x on ReLU, 1.32x on GELU, and 0.26x on SiLU/SwiGLU
— and 0.21x on post-LayerNorm tensors. BlockDialect's own profiling quantises
attn_input and mlp_input, which are post-normalisation and symmetric by
construction, and modern LLMs use SwiGLU. The lever is negative exactly where
MXFP4 matters. An asymmetric 12/3 allocation recovers GELU to 1.97x but still
loses on SwiGLU at 0.43x.
So the search for a leading format through the element format is closed: the
universal 4-bit format sits within 1.5x of the MSE optimum, block size is a
straight exchange, minifloat shape is exhausted, the selector idea is published,
and the largest lever found does not transfer to the architectures that use MXFP4.
One structural result stands on its own. The levers partition rather than compose:
a quantiser's error comes from one source, the mismatch between its level set and
the distribution's support, and every lever addresses that same mismatch, so two
cannot compose and the better dominates. Measured on three workloads, the
combination is always worse than the best single lever. The consequence is that a
format selector is a classifier and not an accumulator, and its ceiling is the
maximum over levers rather than their product — an upper bound on the whole
BlockDialect / MixFP4 / dMX family that I have not found stated anywhere.
What is not done is stated too: perplexity was never measured. Everything here is
NRMSE on synthetic tensors passed through real activation functions, and no claim
about network quality is defensible without quantising a live model.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Energy asymmetry: the unclaimed ceiling sits on GELU and SwiGLU
One-sidedness is a property of a tensor's energy, not of how many of its values are
negative. GELU sends half its values below zero and 1.8% of its energy; SwiGLU
6.2%. A format that splits its levels evenly by sign spends eight of sixteen codes
on 1.8% of the work, and an unsigned format discards that 1.8% entirely. Both are
wrong, differently.
Measuring the ceiling makes the consequence visible. With D* the distortion floor
over all 16-point quantisers, existing levers saturate it on symmetric tensors --
weights 1.50x against a 1.51x ceiling, post-LayerNorm 1.35x against 1.38x -- and
leave 29-33% unclaimed on exactly the two activations modern LLMs use: GELU 1.74x
against 2.60x, SwiGLU 2.02x against 2.85x.
The optimum's shape says why. Free Lloyd-Max on 16 levels spends 5 of them below
zero on GELU, reaching only -0.65 while the positive side runs to +5.13; on SwiGLU
6 levels to -1.83 against +5.26. Neither symmetric nor one-sided: a narrow dense
lobe down, a long sparse ladder up.
A one-integer family captures it. With k levels down to the block minimum and
16-k up, SwiGLU reaches 2.56x at k=4 -- 90% of the codebook optimum and 1.27x
better than the best existing lever -- and GELU 2.01x at k=4 against 1.74x. ReLU
takes k=1 at 2.21x.
Why this is not in the published methods, read rather than assumed: BlockDialect's
sixteen dialects are each a set of magnitudes stored as "1-bit sign and 3-bit
index" (section 3.1, Figure 4). Every dialect is symmetric in sign. The formatbook
varies magnitudes, not how levels are distributed between the two sides, so an
asymmetric allocation is not expressible in their representation. MixFP4 switches
between E2M1 and E1M2, both signed and symmetric.
The shared-ceiling theorem is now checked numerically on five workloads: no lever
and no combination exceeded D(F0)/D*. Its corollary bounds the whole selector
family -- a selector is a classifier, not an accumulator, and its ceiling is the
maximum over levers rather than their product.
One error found on the way and recorded. The first version had
silu(x) = x/(1+exp(-(-x))), a doubled negation that mirrored the function. It
surfaced because the optimum showed ten negative levels reaching -5.24, which SiLU
cannot produce -- it is bounded below by -0.2785. The number contradicted theory
and the bug was mine, not a discovery. Every SwiGLU row is recomputed.
Perplexity is still not measured, and the hardware cost of an asymmetric level
allocation is not estimated -- BlockDialect chose symmetric dialects to keep MAC
arithmetic integer, and asymmetry may break that.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* An asymmetric dialect beats DialectFP4 on activations at no arithmetic cost
Head to head against the real formatbook from Figure 4 of arXiv:2501.01144, on
their 0.5 grid, block 32, their scale rule floor(log2 max) - 2, with the best
dialect chosen per block by exhaustive MSE — which is the upper bound of their
approach, above their own two-stage heuristic, so the comparison is generous to
them:
GELU 1.80x -> 2.11x (1.17x)
SiLU/SwiGLU 2.51x -> 3.02x (1.20x)
ReLU 1.64x -> 2.39x (1.46x)
weights 1.59x -> 1.49x (0.94x)
The reason is the energy asymmetry. GELU sends half its values below zero and 1.8%
of its energy; SwiGLU 6.2%. All sixteen DialectFP4 dialects are symmetric — eight
magnitudes stored as "1-bit sign and 3-bit index" — so eight of sixteen codes go to
1.8% of the work. The asymmetric variant spends k levels down and 16-k up with the
depth taken from the block: one integer, chosen by the same mechanism that already
chooses a dialect.
The hardware cost is zero by their own argument. DialectFP4 keeps every magnitude a
multiple of 0.5 so the index maps to an integer 0..15 and the multiply stays 4-bit
integer. The asymmetric shape stays on that grid — what changes is which integers,
not that they are integers. Rounding to the grid costs 6% on SwiGLU and ReLU, 14%
on GELU, and the formatbook grows by five entries against the sixteen it has.
Asymmetry helps on asymmetric tensors and mildly hurts on symmetric ones, so the
selector picks symmetric for weights and asymmetric for activations on a cheap
feature: the fraction of energy below zero. That matches the shared-ceiling
theorem — levers do not compose, exactly one wins, and the selector's job is to
say which.
What this is not: a leading format. It is a one-parameter extension to someone
else's formatbook worth 1.17-1.46x on activations, measured as NRMSE on Gaussian
inputs passed through real activation functions. Perplexity is still unmeasured,
real activations carry heavy tails and per-channel outliers that will move both the
optimal k and the size of the gain, and a claim this size needs a live model before
it is a result.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Real trained tensors: the criterion predicts correctly, the format wins once
Every earlier measurement fed x ~ N(0,1) through an activation function, which
assumes the distribution rather than observing it and cannot produce the
per-channel outliers training creates. This trains a small SwiGLU transformer to
convergence and captures what an inference engine actually quantises. Local models
were no use: ollama and LM Studio hold only GGUF, already quantised to 4 bits, so
measuring 4-bit quantisation error on them is circular.
The energy asymmetry survives training -- 3.4% of energy below zero on the SiLU
gate against 6.2% on synthetic, the same order. But it lives only there. The gate
multiplied by w3(x) restores symmetry, so swiglu_hidden is 50.4%, as are both
post-LayerNorm tensors and the weights.
Head to head on those tensors, the asymmetric format wins on exactly one:
silu_gate 2.77x against BlockDialect's 2.32x. Elsewhere BlockDialect wins --
attn_input 1.64 vs 1.44, mlp_input 1.61 vs 1.44, swiglu_hidden 2.01 vs 1.63 -- and
weights are a tie at 1.49 vs 1.41.
That is a validated theory with a narrow application. The energy criterion
correctly picked the one tensor where asymmetry pays and correctly predicted the
loss on the other four. But the SiLU gate is an intermediate tensor that many
implementations fuse and never materialise in low precision, so a 1.19x on it may
apply to nothing.
It also corrects my own synthetic result. The 1.20x I reported for "SwiGLU" was
measured on the SiLU output, not on the SwiGLU block output. On the real block
output the asymmetric format loses, 1.63x against 2.01x.
One finding recorded but not pursued: swiglu_hidden carries a per-channel outlier
ratio of 188x between maximum and median, against 2x for weights and 5-7x for the
post-LayerNorm tensors. That is the known 4-bit problem and it is concentrated in
one MLP tensor.
Of eight levers tried, one reached real data, and it wins on one tensor of five.
What stands is the shared-ceiling theorem and the energy-asymmetry criterion as a
predictor. What does not exist is a leading format, and perplexity is still
unmeasured.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Transforms move the ceiling that formats share -- and can destroy what a format exploits
The shared-ceiling theorem says format levers compete for a fixed budget D*(b,p)
and therefore do not compose. An orthogonal transform changes p itself, so it
changes D*. That predicts transforms and formats should compose where format levers
do not, and the prediction is testable on real trained tensors.
A Hadamard rotation over groups of 32, as QuaRot does, on swiglu_hidden: the
outlier ratio falls 157x to 10x, MXFP4's error falls 0.216 to 0.125 -- a 1.72x gain
with the format untouched -- and the ceiling falls with it, 1.72x to 1.35x. The
rotation takes the gain the format levers were competing for and leaves less
behind. That is a quantitative account of why the field went to rotations rather
than formats.
They do compose, but only sometimes, and the exception is the finding.
swiglu_hidden format 1.72x rotation 1.72x together 2.33x 79% of product
silu_gate format 2.68x rotation 1.33x together 2.06x 58%, and WORSE
than format alone
attn_input format 1.40x rotation 1.02x together 1.38x nothing to remove
On silu_gate the combination is worse than the format by itself. Rotation removes
outliers, a property of the tail; the asymmetric format exploits one-sided energy, a
property of the support. On swiglu_hidden the structure was outliers, rotation
removed them, and a symmetric tail still needed the format -- so they compose. On
silu_gate the format lived on asymmetry and rotation symmetrises the distribution,
destroying exactly what the format was earning from.
The practical consequence: applying QuaRot and then optimising the format is not the
same as optimising the format on the original data. The order changes which format
is optimal and can cancel the gain the format was chosen for. I have not found this
stated anywhere.
Refined statement: transforms and formats compose if and only if they address
different structure -- a transform that reduces the statistic a format exploits
anti-composes with it. Measured at 79% of the product where the structures differ
and 58% where they collide.
Caveats stand: the rotation is applied along a vector in groups of 32 rather than
along the hidden dimension with weight compensation as QuaRot does, perplexity is
still unmeasured, and the model is small with a synthetic task.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Order of operations changes the result by up to 2.78x, with a clean control
The composition condition says a transform and a format compose iff they address
different structure. That has a directly testable consequence: the order of
"fit the format" and "rotate" should matter, and should not matter where the
rotation has nothing to remove.
swiglu_hidden format-then-rotate 1.82x rotate-then-format 2.33x 1.28x
silu_gate format-then-rotate 0.74x rotate-then-format 2.06x 2.78x
attn_input 1.40x / 1.38x 0.99x
mlp_input 1.36x / 1.36x 1.00x
The control is clean: on the two tensors where rotation removes nothing, the ratio
is 0.99 and 1.00, so the effect is not noise. And the worst case says the most --
on silu_gate a format fitted before rotation gives 0.74x, worse than doing nothing
at all, because it was tuned to an asymmetry the rotation then destroyed.
Stated as non-commutativity: the distortion of F(p) applied to T(p) differs from
that of F(T(p)) applied to T(p), with the gap growing in how much T changes the
statistic F exploits, and equality iff T leaves that statistic alone.
The practical consequence is that quantisation pipelines treat rotation and format
selection as independent stages, and they are not. A format chosen from the
original tensor's statistics -- including any selector from BlockDialect, MixFP4 or
dMX calibrated before rotation -- can end up worse than the baseline once rotation
changes the distribution. Transform first, then select. Obvious in hindsight, but it
follows from the theorem rather than from intuition, and the size of the effect is
not obvious at all.
Caveats: the rotation is applied in groups of 32 along a vector rather than along
the hidden dimension with weight compensation as QuaRot does, the model is small
with a synthetic task, and the "format fitted on original data" is free Lloyd-Max
on 16 levels, the strongest available -- a weaker selector would suffer less
because it has less to lose.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Perplexity on a task that can degrade, and the summary of what this line produced
The first perplexity attempt was useless and the fault was mine: the model learned
the task to loss 0.0000, and a task with zero error has infinite margin, so
quantisation either breaks the argmax or does not. That is a binary, not a quality
measure. Adding 30% random labels gives the task an irreducible floor of 1.913 and
makes degradation graded.
fp32 loss 2.3350 ppl 10.330
MXFP4 on all weights loss 4.5092 ppl 90.849 8.79x
Lloyd-16 on all weights loss 2.3239 ppl 10.215 0.99x
Fitting levels to the distribution takes essentially all of it. The absolute
numbers do not transfer to real LLMs, which have redundancy and calibration
pipelines, but the ordering does and that is what is claimed.
SUMMARY.md consolidates the line. No leading format was found, and the measurements
say why: of eight format levers, one reached real data and wins on one tensor of
five -- an intermediate gate many implementations never materialise. What stands is
three statements, all verified, all about boundaries rather than about a format.
The shared ceiling: format levers compete for D*(b,p) and do not compose, so a
selector is a classifier and not an accumulator. That bounds BlockDialect, MixFP4
and dMX from above.
The energy criterion: one-sidedness is a property of energy, not of the count of
negatives, and it predicted the winner on all five real tensors.
The composition condition and its non-commutativity: a transform moves D* while a
format divides it, so the two compose iff they address different structure --
Hadamard rotation composes with format on outliers (2.33x from 1.72x and 1.72x) and
anti-composes on asymmetry (2.06x against the format's own 2.68x). Order therefore
matters, by 2.78x in the worst case, and a format fitted before rotation can give
0.74x, worse than doing nothing. The control on two flat tensors returns exactly
1.00.
Six errors caught along the way are listed rather than buried, each in the file
where it happened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The ternary-network datapath, measured as one design rather than summed
In a ternary network the weight is in {-1, 0, +1}, so w*a is a select, a negate or
a zero -- not a multiply. The multiplier disappears for every format at once, and
the denominator the decoder is compared against stops being thousands of LUTs of
multiplier and becomes hundreds of LUTs of adder. What was overhead in an ordinary
network becomes the body of the datapath here.
Synthesised as one design on XC7A200T, no DSP, differing in exactly one block --
posit needs a regime decode before the add, TEF reads its fields directly:
TEF 440 LUT 80.73 MHz 0.184 MHz/LUT
posit 895 LUT 31.81 MHz 0.036 MHz/LUT
2.03x in area, 2.54x in frequency, 5.16x in throughput per LUT. Measured end to
end rather than assembled from separate decoder and adder figures.
The decomposition explains the size: the decoder is 20% of an ordinary datapath for
posit and 54% of a ternary one, and 85% against 97% for takum, which also brings 84
block-RAM tiles.
Three things this does not claim, stated in the README because each would be found
in minutes otherwise. It is not a format without multiplications -- TEF pays the
significand multiply like everyone else when one is needed, and the quadratic term
of its own area law is that multiplier; what vanishes in a ternary network vanishes
for the network, not for the format. It is not best-for-FPGA in general -- the
exponent decode of TEF and of an ordinary binary fixed field measure identically at
32 LUTs each, so the advantage is over tapered formats and not over fixed-field
ones. And it is not a measurement on a ternary fabric, which does not exist to buy;
everything here is a binary FPGA where our own theorem says the ternary encoding
earns nothing, and it does not need to -- the gain comes from the absent decoder.
The posit decoder is taken from an open verification suite and brought to the same
fields through fp32, which is structurally fair but is not an optimised posit adder
datapath.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The full matrix: one ternary neuron across formats, scored against every theorem
Seven formats through one synthesised datapath on XC7A200T, differing in exactly
one block -- the activation decode. Adder, weight application and accumulator are
identical.
TEF 440 LUT 0 BRAM 80.73 MHz 0.183 MHz/LUT
binary32 479 0 78.23 0.163
binary16 552 0 62.69 0.114
posit8 560 0 43.55 0.078
takum16 817 57 64.16 0.079
posit16 774 0 36.30 0.047
posit32 955 0 28.33 0.030
The ordering is set by one column, scan-at-decode, and one property, the value
law, and the group boundaries fall exactly where the theorems put them. No scan
and a linear value law: TEF, binary32, binary16, spread 9-25% apart, which is
field widths rather than format. With a scan: posit at every width, 2.35x to 6.10x
behind in throughput per LUT, which is Theorem 14 -- decode cost is set by the
scan, not by the staircase. Logarithmic value law: takum has no scan but brings 57
block-RAM tiles to evaluate 2^f, which is Theorem 20.
The 9% between TEF and binary32 is Theorem 6 in action, and stating it is what
keeps the claim honest: on a binary fabric a packed ternary exponent never carries
more values per bit, so it earns nothing here and should not. The gap comes from
the absent decoder, not from ternarity. The defensible sentence is therefore not
"TEF is the best format" but "a fixed field beats a tapered one by 2.4-6.1x on a
ternary network, and TEF is the best of the fixed fields".
Where ternarity starts to pay is the last column, and it is computed rather than
synthesised because no such FPGA exists: on a fabric whose positions hold trits, a
binary exponent field is either dense and unaddable without a radix conversion, or
addable at one bit per trit and 25-50% wasted. TEF pays neither.
Twelve more formats are still routing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Eighteen formats through one ternary neuron, including the GF ladder
The matrix now covers every family in the catalogue that has RTL: the GF ladder
from arXiv:2606.05017 (GF10, GF14, GF+8, GFTernary), IEEE, VAX, fp8, minifloat,
int8, posit at three widths, takum, LNS and IBM hexadecimal. One synthesised
datapath, XC7A200T, no DSP, differing in exactly one block.
Ranked by throughput per LUT, the boundaries fall where the theorems put them and
not where a favourite would put them:
int8 448 LUT 84.86 MHz 0.189 no exponent field at all
TEF 440 80.73 0.183
GFTernary 466 82.51 0.177
binary32 479 78.23 0.163
... fixed fields cluster 0.109-0.183 ...
takum16 817 + 57 BRAM 64.16 0.079 logarithmic value law
posit8 560 43.55 0.078 regime scan
IBM hex32 683 49.71 0.073 radix 16
LNS16 659 43.11 0.065
posit16 774 36.30 0.047 regime scan
posit32 955 28.33 0.030 regime scan
int8 leads TEF by 3%, and the file says so first. It has no exponent field: its
decode is a sign extension and its entire dynamic range lives in the block scale.
That is a different trade, not a better format, and the defensible sentence is
that TEF leads among formats that carry an exponent.
The group boundaries are the finding. Fixed fields span 1.7x among themselves;
regime-scanning formats sit 2.4x to 6.4x behind; logarithmic value laws pay in
memory as well as logic. That is Theorem 14 and Theorem 20 measured on eighteen
points rather than argued on three.
Skills updated with the matrix and with three things not to say: not "TEF is the
best format" (int8 is 3% ahead), not "no multiplications" (the network removes
them, not the format), not "best for FPGA" (TEF and a binary fixed field measure
identically at 32 LUTs on the exponent decode).
Two synthesis traps recorded: takum's 64K tables route for hours and belong last,
and posit8_es2_decode instantiates posit16_decode, so yosys fails without both.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The TEF ladder in the matrix, and the unfair comparison it exposed
Adding TEF4/8/16/32 through the same datapath as every other format made my own
headline collapse, and the cause is a mistake in how I set the comparison up.
The first tnet_tef received a_off and a_mant already widened to the accumulator's
fields. It never decoded a packed word at all. Every competitor received a packed
word and decoded it. TEF was running a shortcut nobody else had, and the 440 LUT
against posit's 895 was not apples to apples.
Through the same fp32 path the ladder measures 469, 487, 495 and 499 LUTs for
TEF4/8/16/32, at 0.167, 0.147, 0.145 and 0.151 MHz per LUT. That puts TEF in the
MIDDLE of the fixed-field group, not at its head: int8 0.189, GFTernary 0.177,
TEF4 0.167, binary32 0.163, TEF32 0.151, TEF16 0.145. The spread inside the group
is 1.7x and it tracks field widths rather than family.
So "TEF is the best fixed field" does not survive a fair comparison, and the file
now says so where the claim used to be.
What survives intact is the group separation, which was always the actual finding:
fixed fields 0.109-0.189, regime-scanning formats 0.030-0.078, logarithmic value
laws paying in memory as well as logic. That is Theorem 14 and Theorem 20, and
neither depends on whose format leads within a group.
This is the seventh error of my own that a measurement has caught in this campaign
and the most consequential, because it was in the headline rather than in a
footnote.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The BNF/TNF control in silicon, and the frontier across all four families
BNF exists to measure what the ternary encoding is worth rather than to assert it,
and the pair differs in exactly one thing: the radix the exponent field is encoded
in. Through one datapath on XC7A200T, no DSP: BNF16 at 519 LUTs and 71.05 MHz
against TNF16 at 514 and 74.37. One percent in area, under five in frequency --
routing noise.
That is what the no-free-range theorem requires. A ternary exponent packed into
bits never carries more values per bit than a binary one, so on a binary fabric
the pair must tie. It ties, in silicon rather than only in arithmetic, and the
1.00x was computed before the synthesis rather than after.
The frontier across all four families is measured too, and the effective mantissa
matches the declared M within 0.15 on both axes, so the precision law holds on the
phi axis as well as the theorem axis.
One clean result falls out of it. GF and GF-T carry the SAME mantissa at every
rung -- 4 at 8 bits, 9 at 16, 19 at 32, 39 at 64 -- because both take
round((N-1)/phi^2) positions for the exponent. But GF-T's positions are trits, so
it spans 728 binades at 16 bits against GF's 62, and 531440 at 32 bits against
4094. On a ternary fabric, where a trit is a position, GF-T strictly dominates GF:
same width, same precision, 11.4x the range at 16 bits and 130x at 32. That is the
ternary encoding's payoff with no trade at all, and it lives on the phi axis rather
than on ours.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The frontier and the silicon, which contradict each other only apart
Two results that read as opposed until they are put side by side, because they
answer different questions.
By the numbers, TNF is not a format at all. It is a one-parameter family: the
width rule 1 + E_t + M = N, swept over E_t, gives a chain of points every one of
which lands exactly with no position unspent. At 16 bits that is twelve points
from (13.06, 8) to (2.13, 1594322), and all seven catalogued competitors are
strictly dominated -- binary16, gf16, posit16, bfloat16, afp, tekum16, takum16, no
exceptions. At 8 bits the same, seven competitors, none surviving.
That is not a trick of density. Each competitor leaves something on the table for
one of three reasons: a taper, which measures below its declared mantissa; unspent
positions, which the historical GF-T had at 2, 4 and 6 and the width rule
reclaims; or a binary exponent at equal position count, where a trit carries
log2(3) bits against a bit's one. Every competitor loses on at least one. The
family loses on none.
The wide classes needed the probe's ceiling removed -- a binary search rather than
a linear scan to 400 -- and TNF then holds the top-precision point at every width:
11.20 at 16 bits, 25.13 at 32, 55.82 at 64 against binary64's 52.27, 119.12 at 128
against binary128's 111.98. GF-T's ranges at 64 and 128 still saturate the probe
and are marked as such.
By the silicon, TNF sits mid-pack. Through one ternary-neuron datapath across 21
formats it measures 0.145 to 0.167 MHz per LUT against int8's 0.189 and
binary32's 0.163.
The two do not conflict. The frontier asks what a word of a given width carries,
which is a property of the number, and there the family dominates because it
wastes nothing. The silicon asks what it costs to unpack that word, which is a
property of the decoder, and there every fixed field is nearly equal because they
all decode by reading fields. The gap in silicon runs between fixed and tapered,
2.4x to 6.4x, and it is Theorem 14 and Theorem 20 rather than anything about us.
So the honest position is that we win on the axis of the number and do not win on
the axis of the decoder, because on that axis there is nothing to win -- and the
file says both, along with what may not be claimed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(research): dominance verified across all five width classes
28 catalogued competitors at 8/16/32/64/128 bits, zero survive the
one-parameter TNF family. Range is analytic for a fixed field; the
boundary check is impossible above E_t~30 (2^1.4e23) and is recorded
as a limit of the measurement rather than dropped.
Adds the frontier theorem and no-survivors corollary to the paper,
with the silicon reconciliation stated inline: dominance in
(M_eff, range) is not dominance in LUTs, and the paper says so.
Refs #516
* docs(skills): record what the frontier result does and does not close
The one-parameter TNF family dominates all 28 catalogued competitors
across five width classes. That closes the number axis and closes
nothing else: silicon puts TNF mid-pack among fixed fields, the block
axis (MXFP4) is still uncovered, and a dense family dominates any
finite point set for free unless each competitor's shortfall is a
theorem. The publication stop rule therefore stands.
Also records the killed boundary experiment rather than dropping it,
and fixes the stale arxiv_tef path.
Refs #516
* paper: retitle to Ternary Network Floats and prove the trit is the slack
The no-survivors list becomes a statement about the design space. The
family covers (m,b) iff m + log_3(b+1) <= N-1; a uniform format obeys
m + log_2(b+1) <= N-1 in its own accounting. Same books, two
currencies, and the exchange rate is the whole result: every uniform
binary format is covered with slack E*(1 - log_3 2) = 0.3691*E
positions, growing linearly in the exponent width it spends.
This is stated against T6, which says a packed ternary exponent
carries no more per bit and which silicon confirmed (BNF16 vs TNF16,
1% apart). Both hold on different axes: per position the trit spans
log2(3) binades, per bit of binary fabric the packing gives it back.
The bound is tight -- posit16 sits within 1% -- and not vacuous: the
same posit16 scored where a taper is designed to live violates it at
16.30 against 15. So escape needs non-uniformity, and there are
exactly two routes: taper, or a scale outside the word. Both are
named as frontier, not as territory held.
Abstract rewritten to arXiv's 1826/1920 character metadata limit with
no TeX-isms, following the motivation/problem/approach/results
structure; abstract_arxiv.txt is the submission field verbatim.
Refs #516
* paper: width-rule optimality, regret, composition, and the block gap
Three results the no-survivors table needed to stand on.
Theorem (optimal member): once a workload's visited range is named,
E_t* = ceil(log_3(b+1)) maximises the mantissa uniquely. The rule has
no free parameter, so it names a winner before measurement -- which is
what makes it falsifiable.
Theorem (regret): mis-sizing is asymmetric. Over-sizing costs precision
logarithmically on every value; under-sizing costs range linearly on
the tail. A wider exponent is not the conservative choice by default.
Theorem (composition): transforms and formats act on different
arguments, so their order matters -- measured at 2.78x. Recorded
because an earlier version of this work reported the levers as nearly
multiplicative, computing a ratio against a losing baseline.
Related work now states where the block literature is not looking: the
2026 papers vary the transform (learnable block optimisation,
format-aware rounding, residual channels) and hold the element at E2M1.
That is the variable the width rule speaks to.
Six dangling cross-references repointed to their real labels.
Refs #516
* research(block): MXFP4 was given six magnitudes where the spec gives eight
fp_levels reserved the top exponent code for Inf/NaN as IEEE binary
formats do. OCP microscaling element formats reserve nothing -- every
E2M1 code is finite -- so the reservation cost the standard the values
4.0 and 6.0 and ran it at 2.58 bits against int4's 3.
Uncorrected, the run reads 'uniform int4 beats MXFP4 by 27%'. It does
not; the rows had different level counts.
Same class as the earlier amax defect: both underfed the competitor and
both would have produced a headline in our favour. Adopted rule -- a
competitor's cardinality is a specification fact, asserted against the
spec in the harness, never derived from a shared helper.
Refs #516
* paper: name the first principles -- KKT, radix economy, Kraft
The width rule was stated as a rule. It is the solution of a
constrained maximisation: maximise M subject to 1+E+M=N and r^E >= b+1.
The multipliers come out lambda = mu = 1, so complementary slackness
makes the range constraint active -- the exponent is exactly as wide as
the range demands and not one position wider. Theorem (optimal member)
is now derived, and lambda=1 is the shadow price: one position of
exponent costs exactly one of mantissa.
Radix economy rho(r)=r/ln r, minimal at e, explains why three and not
two -- a 1950s result we claim no credit for. Its use here is the
reconciliation of our two axes, which are the two factors of
r * log_r(V): the number axis counts positions and ternary wins
unconditionally by 0.3691*E; the silicon axis restores the cost per
position and charges kappa(3)/kappa(2)=1.68, which is why BNF16 and
TNF16 land within 1%. T6 and the slack corollary are one product
reported factor by factor, not a contradiction.
Least action and conservation of energy are named as resemblances and
explicitly not used as arguments: no time, no trajectory, and 1+E+M=N
is a designer's budget, not a law of nature.
Applied to the block axis the multiplier problem returns E2M1 at the
measured 99th-percentile within-block span of 3.04 binades -- the OCP
microscaling element, derived rather than beaten.
Refs #516
* research(ternary): the accumulator law, derived and measured
In a ternary layer the weight has no format -- 1.58 bits, no exponent,
no mantissa, no multiply. Every published ternary method quantises the
weight, so none of them was ever a competitor to a number format; the
object a format describes here is the accumulator, and that niche is
empty.
Law: range visited grows as log2(pK), error after pK roundings as
sqrt(pK)*2^-(M+1), so through the KKT solution E* = ceil(log_3 B(K))
and the mantissa side outruns the exponent side.
Measured on real ternarised weights: fitted exponent in K is +0.476 and
+0.435 against a predicted +0.5. Et=2 fits at +0.207 and that is the
range constraint being active, not a refutation -- 9 binades cannot
hold an accumulator visiting 13.9, so its curve is saturation-dominated
then rounding-dominated. The measured span picks Et=3 by the rule, and
the measurement picks Et=3 independently at every fan-in.
Two more self-caught defects recorded. The first instrument normalised
each value by its own magnitude and returned error identically zero
everywhere -- a format that never rounds is not a format. The second
built level sets from the position count alone, ignoring that binary
fabric addresses only 2^(N-1) magnitudes, and produced a 100x
perplexity artefact against our own format.
Refs #516
* paper: lead with the datapath, and split trainable from rejected
Reframed on what the measurements actually support. A ternary node has
three format-bearing sites and only one needs a format: the weight is a
code (sign-select, not multiply), the sample is ADC-native, and the
accumulator is the only object with a range to spend. We close both
halves that need closing -- GFTernary the weight, TNF the accumulator --
and no other pair in the literature does.
New theorem: the golden alphabet is unique. Requiring the product of two
weights to fall back into the lattice the datapath already adds in means
r^2 = r + 1, whose only positive root is phi. Corollary: a k-layer gain
is exactly F_k*phi + F_(k-1), two integers, so rescaling between layers
is shift-and-add and depth never reintroduces a multiplier. This is what
separates the phi alphabet from {-1,0,+1}, which needs a learned real
alpha per layer -- and multiplying by alpha puts the DSP back.
Comparison split in two by a checkable test: does the word carry the
range, or must it be brought from outside? int8, int4, integer
accumulators and the bare E2M1 element go in the rejected table -- e=0,
no range in the word, and every published sub-8-bit training result
using them carries block scales AND a higher-precision master weight.
int8's 0.189 MHz/LUT is reported there, not as a competitor: it leads by
declining the task. GFTernary stays in the main table despite e=0
because its scale is intrinsic and exact rather than external and
learned.
Conclusion stated at the strength the measurements carry: for a ternary
datapath the pair {GFTernary, TNF} is a reference format -- complete,
forced rather than chosen, and predictive. Explicitly not a claim about
datapaths that multiply, where the block axis holds the ground.
Abstract rewritten to exactly 1920/1920 arXiv metadata characters.
Refs #516
* research(block): the block axis is decided, against us
MXFP4 21.94 and MXFP6 14.73 beat TNF4 36.72 and TNF6 18.03 on
wikitext-2 with the MX spec's own E8M0 scale and a verified baseline.
The reason is structural, not tuning: 3^E_t never divides 2^k, so a
ternary exponent packed into a binary word always wastes codes, and
where the alphabet is short the waste dominates. TNF4 gets 7 of 8
magnitudes against E2M1's 8, and 3 binades against 4 -- strictly worse
on both counts at once. At 6 bits the loss reaches 41% of the alphabet.
This is T6 carried to its conclusion: on the number axis a trit is a
position and wins; in a packed word the codes are counted and it pays.
Two things survive. The KKT law returned E2M1 -- a binary exponent --
from the measured within-block span, recommending the industry standard
over our own family; a rule that only ever recommends its author is not
a rule. And the range constraint is visibly active: TNF6 E_t=2 with 19
magnitudes beats E_t=1 with 25, so fewer levels with more range wins,
which is complementary slackness appearing in perplexity.
The reference-format claim is bounded accordingly: it is a claim about
ternary datapaths without multipliers, not about block-scaled binary
ones. The publication stop rule stands -- the block axis was the named
condition and the measurement went against us.
Refs #516
* paper: place the work inside a sixty-eight-year argument
Radix economy and the case for three are not ours, and saying whose
they are makes our contribution smaller and defensible. Fowler's
mechanical balanced-ternary machine around 1840; Brusentsov and Sobolev
building Setun at Moscow State University in 1958 on paired ferrite
cores, ~50 machines 1959-65; Setun-70 in 1970 anticipating RISC
arguments and ended administratively; Knuth keeping the idea alive;
CNTFET, memristor and photonic ternary devices continuing it.
New theorem states the boundary rather than the claim. Written as
cost = r x log_r V: where the position is physical, only the second
factor is compared and ternary gains 0.3691E positions unconditionally;
where the position must be encoded in bits, the format is bounded by
3^E_t * 2^M <= 2^(N-1) and, since 3^E_t never divides a power of two,
the remainder is lost -- 25% at 4 bits, 15.6% at 6 and 8, 5.1% at 16.
We add no support to 'ternary beats binary' as a general statement. We
measured it three times and it went against us each time. What we add
is the condition under which the old argument applies, which reads as a
prescription: it says what fabric must exist for the advantage to be
collected, and that fabric is not the one currently purchasable.
Parhami's binary-encoded balanced ternary anticipated the mechanism and
is credited; our part is measuring its cost in a live network and in
placed-and-routed silicon rather than in operation counts.
Refs #516
* paper: complete the truncated table and withdraw stale headline numbers
Audit of every table against its own caption found two defects.
tab:carryrange claimed 21 formats and showed six rows with an ellipsis.
Now complete: 20 rows with storage, LUTs, Fmax and MHz/LUT, int8
excluded and pointed to the rejected table with the reason.
tab:tnet still carried 440 against 895 LUTs at 0.184 MHz/LUT -- the
comparison this same paper retracts elsewhere, where TNF received
pre-widened fields while the competitors unpacked theirs. Replaced with
matched width against matched width on packed words: 3.1x at 16 bits
and 5.6x at 32, with the withdrawal stated in the caption rather than
buried. The gap belongs to the regime scan, not to the ladder.
Closing section sets the two independent lines side by side: the
arithmetic one, 68 years old, which stalls because 3^E_t never divides
a power of two; and the geometric one, where requiring the product of
two weight symbols to fall into the sum the datapath already forms
gives r^2 = r + 1, whose single positive root determines the alphabet.
Stated as a reading of the results, with every component measured
above, and what they amount to together left to the reader.
Refs #516
* research(families): the two axes return the same format, measured
GF8 E=3 and BNF8 E=3 both give perplexity 14.6130 -- the same number,
not a close one. The golden-ratio rule E = round((N-1)/phi^2) and the
width rule 1+E+M=N were derived independently and for unrelated
reasons, and at eight bits on this workload they name the identical
format. Neither derivation predicted that.
Twelfth self-caught defect, and it falsifies our own prediction: the
width rule named BNF8 E=4 and TNF8 E_t=3, and the winners were E=3 and
E_t=2 -- both predictions one step too wide. The rule's form survives
and is visible in the sweep (single optimum, asymmetric penalty exactly
as the regret theorem states: under-sizing gives 4.5 million, over-
sizing costs 0.3%). What was wrong is the estimator of the visited
range: we measured 0.1st percentile to maximum, crediting a tail that
carries almost no energy. Recorded rather than quietly re-tuned.
Ternary loses on binary fabric for the third independent time: GF-T8
carries 109 magnitudes against GF8's 129 and pays 15.51 against 14.61.
Refs #516
* paper: the multiply-free path is not merely cheap, it is exact
Applying a phi weight to an integer pair (a,b) representing a + b*phi
is (a,b) -> (b, a+b): the Fibonacci recurrence, one integer addition,
no shift. Z[phi] is a ring and the alphabet lies inside it, so for
inputs in Z[phi] the entire linear part of a ternary network -- every
weight application and every accumulation, to arbitrary fan-in and
depth -- stays in Z[phi] and is computed with no rounding error at all.
This is a different kind of claim from the rest of the paper. Elsewhere
we compare error magnitudes between formats; here there is no error to
compare. Measured at fan-in 512 the integer pair reproduces the real
sum to the precision of the checker, not of the datapath. Components
grow logarithmically, eight bits over those 512 terms, and the cost is
two integer accumulators instead of one float.
The base is a minimum rather than a choice: closure needs r^2 = pr + q
with integer p,q, and any p > 1 adds a shift to the addition. p=q=1
gives phi. 1+sqrt(2) satisfies r^2=2r+1 and pays the shift; sqrt(2) has
r^2=2 and loses the scale out of the lattice.
Scope stated so the claim does not overreach: this is arithmetic in a
lattice. It covers the linear algebra that dominates a network's work
and its DSP cost, and says nothing about control flow or addressing.
Refs #516
* research(phi): the phi scale grid halves the cost of removing the multiplier
BitNet stores ternary weights plus a real per-layer scale alpha =
mean|W|, and multiplying by that alpha puts the multiplier back at the
layer boundary. Snapping the scale to a grid removes it. The phi grid
is denser than powers of two by log(2)/log(phi) = 1.440 at the same
cost class, so the prediction made before measuring was that its excess
error over the unreachable exact alpha would be about half.
Measured over 210 layers: exact 0.476781, phi^k 0.488424 (+2.4420%),
2^k 0.499943 (+4.8579%). Ratio of excesses 0.501 against a predicted
0.500. phi wins 163 of 210 layers -- not all, since a layer whose
optimum lands near a power of two is better served by the coarser grid.
Together with dot_exact this closes the multiplier out of the entire
layer: weights, accumulation, and now the scale.
Defect #13 recorded rather than reported as a result. The first attempt
asked this through perplexity, where post-hoc ternarisation destroys a
model not trained for it: every arm including BitNet's exact alpha
landed at ppl ~2.2e7 against a baseline of 14.49. Read naively that
says phi is refuted by 2x. It says nothing -- the tell was that the
control arm was destroyed too, and a comparison whose control fails is
not a comparison.
Refs #516
* fpga(phiscale): what the layer scale costs with and without a multiplier
BitNet's per-layer alpha = mean|W| is a real number, so applying it is
a genuine multiply -- the multiplier the ternary weights removed comes
back at the layer boundary. Carrying the value as an integer pair makes
the scale phi^k into k Fibonacci steps, one adder each.
Synthesised through yosys synth_xilinx: the multiplier arm costs 2
DSP48 blocks, or 1215 LUTs with DSP inference off. The phi arm costs
171 LUTs and zero DSP, and is DSP-invariant -- identical numbers with
and without, because there is no multiply to map.
Costs stated rather than omitted: 4x the registers (135 FF against 33),
k cycles instead of 1 (about 8 for a typical alpha, so ~1.6% of a
fan-in-512 layer), no Fmax because nextpnr-xilinx is not on this
machine, and only k >= 0 -- the inverse step is (a,b) -> (b-a, a) since
phi^-1 = phi - 1, one subtraction, not yet built.
Area reported only after correctness: 200 randomised cases against a
golden model computed independently in the testbench, 0 errors, with a
negative control confirming the bench can detect a mismatch. A circuit
that computes the wrong thing is smaller still.
Refs #516
* fpga(phiscale): name which direction was actually built
Real layer scales are below one -- alpha = mean|W| ~ 0.02 gives
k ~ -8 -- so a deployed layer needs the inverse Fibonacci step
(a,b) -> (b-a, a), from phi^-1 = phi - 1. That is one subtraction where
the built circuit does one addition, so the area transfers and the
count stands, but the forward direction is what was synthesised and
simulated. Recorded because a reader would otherwise take the measured
circuit for the deployable one.
* research: APoT refutes our scale-grid claim, and exactness was never ours
Two of our own claims are withdrawn or bounded, both by measuring a
competitor properly rather than by a new experiment.
The phi^k grid does beat 2^k by exactly the predicted density ratio,
and that reproduces. It was the wrong baseline. The deployed state of
the art for multiplier-free scales is APoT -- sums of power-of-two
terms, ICLR 2020. On the same 210 layers: APoT-2 costs 0.1651% excess
against exact alpha where ours costs 2.4420%, and it does so in one
cycle against our k. APoT-3 costs 0.0054%. Ours wins 17 layers of 210
against APoT-2. The claim that ours is the right grid for a
multiplier-free scale is withdrawn.
dot_exact is machine-checked and remains true, but the significance
attached to it does not survive. An APoT scale is a dyadic rational,
and the dyadic rationals are also a ring closed under the datapath's
operations: checked in exact arithmetic, a fan-in-512 layer with an
APoT-2 scale and Q8 inputs lands on a denominator of 2^15, exact in
ordinary fixed point. Binary has had this since fixed point existed.
What remains unique to phi is term-count non-growth -- a Z[phi] value
stays a two-component pair through any depth, while APoT-2 squared has
four terms and the count compounds. Stated as a conjecture; it needs an
experiment, not an assertion.
Also lands the deployable signed scale path, closing the hole named in
the previous iteration: both directions, since real layer scales give
k about -8. 360 randomised cases including a multiply-then-divide
round-trip identity, 0 errors. Costs two LUTs over the forward-only
version, 173 against 171, still zero DSP and still DSP-invariant.
Refs #516
* research(scale): three vertices, no winner, and the shifter we ignored
Measuring area rather than only accuracy shows the scale applier was
never a one-dimensional question. phi^k costs 173 LUTs, APoT-2 costs
384, the multiplier 1215 or 2 DSP48. Ratio phi to APoT is 2.22x at
d=0, 3.81x at d=1, 7.03x at d=2.
The reason is not the one predicted. Term-count non-growth is the
second effect; the first is that APoT's shift is by a runtime value, so
it is a barrel shifter at 192 LUTs each on 32 bits. We had been calling
APoT shift-add and treating that as free. A constant shift is free; a
variable one is a barrel, and a scale serving any layer is variable by
construction.
Four theorems: term growth n^(d+1) against a constant pair; shifter
cost Theta(W log W) against zero for a fixed permutation; a crossover
depth d* = 2 at which APoT costs what it was introduced to avoid,
measured at 1217 against the multiplier's 1215; and no domination in
(area, latency, error), all three being vertices.
The accuracy withdrawal stands. APoT-2 remains 15x more accurate at one
cycle against k, and anyone bound by accuracy or latency should use it.
The honest output is a hybrid rule rather than a winner -- the previous
two iterations each defended a single answer and each was wrong in a
different direction.
Refs #516
* paper: the scale applier has three vertices, and two of our claims move
Records both corrections in place rather than quietly restating the
position.
Withdrawn: that the phi^k grid is the right one for a multiplier-free
scale. The 0.501 density ratio reproduces and is correct, but the
baseline was wrong -- APoT is what the field deploys, and it costs
0.1651% against our 2.4420%, in one cycle against k.
Bounded: dot_exact is machine-checked and true, but an APoT scale is a
dyadic rational and Z[1/2] is also a closed ring, so a fan-in-512 layer
with an APoT-2 scale and Q8 inputs is already exact in ordinary fixed
point. Binary has had this since fixed point existed.
Four new theorems from measuring area rather than accuracy: term growth
n^(d+1) against a constant pair; shifter cost Theta(W log W) against
zero, which is the effect we missed and is larger than the one we
predicted -- APoT's shift is by a runtime value, so it is a barrel at
192 LUTs each on 32 bits, while a Fibonacci step has no shifter at all;
a crossover at d* = 2 where APoT costs what it was introduced to avoid;
and no domination in (area, latency, error).
Output is a rule, not a winner: phi^k where the path is area-bound or
composed without requantisation, APoT where it is latency- or
accuracy-bound. Stated that way because the two previous iterations
each defended a single answer and each was wrong differently.
Refs #516
* research: the area ordering inverts with regime, and ours was the wrong one
Two attacks on our own area claim, both successful.
Freeze the scale and APoT's shifts become compile-time constants --
wiring, not logic. Measured: APoT-2 at 26 LUTs against an unrolled
phi^k at 64, 128, 256 for K = 2, 4, 8. An unrolled recurrence is linear
in K at about 32 LUTs a step; a constant-shift applier is one adder
regardless. The lines never cross, not even at K = 1.
Then the barrel. Our regime-2 advantage used a 5-bit shift field, but
across 210 layers the scales span only 3.15 octaves, so two bits
suffice. At SW=2 APoT costs 130 LUTs against our 199. The 2.22x
advantage was an artefact of giving the competitor a wider field than
it needs, and it is withdrawn.
Instrument limitation recorded: the APoT sweep is non-monotonic (130,
380, 230, 384) with the parameter demonstrably applied and yosys
deterministic, so this is abc mapping heuristics. LUT counts from yosys
alone are not reliable at this granularity -- a claim resting on 30%
between two such points is unsafe; one resting on regime 1's 10x is.
Three theorems: an unrolled recurrence is Theta(KW) against Theta(W)
for constant shifts, so it loses at every K; the ordering inverts with
compile-time versus runtime scale, so neither family is better and the
architecture decides; and a barrel is priced by the workload's range,
not by a convenient field width.
What survives is narrower: phi^k is the only applier whose area is
independent of composition depth, which is real but uncommon. The
alphabet uniqueness and Z[phi] closure are machine-checked and were
never area claims.
Refs #516
* paper: withdraw the area advantage as well, by our own attack
Three theorems and one withdrawal. An unrolled recurrence is Theta(KW)
where a constant-shift additive applier is Theta(W), so with a frozen
scale the recurrence loses at every K -- measured 26 LUTs against 64,
128, 256 at K = 2, 4, 8. The ordering therefore inverts with regime and
is a property of neither family. And a barrel is priced by the
workload's range: 3.15 octaves across 210 layers means two bits, where
APoT costs 130 against our 199, so the 2.22x was an artefact of a
five-bit field.
Instrument limitation stated rather than smoothed: the APoT sweep is
non-monotonic with the parameter applied and the tool deterministic, so
logic-synthesis LUT counts are not reliable at this granularity.
What survives is term-growth independence alone, in a regime that is
real but uncommon. The alphabet uniqueness and Z[phi] closure are
untouched, having never been area claims.
Refs #516
* research: the operating point, which is what survives five withdrawals
Z[phi] is the only one of three systems cheap under BOTH operations.
LNS makes multiplication free and pays for addition with a log(1+2^x)
table -- our own takum32_decode measurement, 10,967 LUT and 84 RAMB36.
Fixed point and APoT make addition cheap and pay for scale
multiplication unless the scale is frozen. Z[phi] costs one adder for
multiplication by a power of phi and 64 LUTs for componentwise
addition.
The price is that its free multiplication is restricted to powers of
phi. Four theorems say why that is not a restriction here: in a
datapath where weights are codes applied by sign-select and the scale
is a power of the base, the required multiplication set is exactly
{base^k}; a ring closed under addition and under multiplication by
generator powers is sufficient; LNS is over-provisioned and fixed point
under-provisioned for that profile; and compile-time composition is
free, so term growth occurs only where d or the scales are runtime
quantities.
That last theorem removes three of the four depth cases we had claimed:
low-rank W=UV, folded conv+BN and residual branch scalars are all known
after training, so the product is precomputed and no composition
happens in hardware. One survives -- accumulation along a mesh route,
where the hop count is runtime. That is the tri-net datapath, not the
network.
States its own boundary rather than waiting to be asked: not an area
win, not an accuracy win, and nothing to beat where the scale is
frozen. A claim about which system matches a datapath's operation
profile.
Refs #516
* paper: the operating point, and four theorems that bound it
Z[phi] is the only one of three systems cheap under both operations.
LNS buys free multiplication and pays for addition with a log(1+2^x)
table -- our own takum32_decode figure, 10,967 LUT and 84 RAMB36.
Fixed point and APoT buy cheap addition and pay for scale
multiplication unless frozen.
Four results say why the restriction to powers of phi costs nothing
here: the required multiplication set of a sign-select datapath is
exactly {base^k}; a ring closed under addition and generator-power
multiplication suffices; LNS is over-provisioned and fixed point
under-provisioned for that profile; and compile-time composition is
free, which removes three of the four depth cases we had claimed and
leaves only mesh-route accumulation, where the hop count is runtime.
Boundaries stated in the section rather than extracted later: not an
area win, not an accuracy win, and nothing to win where the scale is
frozen.
* research: seven iterations, seven withdrawals, and what is left
An accounting. Every phi claim about silicon was tested tonight, mostly
by us, and none survived.
Withdrawal 6: our LNS row cited takum32_decode at 10,967 LUT. That is a
format decoder, not an adder. An honest LNS-32 adder with a
4096-entry table costs 275 LUT -- we were off by two orders of
magnitude, and the over-provisioning theorem shrinks from 170x to 8.6x
at matched storage.
Withdrawal 7: the mesh case, the last place a depth advantage could
live. Matched combinational comparison gives APoT requantisation 103
LUT against a Fibonacci step at 128 -- phi loses by 25%. Reached first
through an unmatched comparison reading 0.60x against us, because our
side carried a controller the other did not. Same defect as the six
wins before it, pointed the other way. A loss deserves the same audit
as a win.
What survives: the machine-checked mathematics, untouched and never a
hardware claim; the LNS comparison rebuilt honestly, where Z[phi]
addition is 32 LUT against 275 and LNS additionally cannot represent
zero while a ternary alphabet is 46% zeros; the number-axis frontier;
and zero DSP, which belongs to ternary weights rather than to phi.
Closing theorem: the terms of a comparison are part of its result.
Seven times a ratio changed sign or magnitude when the competitor was
rebuilt as its own advocate would build it. The measurements were never
wrong; the comparisons were.
Refs #516
* research(frontier): withdrawal 8 -- the headline was a positions-vs-bits artefact
The claim that 28 catalogued formats are dominated with slack 0.3691*E
counts positions. A format is stored in bits, and a realisable member
obeys 3^Et * 2^M <= 2^(N-1). Substituting the minimal Et gives
m + log2(b+1) <= N-1, identically the uniform binary budget, since
log2(3) log3(x) = log2(x). The slack vanishes exactly, for every E.
Measured at equal storage: 6 of 17 dominated, not 28 of 28. Uniform
binary formats that spend all their bits tie exactly -- binary32 at
31.10 against 31, binary64 at 63.27 against 63. Those dominated are
wasting budget: posit32 by 5.12, posit64 by 14.24, takum32 by 0.82,
cray_float by 0.65, and our own gf8 by 0.70. Several tapered formats
escape outright -- posit16 by 2.37, takum16 by 1.53 -- which is
consistent with our own corollary that escape requires non-uniformity.
Corrected statement: at equal storage the family ties every uniform
binary format spending all its bits, dominates those wasting budget,
and is escaped by tapers measured where they concentrate precision.
The same defect was recorded twice yesterday, as T6 and as the level
table where TNF4 with Et=2 does not exist in four bits. Both were
treated as local facts about packing and neither was carried back to
the headline. Rule adopted: a correction that invalidates a comparison
must be applied to every claim resting on the same quantity, not only
where it surfaced.
Refs #516
* paper: correct the headline -- the frontier at equal storage
The abstract claimed 28 catalogued formats dominated with slack
0.3691*E. That counts positions. A realisable member is stored in bits
and obeys 3^Et * 2^M <= 2^(N-1); substituting the minimal Et collapses
the condition to m + log2(b+1) <= N-1, identically the uniform binary
budget, since log2(3) log3(x) = log2(x).
Abstract rewritten to state both counts and which one holds. New
theorem and corollary in place, with the measured outcome: 6 of 17
dominated, uniform binary formats spending all their bits tie exactly,
and tapers escape -- consistent with our own corollary that escape
requires non-uniformity.
Also records why it went unseen: the same packing defect appears twice
elsewhere in the paper, as the no-free-range theorem and as TNF4 with
Et=2 not existing in four bits, and neither was carried back to the
result resting on the same quantity.
Refs #516
* research: withdrawal 9 -- the ternary rungs are wider than their names
Found by applying withdrawal 8's rule systematically rather than
locally, and it is the deepest of the nine.
The SSOT specifies each ternary rung as 1 + Et + M = N with Et in trits
and M in bits, and declares storage=uN. The oracle stores the exponent
offset as an integer in [0, 3^Et - 1], so the word is
1 + ceil(Et log2 3) + M bits. GF-T16 encodes into 20 bits, GF-T32 into
40, GF-T64 into 79. Confirmed by running the encoder. All nine ternary
rows carry a wrong storage field.
The RTL agrees with the oracle -- tnet_tef #(MW=25, OW=10) is 36 bits
for what …
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #7 (Phase 3: Synthesis + Demo)
Complete FPGA synthesis flow with benchmark results and release preparation.
Verilog — Synthesizable Top
vsa_matmul_top.v— Autoregressive ternary inference for QMTECH XC7A100T:vsa_matmulinstance: 64×64 ternary matmul, 0 DSP48[AA BB FE pass token best_score gen_count FF]vsa_matmul_top.xdc— Pin constraints for XC7A100T-1FGG676CRust — CLI + Benchmark
SynthConfig::vsa_matmul()— one-command VSA synthesistrios-fpga synth-vsa— full synthesis pipelinetrios-fpga bench— CPU benchmark with FPGA projectionBenchmark Results
Release + Paper
release/RELEASE-v0.2.md— release manifest with SHA-256, commandsdocs/arxiv-trinity-stack-draft.md— arXiv paper draft: Trinity Stack: φ-Structured GF16 Inference on $30 FPGAVerification
cargo test --workspace: 21/21 PASScargo clippy --workspace: 0 warningsφ² + φ⁻² = 3 | TRINITY