Repository navigation
Packed-quant byte-order (block layout) is an unwritten, contradictory contract across the engine and downstream converters #973
Description
Activity
- addedbugSomething isn't workingSomething isn't workingtensorsTensor operations and data structuresTensor operations and data structures
on Aug 12, 2026 This belongs under the broader tensor-storage unification effort tracked at #932 / SKEEP-003 (`docs/modules/skeep/pages/003-unified-tensor-storage.adoc`).
SKEEP-003's own hard constraint — "SKaiNET's packed-encoding system is ahead of comparable frameworks and must survive any refactor with bit-identical behavior" — and its dtype/encoding-coherence goal are exactly the gap this issue documents: there is no type-level representation of packed-quant byte order (canonical row-major vs. kernel-native input-block-major), so nothing can enforce "bit-identical" across a refactor because the two orderings are indistinguishable at the type level today. The 7-contradiction census here is concrete evidence for SKEEP-003's problem statement, and the
KernelPackedWeightData-as-distinct-type proposal in this issue is a narrower, packed-quant-specific instance of SKEEP-003's "dtype/encoding coherence" cross-cutting improvement (and possibly its "single view mechanism" one, givenops.transpose's role here).Recommend either:
- folding this issue's packed-quant byte-order analysis into SKEEP-003 as a concrete case study feeding whichever end-state (storage-first / data-first) gets chosen, or
- keeping this as a standalone issue but blocking/tracking it under TensorData and TensorStorage are parallel layers — unify the storage model (ownership, views, dtype/encoding, placement): SKEEP-003 discussion anchor #932 so the two efforts don't independently reinvent the same distinct-type mechanism.
Not proposing a specific choice here — that's the maintainer decision SKEEP-003 already frames as open.
Cross-reference from the SKEEP-003 kernel-dispatch migration (#1029, PR #1072):
The view-keyed dispatcher can now host platform-pack kernels, but the packed (Q4_0…Q6_K) SPI kernels could not be bridged, for exactly the reason this issue describes. They read their weight bytes block-major — the order
DefaultCpuOpsBase.transposePackedBlocksproduces — while a packedTensorViewdescribes the file's canonical row-major block order. Registering them under a view key without a written byte-order contract would risk the silent-wrong-numbers class of #968/#971 rather than a loud failure.So #1029 bridges only the dense FP32 kernel and leaves the packed fast paths on their existing ladder, with the registry serving packed operands through the decoding reference kernel (correct for any layout, slower). This issue is the prerequisite for finishing that migration: once the block order is pinned down and asserted, the packed kernels can be registered under view keys and the ladders in
DefaultCpuOps/DefaultCpuOpsJvmdeleted under the golden parity gate.- added sub-issues
on Aug 24, 2026 - added 5 commits that reference this issue
on Aug 24, 2026 Closing: the contract is written down, in the type system and on paper
All five slices are merged — #1094, #1095, #1097, #1098, #1096 — and the normative document lives at
docs/design/memory/packed-weight-layout.md.The principle this issue asked for was "one convention, one owner, explicit in the type, loud on violation." Here is where each part landed.
Explicit in the type.
Layout.blockOrdercarriesROW_MAJORorINPUT_BLOCK_MAJOR, and the order is expressed in the strides rather than in a branch — sonarrow,transpose,getandtoFloatArraykeep working on a view in either order.LayoutClasssplits the same way, so a kernel declares the order it reads in itsKernelKeyinstead of assuming it (#1094).One owner.
PackedWeightsis the only sanctioned implementation of the permutation, for views and for raw bytes, and it is idempotent by construction (#1097).Loud on violation.
ops.transposethrows for a heap packed weight, naming the primitive and the reason;requireOutInrefuses a weight whose label looks transposed rather than permuting the wrong grid (#1096, #1098).The operation, renamed to what it is.
matmulWeightTransposed(x, W[out, in])is the primitive ggml and BLAS actually have, andDefaultCpuOpsrelayouts once per weight rather than once per call — which removes the per-forward copy #969 introduced (#1096).The census, seven items later
# Contradiction Status 1 Q5_0/Q5_1TensorDatakdoc claimed input-block-major bytes and a shape-swap transpose✅ corrected, pointing at the normative doc 2 Kernel SPI kdoc claimed byte-identity with a storage type holding the other order ✅ corrected 3 Heap tier kernel-native vs MemSeg canonical, distinguished only by marker interfaces ✅ stated in code: the refusal is scoped to the heap tier because MemSeg genuinely reads canonical bytes 4 Q4_K-over-MemorySegment: canonical dead code vs kernel-native live ⚠️ not done — I did not find the dead implementation the census names; whoever knows where it is should delete it5 Two competing declared-shape conventions in the engine's own tests ✅ each of the 27 migrated call sites now names which operation it wants 6 GGUF neorder vstransposePackedBlocks's[out, in]assumption✅ WeightOrientation.OUT_INat the boundary, plus the guard7 toFloatArray()/get()hard-wired canonical◐ partial: correct for views (the order is in the layout, and RelayoutedBlockDecoderkeepsget()honest across a relayout); a legacyQ*BlockTensorDataholding kernel-order bytes still decodes as if canonical. It is now at least named: onlyrelayoutPackedWeightForKernelsproduces one, so a caller knows what they holdBoth latent hazards are gone: double application is impossible (the prepack is idempotent, and
transposerefuses), and the per-forward copy is a per-weight one.What this does not fix
- Downstream. The break is deliberate and confirmed acceptable — versions are pinned. SKaiNET-transformers needs to move to
matmulWeightTransposed/PackedWeights, andPackedLayoutFixturesis published so its tests can assert against the same bytes this repository asserts against. That is the guardrail against a byte-layout change shipping green on both sides again. - Census item 4, above.
- Making
WeightOrientation.OUT_INthe default, which is a second breaking change and wants downstream to have moved first.
Why it is worth the churn
Two of the seven contradictions had already produced wrong numbers in production six weeks apart (#968, #971), and neither failed loudly: mixing the two block orders yields a block-permuted matrix of finite, plausible values. Every test added here uses a three-block-wide weight for that reason — at one block per row the two orders coincide, and a test built that way passes whichever convention the code happens to hold. That is the shape that hid the original bug.
Closing.
- Downstream. The break is deliberate and confirmed acceptable — versions are pinned. SKaiNET-transformers needs to move to
Summary
The 0.40.1 hotfix (#969, closing #968) is locally correct and globally wrong: it fixed
ops.transposefor one population of packed-quant tensors (canonical/row-major bytes) while silently corrupting the other population that legitimately exists today (already kernel-native bytes). This just caused a real downstream regression, independently reproduced and fixed at the call-site level in SKaiNET-transformers#311 — but that PR is a point fix, not a fix for the underlying gap. Filing this to track the actual root cause: nothing in the type system, API contract, or test suite says which physical byte order aQ*BlockTensorData.packedDataholds, and the engine's own code contradicts itself about it in at least two places.Full analysis, evidence index, and a proposed structural design live in
packed-quant-layout-postmortem.md, written this session while investigating the transformers-side regression (happy to attach/paste the full doc into this issue or a linked gist if useful — trimmed here to the engine-actionable parts).The regression that surfaced this (already fixed downstream)
Two block orderings exist for a 2-D
[out, in]packed weight's quant blocks (grid =outDim × blocksPerRow):o * blocksPerRow + b. What GGUF stores, whatQ4_0Quantizeremits, whattoFloatArray()assumes.b * outDim + o. What every heap matmul kernel (scalar, Panama, native C, JNI) actually reads.#969's fix makes
ops.transposeperform a real canonical→kernel-native block-grid permutation — correct for canonical input. ButSKaiNET-transformers' classic packed path (BlockQuantPacking.pack(), pre-#311) eagerly relayouts GGUF bytes to kernel-native at load time, relying on the pre-0.40.1 shape-swap-only transpose to pass them through unchanged at forward time. On 0.40.1, that transpose now applies the canonical→kernel-native permutation to bytes that are already kernel-native → double permutation → garbage (not a crash —transpose(transpose(W)) ≠ Wfor non-square block grids, so nothing detects it).Both sides were "right" by their own local documentation. The engine's own type kdoc (
Q5_1TensorData.kt:24-27, still true in 0.40.1) states blocks are input-block-major and the CPU-ops lazy transpose is a pure shape swap. The 0.40.1 transpose's own comment says the opposite. The engine now contradicts its own type contract, and no test crosses the repo boundary to catch it — engine CI builds canonical-only fixtures (NativeLazyTransposeGroundTruthReproTestfrom #968), transformers CI (pre-#311) built kernel-native fixtures. Each suite proved its own convention; neither was wrong about its own inputs.Census: 7 contradicting conventions already live in this repo alone
Q5_0/Q5_1TensorDatakdoc says bytes are kernel-native and transpose is a shape swap; the in-repo GGUF loader emits canonical for those same types, and transpose now permutesQ5_1TensorData.kt:24-27,Q5_0TensorData.kt:21-22vsStreamingGgufParametersLoader.kt:195-206Q5_0MatmulKernel.kt:29"Matches Q5_0BlockTensorData.packedData."JvmQuantizedVectorKernels.kt:629,816(canonical) vs heap kernels;DefaultCpuOpsJvm.kt:211-226JvmQuantizedVectorKernels.kt:667vsQ4KMemSegMatmulKernel.kt:37+q4k_matmul.c:245[in,out]+ kernel-native, no transpose vs[out,in]+ canonical + transposeQ8_0MatmulDispatchTest.kt:68vsPackedMatmulDispatchTest.kt:130[in,out](un-reversed) whiletransposePackedBlocksreads the grid assuming[out,in]— transposing a verbatim-loaded GGUF tensor computes the wrong permutation, orrequire-fails whenout % blockSize != 0StreamingGgufParametersLoader.kt:96,StreamingGGUFReader.kt:375vsDefaultCpuOps.kt:466,800toFloatArray()/get()/matmulGenericare hard-wired canonical — any post-transpose (kernel-native) tensor silently dequantizes to a block-permuted matrixPackedBlockStorage.kt:52-61and every*TensorData.toFloatArrayAlso: the engine's only documented canonical→kernel-native MemSeg relayout lives in the downstream repo (
JvmQuantizedVectorKernels.kt:288literally namesGemmaMemSegConverteras responsible for honoring the kernel's layout) — layout knowledge owned by the wrong repo.Two more latent hazards:
transpose(transpose(W)) ≠ Won 0.40.1 for non-square block grids — nothing detects double-transposition.Linear.onForward(Linear.kt:76-77) doesweight.t()per forward; on 0.40.1 that's now an O(bytes) copy of the whole packed weight on every forward call, uncached.(SKaiNET-transformers has its own parallel census — three different conventions applied to the same GGUF K-quant tensor depending on which converter loads it, an inlined Apertus relayout that duplicated and diverged from the shared packer, a mixed-convention legacy weights object disambiguated by an unsafe shape heuristic. Tracked/fixed on that side via #311; not re-listed here since it's downstream-repo scope.)
The deeper semantic problem
ops.transposeon a block-quantized tensor is not a representable operation. Blocks quantize 32/256-element runs along the input dimension; a true transpose needs runs along the other axis, i.e. requantization. What the engine calls "packed transpose" is really a layout conversion into kernel feed order wearing transpose's name and a swapped shape label that lies about the data:shape=[in,out]+ bytes have no self-consistent packed interpretation →toFloatArray()on it is garbage (contradiction Flatten #7 above);mul_matcomputesx·Wᵀnatively against[out,in]weights — quantized data is never "transposed."Proposed direction (not a full spec — for discussion)
Principle: one convention, one owner, explicit in the type, loud on violation.
NarrowFloatInputMajorTensorData): a packed weight converted to kernel feed order becomes a distinct type (e.g.KernelPackedWeightData, logical shape stays[out,in]), not the same class with secretly different bytes. PlainQ*BlockTensorDatathen has exactly one meaning: canonical row-major. Kernels/dispatch accept only the kernel-packed type;toFloatArray()/get()on it either de-permute correctly or refuse.ops.transposefrom the packed hot path. Give the engine a weight-transposing matmul as the primitive (matmulWT(x, W[out,in]), the ggml/BLAS-op(B)shape).ops.transposeon packed data becomes a loud error pointing at the primitive (or an explicit dequant). This deletes both the semantic lie and the per-forward O(bytes) copy fix: physically reorder packed-quant blocks in ops.transpose (all-zero matmul, not just Q5_0/Q5_1) #969 introduced.prepackForMatmul(weight): KernelPackedWeightData, so every downstream converter (transformers, and any future consumer) calls one engine-owned function instead of maintaining private copies that can silently diverge (as Apertus' did).[out,in](HF convention), canonical bytes; the GGUF loader reverses ne dims (fixes contradiction MaxPooling2D #6).Q5_0/Q5_1TensorData, kernel SPI "Matches packedData", CHANGELOG:1038).packedDatabyte semantics are public API (breaking change = minor/major, never a patch); one normative doc (docs/packed-weight-layout.md) that every kdoc links to instead of restating.Why this needs to be a real issue, not just closed alongside #311
#311 unblocks the immediate downstream pin bump by making the transformers-side classic path stop relayouting (i.e. it adapts to the engine's 0.40.1 contract). That's a valid tactical fix, but it does nothing about contradictions #2–#7 above, which live entirely in this repo and can bite the next caller (engine-native user, a different downstream converter, a future format) regardless of what transformers does. The guessing-game nature of "which byte order does this instance actually hold" is the actual bug; #968/#969 and this transformers regression are two symptoms of it six weeks apart.
Not recommended: another byte-order guess/patch in
transpose(). Any guess re-creates the same trap for whichever producer population it doesn't anticipate — see this exact pattern already happening twice.