Repository navigation
TensorData and TensorStorage are parallel layers — unify the storage model (ownership, views, dtype/encoding, placement): SKEEP-003 discussion anchor #932
Description
Activity
The SKEEP-003 draft this issue promised is in the docs tree:
docs/modules/skeep/pages/003-unified-tensor-storage.adoc(landed via #933, registered in the SKEEP index/nav; Status: Draft). The end-state question — storage-first vs data-first — is deliberately left open here for maintainer discussion; the two shared prerequisites (two-wayLogicalDTypebridge,StorageSpecdecision) and the packed-encoding preservation constraint hold under either answer.Two follow-ups since the draft:
- docs(skeep): SKEEP-003 amendment — deprecate-don't-delete migration rule + implementation slices (#782, #921) #963 amends the SKEEP with what executing its first slice taught: an explicit migration rule (0.39.0 public API preserved — deprecate, don't delete, with the
GgufParametersLoaderdeprecation as the house pattern), a new Implementation Slices section mapping GGUF DEQUANTIZE_TO_FP32 over-allocates: 1.1B Q4_K_M needs >12 GB heap transiently (~4.4 GB legit) #782 (streaming dequant — the staging stage of the IO-pipeline improvement) and Android tensor storage is ART-heap-bound: no off-heap / mmap path outside jvmMain caps practical model size (SKEEP candidate) #921/SKEEP-002 (Android off-heap/mmap — the mobile destination-placement stage) as slices of this proposal, and the TensorStorageFactory violates its own copy/borrow contracts: "zero-copy" borrowFloatArray copies, "borrowed" fromTensorData owns #927–Memory diagnostics: MemoryTracker.recordCopy discards the source label every caller passes; ActiveMemoryTracker is a mutable global #931 mechanical fixes recorded as shipped in 0.39.0. - fix(io-gguf): stream GGUF dequantization — no boxed payloads, no defensive copies (#782) #965 is that first slice made concrete: the GGUF load path stops copying what it already owns (measured 41x boxed materialization in the legacy reader, 2.1–2.3x copy chain on the streaming path → 1.38x allocated / ~1.05x peak live), with a bit-exact packed-vs-dequant parity gate across all seven supported formats. It also demonstrates the "copy semantics are ambient" cost now written into the SKEEP's motivation: the factory copies defensively because ownership transfer isn't expressible in the creation API.
Discussion of the end-state (or a "B now, A when a device backend is scheduled" sequencing) is the open item on this anchor.
- docs(skeep): SKEEP-003 amendment — deprecate-don't-delete migration rule + implementation slices (#782, #921) #963 amends the SKEEP with what executing its first slice taught: an explicit migration rule (0.39.0 public API preserved — deprecate, don't delete, with the
- added sub-issues
on Aug 22, 2026 Decision (2026-08-22): SKEEP-003 is accepted — end-state A (storage-first), delivered with end-state B's incremental mechanics: new types (
Storageowns bytes ·TensorViewinterprets them ·Tensoris the DSL handle over a view or a graph node · kernels take views) introduced beside the existing ones in packagesk.ainet.lang.memory, everyTensorDataimplementation becomes a façade over the view type, dispatch migrates kernel by kernel on declaredFormat(dtype, encoding)keys, façades are deleted at the next major. All thirteen open design decisions are recorded in the SKEEP (docs PR #1043) and in the committed design recorddocs/design/memory/memory-architecture-proposal.md+memory-architecture-milestones-prd.md.Roadmap (this issue is the umbrella):
- Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001 — M0 · Know before you load (P0 + P1 partial): dtype bridge →
DTypewitness →Format/encoding on every tensor →AllocationSpec→TensorId→ GGUFNameMap→MemoryPlanfrom the GGUF header →skainet-planCLI. No runtime behaviour change. - Memory architecture M1 — Flat decode (Storage/Scope/TensorView, matmul via KernelRegistry, TraceSink) #1002 — M1 · Flat decode (P1 complete, P2, P3 matmul): Phase-2 access-path spike →
Storage/Scope/TensorView→TensorDatafaçades →TraceSink+ exporters → matmul throughKernelKey/registry (closes the Quantized matmul dispatch skips rank-1 (single-token decode) activations, falls through to broken matmulGeneric #993/Quantized matmul dispatch silently falls back to a broken generic kernel for rank>2 attention projections (ClassCastException: Byte cannot be cast to Float) #991 class) → plan-vs-actual →skainet-decodesample (home decided in [S1.11] M1 sample: decide the home ofskainet-decode(SKaiNET core vs SKaiNET-transformers), then build the CLI + Android activity #1032). - Memory architecture M2 — 1.58-bit on a 2 GB board (views, IO pipeline, 2 GB planner, ternary kernels) #1003 — M2 · 1.58-bit on a 2 GB board (P4, P5, P6): one view mechanism, sliding-window KV
(head, tail), IO pipeline + Android mmap (Android tensor storage is ART-heap-bound: no off-heap / mmap path outside jvmMain caps practical model size (SKEEP candidate) #921/createRandomAccessSourcereturns null on Android: full-file heap load OOMs on a 138 MiB GGUF (working ~40-line fix included) #922), 2 GB planner profile, ternary encodings + requant adapter + NEON kernel pack. - P7 (compiled parity) and P8 (façade removal) follow M2 and will be filed then.
Working rules for every slice: one
sub-issue= onefeature/<issue>-<slug>branch = one PR intodevelop; additive or behind a façade / opt-in sodevelopstays usable after each merge; deprecate-don't-delete; BCV dumps updated; full local test gate (scripts/pr-gate.sh: jvmTest, apiCheck, JS/Wasm, linuxX64Test, assemble, Java API tests) before each PR. First code slice: #1006 (two-wayLogicalDType↔DTypebridge).- Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001 — M0 · Know before you load (P0 + P1 partial): dtype bridge →
Closing: SKEEP-003 is implemented, and the contract it depended on is settled
This issue asked why
TensorDataandTensorStoragewere parallel layers, and what a unified storage model would look like. SKEEP-003 answered it, three milestones delivered it, and the byte-order contract that the answer kept running into is now written down.The model.
Storageowns bytes ·TensorViewinterprets them (Shape + Format + Layout + Storage) ·Tensorstays the DSL handle · kernels receive views.Format = (DType, TensorEncoding), so a Q4_K weight isFormat(FP32, Q4_K)and the logical dtype is never erased by the packing.Scopeowns lifetime —Model,Forward(a pre-sized slab, reset per step),Ambient.materialize()is the only copy point, and every conversion the dispatcher inserts is a visibleAdapterInsertedevent rather than a hidden allocation.The milestones.
- M0 — know before you load (Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001):
MemoryPlanfrom a GGUF header alone, structuredTensorId, GGUF↔id name maps, and theskainet-planCLI. - M1 — flat decode (Memory architecture M1 — Flat decode (Storage/Scope/TensorView, matmul via KernelRegistry, TraceSink) #1002):
Storage/Scope/TensorViewon every target,TensorDataas a façade over views, matmul throughKernelKeydispatch,TraceSinkwith Perfetto/JFR/android.os.Traceexporters, plan-vs-actual. - M2 — 1.58-bit on a 2 GB board (Memory architecture M2 — 1.58-bit on a 2 GB board (views, IO pipeline, 2 GB planner, ternary kernels) #1003): ternary encodings with a ggml-faithful codec, an int8 requant adapter and
bitnet_gemv(reference + NEON), one view mechanism, one IO pipeline (quantPolicy × staging), a two-pool device fit check and planner profiles, and a sliding-window KV ring.
The contract (#973): block order is now part of the layout, kernels declare the order they read, the relayout is engine-owned and idempotent, and
ops.transposerefuses a packed weight instead of performing a per-forward copy of a semantic lie.What the work found on the way
Several defects that were live and silent, none of which this issue set out to find:
- both GGML ternary decoders walked elements in the wrong order, and GGUF typed those tensors
Opaque; - a blocked layout assumed its block axis was the last one, so transposing a packed view decoded the wrong elements — zero-copy and wrong;
- the planner assumed bf16 KV against an FP32 ring, understating a dense cache by 2×;
- adopted/mapped weights were invisible to plan-vs-actual;
- three copies of one source factory had drifted apart across format modules;
- and two CI gaps: Android host tests and plain-JVM module tests had never run in CI at all.
What stays open, deliberately
- The
skainet-decodesample in SKaiNET-transformers, which owns a model: M1-A2, the tok/s half of M1-A5 and M2-A1's measured run belong there. - M2-A5 on an Android device — the mechanism and the fit check exist; the measurement needs hardware with ART.
- Downstream migration to
matmulWeightTransposed/PackedWeights, and makingWeightOrientation.OUT_INthe default once that has happened. - P7/P8 of the proposal (device placement, graph-level planning) were always post-M2.
The full record is in
docs/design/memory/— the proposal, the milestone PRD, the M2 acceptance tables, and the packed-weight layout contract.Closing. SKEEP-003 is
Acceptedand implemented; SKEEP-002 keeps its status note about what Android still waits on.- M0 — know before you load (Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001):
Two corrections to the closing summary above, so it does not mislead anyone reading it later. Both are things that changed after it was written; neither reopens the issue.
1.
ops.transposeno longer refuses a packed weight.The summary says the #973 contract ends with "
ops.transposerefuses a packed weight instead of performing a per-forward copy of a semantic lie". The refusal was right about the operation and wrong about the caller: it meantx.matmul(w.t())compiled for a dense weight and threw for a packed one, so the code an author writes depended on which file the user loaded and which device it ran on.#1108 replaced it.
transposenow returnsTransposedWeightTensorData— the shape is the transpose, the payload is still the weight's, and the type is deliberately notPackedBlockStorage, so no kernel can reach the bytes through it.matmulrecognises the marker and asksmatmulWeightTransposedfor the product.w.t().t()iswagain.What the summary says about the contract still holds: transposing block-quantized bytes is still not a representable operation, and nothing performs a per-call copy. Only the way that is expressed changed — from a refusal to a value that knows it is unmaterialized.
2. The design record is not in
docs/design/memory/.That directory was deleted in #1106. The record is Antora-only now, under
docs/modules/—explanation/memory-model.adoc,explanation/packed-weight-layout.adoc,explanation/eager-execution.adoc,how-to/plan-model-memory.adoc, and arc42 inreference/architecture.adoc.
The "stays open, deliberately" list now has issues. It was living only in the prose of a closed issue, which is not tracking:
- The skainet-decode sample, and the measured runs that need a model #1129 — the
skainet-decodesample and the measured runs that need a model - M2-A5: a Q4_K_M model loading under a real heap cap on an Android device #1130 — M2-A5 on an Android device with ART
- P7/P8: device placement and graph-level memory planning #1131 — P7/P8: device placement and graph-level planning
The fourth item, downstream migration to
matmulWeightTransposed/PackedWeightsand thenWeightOrientation.OUT_INby default, has been overtaken: #1109 replaced the three loader policy flags with one resolvedWeightForm, andWeightOrientationis deprecated in favour of itsshapeaxis. The migration is now "adoptWeightForm", which is #1129's concern where it matters (the sample) and otherwise a downstream repository's.- The skainet-decode sample, and the measured runs that need a model #1129 — the
P7/P8 arc (#1131) resolved and closed: placement is resolver-owned (
AllocationResolver, #1142–#1144), eager buffer lifetime is the Scope split wired at creation (#1145), and graph-level memory planning stays downstream in IREE — core decides and carries (#1134 rationale, SKEEP-003a). SKEEP-003's 'scheduled for deletion' items (StorageSpec, the oldMemoryPlanner,@Place/@Weights) are discharged by #1142, and cross-cutting improvement 5 (placement consulted at creation) by #1143/#1145 in resolver form. First façade-removal slice (#1159, legacy loader axes) in flight. This umbrella stays open as the SKEEP-003 anchor for what remains: #1146 (op outputs through the scope), #1147/#1148 (IREE-milestone carriage + parity harness).
SKaiNET carries two storage abstractions:
TensorData(skainet-lang-core/.../tensor/data/) — the live one. EveryTensorholds one; backends dispatch by downcasting to concrete classes/markers (is Q4_KTensorData -> .packedData); all common-code storage is heap arrays.TensorStorage(.../tensor/storage/, 20 files) — a designed descriptor layer with the right concepts (BufferHandle.{Owned,Borrowed,Aliased,FileBacked,DeviceResident},Placement/MemoryDomain,TensorEncoding,MemoryPlanner,StorageSpec) whose own KDoc says "new loaders, planners, and backends should targetTensorStoragedirectly" — but which noTensorever holds.StorageSpechas zero consumers; the planner is never consulted at any allocation;Aliasedis never produced;FileBackedandDeviceResidentthrow in every consumer;LogicalDType.fromDTypehas no inverse, so the layer structurally cannot backTensor<T>today.The result is a set of related, recurring costs:
Placement.Residencymodels but nothing uses)SlicedTensorViewindex remap vsBufferHandle.Aliasedbyte range)DTypeKClass generic,LogicalDTypeenum,TensorEncoding), with packed tensors erasing their logical dtype entirely (Q4_KTensorData : TensorData<DType, Byte>— a logically-FP32 weight is not typed as such; ops find it by class check)createRandomAccessSourcereturns null on Android: full-file heap load OOMs on a 138 MiB GGUF (working ~40-line fix included) #922, SKEEP-002) keeps landing at the edges because the middle isn't wiredAn SKEEP-003 draft (PR to follow) lays out the analysis and two candidate end-states — (a)
TensorStoragebecomes the single byte-owner andTensorDataa typed view protocol over it; (b)TensorDatastays primary and absorbsBufferHandle/Placement, retiring the parallel descriptor — deliberately without a recommendation: the trade-off (dispatch rewrite + dtype coherence vs minimal churn + status quo dispatch) is a maintainer decision. Both directions share two prerequisites (two-wayLogicalDType ↔ DTypebridge; decideStorageSpec's fate) and one hard constraint: the packed-encoding system (7 GGML block formats, ternary, TurboQuant, kernel dispatch, StableHLOskainet.tensor_encodingsexport) must survive bit-identically.Mechanical bugs found during the same audit are filed separately (TensorStorageFactory contract violations, GGUF encoding mapping dropping five formats, transfer-API gaps, rank-broken
copyToFloatArraydefault, memory-diagnostics paper-cuts) — they are fixable under either end-state.