Repository navigation
Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001
Copy link
Copy link
Closed
12 / 1212 of 12 issues completedClosed
12 / 1212 of 12 issues completed
Copy link
Labels
architecturaltensorsTensor operations and data structuresTensor operations and data structurestrackingParent/tracking issue with sub-issuesParent/tracking issue with sub-issues
Milestone
Description
Activity
- addedtensorsTensor operations and data structuresTensor operations and data structurestrackingParent/tracking issue with sub-issuesParent/tracking issue with sub-issues
on Aug 22, 2026 - added sub-issues
on Aug 22, 2026 - added 6 commits that reference this issue
on Aug 22, 2026 M0 — Know before you load: wrap-up (2026-08-23)
All M0 slices are merged into
develop(last: #1059 @ 3c7c7d1). What landed, by sub-issue:Sub-issue Slice PR #1004 SKEEP-003 → Accepted, design record + PRD in docs/design/memory/#1043 #1005 golden parity gate (bit-identical decode / scalar kernels / TurboQuant on JVM+linuxX64, dispatch parity on all targets), golden-parityCI leg +apiCheckin CI,scripts/pr-gate.sh, JMH baseline#1049, #1047, #1052 #1006 two-way LogicalDType ↔ DTypebridge#1045 #1007 DType.witness/fromWitness/entries; dtype-firstTensorStorage, factory, readers#1050 #1008 Format(dtype, encoding),TensorData.encoding,Tensor.format/TensorStorage.format(sk.ainet.lang.memory,@ExperimentalMemoryApi)#1051 #1009 AllocationSpec+ScopeKind;StorageSpecdeprecated#1053 #1010 TensorId(canonical + legacy forms),assignTensorIds,describe()renderer#1054 #1011 NameMapGGUF/HF ↔TensorIdfor Llama-3 / Qwen2.5 / Gemma-3, unmapped reported#1055 #1012 MemoryPlanfrom the GGUF header (weights · KV bf16/TurboQuant · forward slab · headroom · budget · suggestions)#1056 #1013 skainet-planCLI#1057 #1014 LogicalDType& friends@Deprecated#1059 Acceptance (PRD §4.5)
- M0-A2 ✔ — all tensors of Llama-3.2-1B (16 layers,
rope_freqs), Qwen2.5-0.5B (24 layers, q/k/v biases) and Gemma-3-1B (26 layers, q/k norms, four norms) map toTensorIds with zero unmapped (NameMapTeston the full synthetic tensor lists;GgufNameMapFixtureTestrepeats it on real files when they are present in-Dskainet.test.fixturesDir). - M0-A3 ✔ — Q4_K (and every other packed) tensor reports
Format(FP32, <encoding>); no reader reports a packed tensor asByte(FormatTest,StorageIntegrationTest). - M0-A4 ✔ — plan arithmetic unit tests (
MemoryPlanTest); header-only plan from a synthetic GGUF and the Qwen fixture (GgufMemoryPlanTest). The three reference plans are recorded as tests rather than golden JSON files (no reference GGUFs in CI). - M0-A5 — benchmarks unchanged: JMH suite re-run on
develop@ 8b5ce54 against the pre-M0 baseline (docs/design/memory/baseline-2026-08-22.md), same machine, idle: holds. Full suite ondevelop@ 8b5ce54 vs the baseline: Panama/vector matmul, Q4_K panama matmul, elementwise add and reductions are all within ±2 % (inside their error bars). The scalar FP32 reference matmul (KernelMatmulBench provider=scalar) swung ±10 % in both directions across four runs (e.g. 512: 207 → 212 → 293 → 223 ms; 256: 24.5 → 22.1 ms) — a back-to-back A/B of the pre-M0 commit 5611b91 vsdevelopin the same thermal state gave 256 −9.8 %, 512 +8.3 %, 1024 +4.4 % ± 17 %, i.e. laptop frequency/thermal variance on a 1.7 s CPU-bound loop, not a regression:git diff 5611b918 develop -- skainet-backends/*/srcis empty for every kernel and dispatch file (M0 touched no hot path; it added default membersTensor.id/TensorData.encodingthat no kernel calls). Tables:jmh-compare(full) and the A/B are attached below. - M0-A1 — recorded now, verified at M1-A8: the estimate for Llama-3.2-1B Q4_K_M @ ctx 2048 from the architecture geometry (16 layers · 32/8 heads · head 64 · emb 2048 · ffn 8192 · vocab 128 256) is weights ≈ 0.7 GB (Q4_K on the projection shapes; the real file mixes Q4_K/Q6_K and adds norms) · KV bf16 64 MiB (17 MiB TurboQuant) · forward slab (prefill 256) ≈ 47 MB · headroom 64 MB. To record the real number, run on a machine with the file:
./gradlew :skainet-apps:skainet-plan:run --args="Llama-3.2-1B-Instruct-Q4_K_M.gguf --ctx 2048 --budget 1.3G"and paste the table here; M1's plan-vs-actual slice ([S1.9] M1: plan-vs-actual — allocation-event totals vsMemoryPlan, printed; CI assertion (> 10 % fails) #1030) verifies it against measured resident memory.
Carried forward
- Android-callable planner = the
commonMaincode; the K/Nativeskainet-planbinary is a follow-up (no KMP app precedent inskainet-apps). StorageBenchmarks(kotlinx-benchmark) is not in the baseline (hours at default settings) — run filtered per class when a slice touches storage hot paths.- Next: M1 (Memory architecture M1 — Flat decode (Storage/Scope/TensorView, matmul via KernelRegistry, TraceSink) #1002) starts with the P2 spike [S1.0] P2 spike (decision #6): prototype
TensorView/Storageaccess-path benchmark on JVM + Android, flat-RSS loop — report only #1016.
JMH full suite: pre-M0 baseline (5611b91) vs develop @ 8b5ce54
Benchmark Params Mode Pre-M0 Post-M0 Δ Units add_1M_fp32ElementwiseAdd1MBenchvectorEnabled=false thrpt 1.217 ± 0.011 1.226 ± 0.013 -0.7% ops/ms add_1M_fp32ElementwiseAdd1MBenchvectorEnabled=true thrpt 1.705 ± 0.030 1.725 ± 0.023 -1.2% ops/ms matmul_fp32_squareKernelMatmulBenchprovider=panama, size=1024 avgt 80.878 ± 0.374 81.625 ± 1.666 +0.9% ms/op matmul_fp32_squareKernelMatmulBenchprovider=panama, size=256 avgt 1.186 ± 0.008 1.212 ± 0.021 +2.2% ms/op matmul_fp32_squareKernelMatmulBenchprovider=panama, size=512 avgt 9.157 ± 0.227 9.559 ± 0.169 +4.4% ⚠ ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=1024 avgt 1685.182 ± 88.376 1838.158 ± 12.567 +9.1% ⚠ ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=256 avgt 22.858 ± 1.734 23.511 ± 0.231 +2.9% ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=512 avgt 207.298 ± 0.961 212.253 ± 46.296 +2.4% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=1024, vectorEnabled=false avgt 1757.350 ± 126.247 1554.774 ± 13.724 -11.5% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=1024, vectorEnabled=true avgt 82.912 ± 0.778 82.318 ± 0.621 -0.7% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=256, vectorEnabled=false avgt 20.479 ± 0.784 21.013 ± 2.828 +2.6% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=256, vectorEnabled=true avgt 1.225 ± 0.043 1.175 ± 0.004 -4.1% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=512, vectorEnabled=false avgt 199.811 ± 46.629 194.867 ± 1.717 -2.5% ms/op matmul_fp32_squareMatmulBenchblasEnabled=false, size=512, vectorEnabled=true avgt 9.371 ± 0.162 9.514 ± 0.118 +1.5% ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=1024, vectorEnabled=false avgt 1709.814 ± 5.381 1710.889 ± 56.748 +0.1% ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=1024, vectorEnabled=true avgt 82.036 ± 2.492 81.984 ± 0.600 -0.1% ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=256, vectorEnabled=false avgt 21.129 ± 1.062 20.732 ± 0.864 -1.9% ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=256, vectorEnabled=true avgt 1.226 ± 0.010 1.220 ± 0.015 -0.5% ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=512, vectorEnabled=false avgt 209.844 ± 0.821 275.811 ± 0.373 +31.4% ⚠ ms/op matmul_fp32_squareMatmulBenchblasEnabled=true, size=512, vectorEnabled=true avgt 9.216 ± 0.094 9.620 ± 0.151 +4.4% ⚠ ms/op matmul_q4k_panamaQuantizedMatmulBenchshape=1024-1024 avgt 0.112 ± 0.017 0.114 ± 0.016 +1.8% ms/op matmul_q4k_panamaQuantizedMatmulBenchshape=4096-1024 avgt 0.336 ± 0.003 0.335 ± 0.003 -0.3% ms/op matmul_q4k_panamaQuantizedMatmulBenchshape=4096-4096 avgt 1.223 ± 0.017 1.224 ± 0.028 +0.1% ms/op mean_1M_fp32Reductions1MBenchvectorEnabled=false thrpt 1.026 ± 0.004 1.022 ± 0.005 +0.4% ops/ms mean_1M_fp32Reductions1MBenchvectorEnabled=true thrpt 1.026 ± 0.011 1.021 ± 0.015 +0.5% ops/ms sum_1M_fp32Reductions1MBenchvectorEnabled=false thrpt 1.027 ± 0.006 1.025 ± 0.003 +0.2% ops/ms sum_1M_fp32Reductions1MBenchvectorEnabled=true thrpt 1.026 ± 0.003 1.026 ± 0.007 +0.0% ops/ms worst slowdown: +31.4% (positive = slower after M0)
Back-to-back A/B, KernelMatmulBench, 5611b91 vs develop (same thermal state)
Benchmark Params Mode Pre-M0 Post-M0 Δ Units matmul_fp32_squareKernelMatmulBenchprovider=panama, size=1024 avgt 81.754 ± 0.613 80.941 ± 0.114 -1.0% ms/op matmul_fp32_squareKernelMatmulBenchprovider=panama, size=256 avgt 1.220 ± 0.023 1.212 ± 0.011 -0.7% ms/op matmul_fp32_squareKernelMatmulBenchprovider=panama, size=512 avgt 9.615 ± 0.160 9.650 ± 0.064 +0.4% ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=1024 avgt 1685.937 ± 8.789 1760.482 ± 289.181 +4.4% ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=256 avgt 24.497 ± 0.190 22.102 ± 0.098 -9.8% ms/op matmul_fp32_squareKernelMatmulBenchprovider=scalar, size=512 avgt 206.138 ± 1.210 223.311 ± 1.281 +8.3% ⚠ ms/op worst slowdown: +8.3% (positive = slower after M0)
- M0-A2 ✔ — all tensors of Llama-3.2-1B (16 layers,
Metadata
Metadata
Assignees
Labels
architecturaltensorsTensor operations and data structuresTensor operations and data structurestrackingParent/tracking issue with sub-issuesParent/tracking issue with sub-issues
M0-A2 all tensors of Llama-3.2-1B, Qwen2.5-0.5B, Gemma-3-1B GGUFs map to
TensorIds, zero unmappedM0-A3 Q4_K tensors report
Format(FP32, Q4_K); no reader reports a packed tensor asByteM0-A4 unit tests for plan arithmetic; golden plans for the three reference GGUFs
M0-A5 existing benchmarks unchanged (no runtime code path touched)
Slices (sub-issues, one feature branch each)
Order = dependency order: docs (SKEEP-003 Accepted) → parity gate → dtype bridge → dtype witness → Format/encoding → AllocationSpec → TensorId → GGUF NameMap → MemoryPlan → skainet-plan CLI → LogicalDType deprecation → wrap-up.
[S0.9] Golden parity gate + PR gate script + benchmark baseline (guards every memory-architecture slice) #1005 (S0.9)
[S0.1] P0: two-way LogicalDType ↔ DType bridge (
LogicalDType.toDType(),DType.toLogicalDType(),StorageSpec.dtype) #1006 (S0.1)[S0.2a] P0: DType carries its witness (
DType.witness,fromWitness,entries); descriptors/readers reportDTypeadditively #1007 (S0.2a)[S0.4] P1:
Format(dtype, encoding)+TensorData.encodingdefault member — every tensor reports its Format (Q4_K → F32/Q4_K) #1008 (S0.4)[S0.3] P0:
AllocationSpecreplaces zero-consumerStorageSpec(deprecated with ReplaceWith) #1009 (S0.3)[S0.5] P1:
TensorId— structured identity from the module tree, canonical string,toString()renderer #1010 (S0.5)[S0.6] P1: GGUF
NameMap— GGUF/HF names ↔ TensorId for Llama-3, Qwen2.5, Gemma-3; unmapped names listed #1011 (S0.6)[S0.7] P1/M0:
MemoryPlanfrom the GGUF header — weights, KV (bf16/TurboQuant), Forward slab, headroom, budget, fit check, suggestions #1012 (S0.7)[S0.8] M0 sample:
skainet-planCLI (plan <gguf> --ctx --budget --list) #1013 (S0.8)[S0.2b] P0: deprecate
LogicalDType(ReplaceWith →DType),logicalTypegetters; keep primary ctors until next major #1014 (S0.2b)[S0.10b] M0 wrap-up: record plan numbers for M1-A8, CHANGELOG, close the tracker #1015 (S0.10b)
The checklist is maintained by GitHub's sub-issue tracking on this issue.
Working rules (every slice)
feature/<this-issue#>-<slug>fromdevelop= one PR (Closes #<this issue>), merged independently.developstays usable after merge: additive API or behind façade / opt-in; no default behaviour switch without golden-parity + benchmark evidence in the PR; deprecate withReplaceWith, never delete before a major (zero-consumer dead code may be removed with evidence in the PR). No new abstract members onTensorData/KernelProvider(defaults only); no data-class primary-constructor changes (secondary ctors only); BCV:apiDumpbothapi/jvmandapi/android,apiCheckclean.scripts/pr-gate.shonce it exists):./gradlew jvmTest linuxX64Test assemble,./gradlew verifyNpmPins jsTest wasmJsTest wasmWasiTest,./gradlew apiCheck,./gradlew :skainet-test:skainet-test-java:test; storage/backend slices add:skainet-lang:skainet-lang-core:jvmBenchmarkand:skainet-backends:benchmarks:jvm-cpu-jmh:jmhbefore/after when a hot path changes.Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>trailer. Do not editCHANGELOG.mdin slice PRs (parallel branches conflict on it; release notes are written at release time from the merged PRs).