Skip to content

Memory architecture M0 — Know before you load (plan + identity, no runtime change) #1001

Description

@michalharakal
Parent tracking issue for **M0 — Know before you load** of the SKEEP-003 memory architecture (umbrella: #932). Design docs: `docs/design/memory/memory-architecture-proposal.md` (proposal rev. 2, §-numbers below refer to it) and `docs/design/memory/memory-architecture-milestones-prd.md` (PRD, requirement IDs `Mx-Fy`/`Mx-Ay`) — both land with the SKEEP-003 docs PR; until merged they are attached to #932.

## Goal
Answer *will this model fit on this device at this context length?* from the GGUF header alone, give every weight a stable `TensorId`, and make every tensor report a coherent `Format(dtype, encoding)`. No runtime behaviour changes (existing benchmarks unchanged).

## Sample app / acceptance vehicle
`skainet-plan` CLI (`skainet-apps/skainet-plan`, JVM; planner in `commonMain` so Android/K-Native can call it): `skainet plan model.gguf --ctx 2048 --budget 1.3G` prints the resident breakdown, fit verdict and suggestions; `--list layers[3].*` prints `TensorId · Format · Shape · bytes · ← gguf name`.

## Consumes (proposal §9 phases)
P0 (two-way dtype bridge → `DType` merge, `StorageSpec` → allocation spec) · P1 partial (`Format` reported everywhere, `TensorId` + GGUF `NameMap`, `toString()` renderer).

## Acceptance criteria (PRD)
- M0-A1 plan for Llama-3.2-1B Q4_K_M @ ctx 2048 within ±10 % of measured resident memory (recorded now, verified at M1-A8)

The checklist is maintained by GitHub's sub-issue tracking on this issue.

Working rules (every slice)

  • One slice = one branch feature/<this-issue#>-<slug> from develop = one PR (Closes #<this issue>), merged independently.
  • develop stays usable after merge: additive API or behind façade / opt-in; no default behaviour switch without golden-parity + benchmark evidence in the PR; deprecate with ReplaceWith, never delete before a major (zero-consumer dead code may be removed with evidence in the PR). No new abstract members on TensorData/KernelProvider (defaults only); no data-class primary-constructor changes (secondary ctors only); BCV: apiDump both api/jvm and api/android, apiCheck clean.
  • Test gate before the PR (JDK 25, scripts/pr-gate.sh once it exists): ./gradlew jvmTest linuxX64Test assemble, ./gradlew verifyNpmPins jsTest wasmJsTest wasmWasiTest, ./gradlew apiCheck, ./gradlew :skainet-test:skainet-test-java:test; storage/backend slices add :skainet-lang:skainet-lang-core:jvmBenchmark and :skainet-backends:benchmarks:jvm-cpu-jmh:jmh before/after when a hot path changes.
  • Commits: conventional commit subject, Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> trailer. Do not edit CHANGELOG.md in slice PRs (parallel branches conflict on it; release notes are written at release time from the merged PRs).

Activity

  1. michalharakal commented on Aug 23, 2026

    @michalharakal
    ContributorAuthor

    M0 — Know before you load: wrap-up (2026-08-23)

    All M0 slices are merged into develop (last: #1059 @ 3c7c7d1). What landed, by sub-issue:

    Sub-issue Slice PR
    #1004 SKEEP-003 → Accepted, design record + PRD in docs/design/memory/ #1043
    #1005 golden parity gate (bit-identical decode / scalar kernels / TurboQuant on JVM+linuxX64, dispatch parity on all targets), golden-parity CI leg + apiCheck in CI, scripts/pr-gate.sh, JMH baseline #1049, #1047, #1052
    #1006 two-way LogicalDType ↔ DType bridge #1045
    #1007 DType.witness / fromWitness / entries; dtype-first TensorStorage, factory, readers #1050
    #1008 Format(dtype, encoding), TensorData.encoding, Tensor.format / TensorStorage.format (sk.ainet.lang.memory, @ExperimentalMemoryApi) #1051
    #1009 AllocationSpec + ScopeKind; StorageSpec deprecated #1053
    #1010 TensorId (canonical + legacy forms), assignTensorIds, describe() renderer #1054
    #1011 NameMap GGUF/HF ↔ TensorId for Llama-3 / Qwen2.5 / Gemma-3, unmapped reported #1055
    #1012 MemoryPlan from the GGUF header (weights · KV bf16/TurboQuant · forward slab · headroom · budget · suggestions) #1056
    #1013 skainet-plan CLI #1057
    #1014 LogicalDType & friends @Deprecated #1059

    Acceptance (PRD §4.5)

    • M0-A2 ✔ — all tensors of Llama-3.2-1B (16 layers, rope_freqs), Qwen2.5-0.5B (24 layers, q/k/v biases) and Gemma-3-1B (26 layers, q/k norms, four norms) map to TensorIds with zero unmapped (NameMapTest on the full synthetic tensor lists; GgufNameMapFixtureTest repeats it on real files when they are present in -Dskainet.test.fixturesDir).
    • M0-A3 ✔ — Q4_K (and every other packed) tensor reports Format(FP32, <encoding>); no reader reports a packed tensor as Byte (FormatTest, StorageIntegrationTest).
    • M0-A4 ✔ — plan arithmetic unit tests (MemoryPlanTest); header-only plan from a synthetic GGUF and the Qwen fixture (GgufMemoryPlanTest). The three reference plans are recorded as tests rather than golden JSON files (no reference GGUFs in CI).
    • M0-A5 — benchmarks unchanged: JMH suite re-run on develop @ 8b5ce54 against the pre-M0 baseline (docs/design/memory/baseline-2026-08-22.md), same machine, idle: holds. Full suite on develop @ 8b5ce54 vs the baseline: Panama/vector matmul, Q4_K panama matmul, elementwise add and reductions are all within ±2 % (inside their error bars). The scalar FP32 reference matmul (KernelMatmulBench provider=scalar) swung ±10 % in both directions across four runs (e.g. 512: 207 → 212 → 293 → 223 ms; 256: 24.5 → 22.1 ms) — a back-to-back A/B of the pre-M0 commit 5611b91 vs develop in the same thermal state gave 256 −9.8 %, 512 +8.3 %, 1024 +4.4 % ± 17 %, i.e. laptop frequency/thermal variance on a 1.7 s CPU-bound loop, not a regression: git diff 5611b918 develop -- skainet-backends/*/src is empty for every kernel and dispatch file (M0 touched no hot path; it added default members Tensor.id / TensorData.encoding that no kernel calls). Tables: jmh-compare (full) and the A/B are attached below.
    • M0-A1 — recorded now, verified at M1-A8: the estimate for Llama-3.2-1B Q4_K_M @ ctx 2048 from the architecture geometry (16 layers · 32/8 heads · head 64 · emb 2048 · ffn 8192 · vocab 128 256) is weights ≈ 0.7 GB (Q4_K on the projection shapes; the real file mixes Q4_K/Q6_K and adds norms) · KV bf16 64 MiB (17 MiB TurboQuant) · forward slab (prefill 256) ≈ 47 MB · headroom 64 MB. To record the real number, run on a machine with the file: ./gradlew :skainet-apps:skainet-plan:run --args="Llama-3.2-1B-Instruct-Q4_K_M.gguf --ctx 2048 --budget 1.3G" and paste the table here; M1's plan-vs-actual slice ([S1.9] M1: plan-vs-actual — allocation-event totals vs MemoryPlan, printed; CI assertion (> 10 % fails) #1030) verifies it against measured resident memory.

    Carried forward

    JMH full suite: pre-M0 baseline (5611b91) vs develop @ 8b5ce54
    Benchmark Params Mode Pre-M0 Post-M0 Δ Units
    add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=false thrpt 1.217 ± 0.011 1.226 ± 0.013 -0.7% ops/ms
    add_1M_fp32 ElementwiseAdd1MBench vectorEnabled=true thrpt 1.705 ± 0.030 1.725 ± 0.023 -1.2% ops/ms
    matmul_fp32_square KernelMatmulBench provider=panama, size=1024 avgt 80.878 ± 0.374 81.625 ± 1.666 +0.9% ms/op
    matmul_fp32_square KernelMatmulBench provider=panama, size=256 avgt 1.186 ± 0.008 1.212 ± 0.021 +2.2% ms/op
    matmul_fp32_square KernelMatmulBench provider=panama, size=512 avgt 9.157 ± 0.227 9.559 ± 0.169 +4.4% ⚠ ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=1024 avgt 1685.182 ± 88.376 1838.158 ± 12.567 +9.1% ⚠ ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=256 avgt 22.858 ± 1.734 23.511 ± 0.231 +2.9% ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=512 avgt 207.298 ± 0.961 212.253 ± 46.296 +2.4% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=1024, vectorEnabled=false avgt 1757.350 ± 126.247 1554.774 ± 13.724 -11.5% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=1024, vectorEnabled=true avgt 82.912 ± 0.778 82.318 ± 0.621 -0.7% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=256, vectorEnabled=false avgt 20.479 ± 0.784 21.013 ± 2.828 +2.6% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=256, vectorEnabled=true avgt 1.225 ± 0.043 1.175 ± 0.004 -4.1% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=512, vectorEnabled=false avgt 199.811 ± 46.629 194.867 ± 1.717 -2.5% ms/op
    matmul_fp32_square MatmulBench blasEnabled=false, size=512, vectorEnabled=true avgt 9.371 ± 0.162 9.514 ± 0.118 +1.5% ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=1024, vectorEnabled=false avgt 1709.814 ± 5.381 1710.889 ± 56.748 +0.1% ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=1024, vectorEnabled=true avgt 82.036 ± 2.492 81.984 ± 0.600 -0.1% ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=256, vectorEnabled=false avgt 21.129 ± 1.062 20.732 ± 0.864 -1.9% ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=256, vectorEnabled=true avgt 1.226 ± 0.010 1.220 ± 0.015 -0.5% ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=512, vectorEnabled=false avgt 209.844 ± 0.821 275.811 ± 0.373 +31.4% ⚠ ms/op
    matmul_fp32_square MatmulBench blasEnabled=true, size=512, vectorEnabled=true avgt 9.216 ± 0.094 9.620 ± 0.151 +4.4% ⚠ ms/op
    matmul_q4k_panama QuantizedMatmulBench shape=1024-1024 avgt 0.112 ± 0.017 0.114 ± 0.016 +1.8% ms/op
    matmul_q4k_panama QuantizedMatmulBench shape=4096-1024 avgt 0.336 ± 0.003 0.335 ± 0.003 -0.3% ms/op
    matmul_q4k_panama QuantizedMatmulBench shape=4096-4096 avgt 1.223 ± 0.017 1.224 ± 0.028 +0.1% ms/op
    mean_1M_fp32 Reductions1MBench vectorEnabled=false thrpt 1.026 ± 0.004 1.022 ± 0.005 +0.4% ops/ms
    mean_1M_fp32 Reductions1MBench vectorEnabled=true thrpt 1.026 ± 0.011 1.021 ± 0.015 +0.5% ops/ms
    sum_1M_fp32 Reductions1MBench vectorEnabled=false thrpt 1.027 ± 0.006 1.025 ± 0.003 +0.2% ops/ms
    sum_1M_fp32 Reductions1MBench vectorEnabled=true thrpt 1.026 ± 0.003 1.026 ± 0.007 +0.0% ops/ms

    worst slowdown: +31.4% (positive = slower after M0)

    Back-to-back A/B, KernelMatmulBench, 5611b91 vs develop (same thermal state)
    Benchmark Params Mode Pre-M0 Post-M0 Δ Units
    matmul_fp32_square KernelMatmulBench provider=panama, size=1024 avgt 81.754 ± 0.613 80.941 ± 0.114 -1.0% ms/op
    matmul_fp32_square KernelMatmulBench provider=panama, size=256 avgt 1.220 ± 0.023 1.212 ± 0.011 -0.7% ms/op
    matmul_fp32_square KernelMatmulBench provider=panama, size=512 avgt 9.615 ± 0.160 9.650 ± 0.064 +0.4% ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=1024 avgt 1685.937 ± 8.789 1760.482 ± 289.181 +4.4% ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=256 avgt 24.497 ± 0.190 22.102 ± 0.098 -9.8% ms/op
    matmul_fp32_square KernelMatmulBench provider=scalar, size=512 avgt 206.138 ± 1.210 223.311 ± 1.281 +8.3% ⚠ ms/op

    worst slowdown: +8.3% (positive = slower after M0)

  2. michalharakal commented on Aug 23, 2026

    @michalharakal
    ContributorAuthor

    M0 complete — all twelve sub-issues closed, wrap-up above. M0-A1's real-file number is recorded/verified in M1 (#1030). Next: M1 #1002, starting with the P2 spike #1016.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    architecturaltensorsTensor operations and data structurestrackingParent/tracking issue with sub-issues

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions