You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix(models): dedicated root-cause debugging of the DeepSeek-V2-Lite greedy nondeterminism (follow-up to #817) #824
UPDATE (root cause pinned): A dedicated debugging pass has pinned this to the MLX 0.32.1 bump 6c53673 (#704, 2026-07-09) via a clean deterministic bisect. Key reframing of #817's conclusion: with MLX_USE_CUDA_GRAPHS=0 the output is deterministic garbage (not a memory-safety nondeterminism), compute-sanitizer initcheck+memcheck are both clean, and the fault is an upstream MLX CUDA kernel change breaking deepseek_v2's naive-MLA shape (qk 192 / v 128); deepseek_v3's absorbed-MLA is unaffected. Full analysis and the concrete next steps (sub-bisect the upstream MLX range, then patch-overlay / upstream-report / absorbed-MLA-rewrite) are in the comment below. No fix merged yet; the investigation plan below still applies for the remaining upstream-narrowing work.
Summary
This is the dedicated root-cause investigation for the DeepSeek-V2-Lite greedy nondeterminism established in #817. #817 determined that the "regressed since the benchmark" framing was wrong and that a mechanical bisect is not viable (there is no coherent-good anchor to bisect from), so a code fix was intentionally not forced. This issue owns the actual debugging: find and close the memory-safety defect that makes deepseek_v2 produce different greedy output across identical fresh-process runs, then restore deterministic, coherent decoding on both backends.
Nondeterministic greedy decode. At --temp 0, repeated fresh-process runs of the same binary on the same prompt produce different outputs, most of them degenerate (a repeated single char like ! or a repeated punctuation gram like }'_ / </:). Greedy decode should be bit-identical run to run; it is not. That is a memory-safety class defect (uninitialized read, out-of-bounds, or a race), not a stable numeric error.
Pre-existing. At the 2026-05-19 benchmark commit (331ac6e), coherence is already only ~50% (measured 8/12 and 5/12 coherent over fresh greedy runs with fused MoE off). The benchmark's ✅ means only "generates 100 tokens without crashing", not coherent output, which is why this went unnoticed.
Corruption is at the first generated token, so it originates in prefill or weight loading, not late in decode.
Checkpoint is intact (shard sha256 match HuggingFace; unmodified since January).
Specific to the deepseek_v2 geometry. Sibling models are stable and coherent on the same binary: qwen1.5-moe (switch_layers MoE, standard attention), youtu-vl (DeepSeek-V3-style MLA, with q_lora), unlimited-ocr (DeepSeek v1 MoE decoder). The distinguishing config of deepseek-v2-lite: q_lora_rank = null (no query compression, unlike V3 which always has q_lora), MLA head dims qk 192 = 128 nope + 64 rope / v 128, 64 routed + 2 shared experts, 6 active, n_group = topk_group = 1, first_k_dense_replace = 1.
Leading hypothesis
The q_lora_rank = null branch of the MLA attention in src/models/deepseek_v2.rs is exercised by DeepSeek-V2-Lite but not by the working V3 siblings (which always compress the query), which makes it the prime suspect for an uninitialized/oversized scratch buffer, an incorrect mask/shape, or a KV-cache slice read before it is written. The MoE gather/expert-selection path (64 experts, 6 active, grouped routing degenerate at n_group=1) is the secondary suspect.
Investigation plan
Deterministic repro harness. Reuse the coherence probe from the fix(models): DeepSeek-V2-Lite emits nondeterministic garbage on GB10 CUDA (broken since the 2026-05-19 benchmark) #817 investigation: N fresh greedy runs, classify each as coherent vs degenerate (repeated-char / repeated-gram / low-word-count), report a coherence rate. Establish a stable baseline rate on current main for both backends. This is the metric every subsequent step is measured against, because single runs are unreliable at ~50% flakiness.
Confirm the corruption is in prefill. Dump the first-token logits (or the top-k argmax) across several fresh runs. If they differ at --temp 0, prefill is nondeterministic and the search stays in prefill / weight-load; if they are identical but later tokens diverge, move the search into the decode step.
Sanitizer sweep. The CPU build makes this tractable: run the deepseek-v2-lite prefill under a sanitizer (compute-sanitizer --tool initcheck/memcheck for the CUDA path; ASan / Valgrind memcheck for the CPU build) to catch uninitialized reads and out-of-bounds directly. This is the highest-leverage step: a memory-safety defect that flips output run to run should surface here.
Narrow to the q_lora_rank = null MLA branch. Diff the attention forward in src/models/deepseek_v2.rs against the working V3 path (deepseek_v3.rs, used by youtu-vl). Look for a scratch/query buffer, a rope/nope split slice, an attention mask, or a KV-cache region that is allocated but read before it is fully written specifically when there is no q_lora projection.
Rule the MoE path in or out. Force a single-expert / dense configuration or a reference gather to see whether coherence goes to 100%, isolating the attention branch from the expert-gather branch.
Greedy (--temp 0) decode of models/deepseek-v2-lite-4bit is deterministic across fresh processes and coherent (correct answer to a simple prompt) on both CUDA and CPU, verified over many repetitions.
A regression test that pins greedy determinism (identical output across two in-process generations with a fixed seed, or a coherence assertion) for a deepseek_v2 model, so this cannot silently regress again.
The GB10 benchmark harness is updated to validate output coherence, not just "generates 100 tokens cleanly" (docs/benchmark_results/model_tests_gb10.md line 21). The false ✅ from a crash-free-but-garbage run is what hid this for two months; the harness should catch degenerate output.
The corrected docs/benchmark_results/model_tests_gb10.md row for deepseek-v2-lite reflects reality once fixed.
Refs
Root-cause follow-up to #817 (which holds the full measurement data and the corrected diagnosis). The fused decode-MoE default-on interaction is a separate, scoped degradation noted in #817; this issue targets the underlying nondeterminism.
Summary
This is the dedicated root-cause investigation for the DeepSeek-V2-Lite greedy nondeterminism established in #817. #817 determined that the "regressed since the benchmark" framing was wrong and that a mechanical bisect is not viable (there is no coherent-good anchor to bisect from), so a code fix was intentionally not forced. This issue owns the actual debugging: find and close the memory-safety defect that makes
deepseek_v2produce different greedy output across identical fresh-process runs, then restore deterministic, coherent decoding on both backends.Established facts (from #817)
--temp 0, repeated fresh-process runs of the same binary on the same prompt produce different outputs, most of them degenerate (a repeated single char like!or a repeated punctuation gram like}'_/</:). Greedy decode should be bit-identical run to run; it is not. That is a memory-safety class defect (uninitialized read, out-of-bounds, or a race), not a stable numeric error.MLXCEL_DEVICE=cpu), so it is not a CUDA-kernel-only fault. The fused decode-MoE default-on (perf(moe): validate fused decode-MoE on M5, then enable MLXCEL_FUSED_MOE by default #282/perf(moe): default-on MLXCEL_FUSED_MOE, validated on M5 (#282) #285) is an additional CUDA-specific degradation stacked on top (it drops coherence further), but it is not the root.deepseek_v2geometry. Sibling models are stable and coherent on the same binary: qwen1.5-moe (switch_layers MoE, standard attention), youtu-vl (DeepSeek-V3-style MLA, with q_lora), unlimited-ocr (DeepSeek v1 MoE decoder). The distinguishing config of deepseek-v2-lite:q_lora_rank = null(no query compression, unlike V3 which always has q_lora), MLA head dims qk 192 = 128 nope + 64 rope / v 128, 64 routed + 2 shared experts, 6 active,n_group = topk_group = 1,first_k_dense_replace = 1.Leading hypothesis
The
q_lora_rank = nullbranch of the MLA attention insrc/models/deepseek_v2.rsis exercised by DeepSeek-V2-Lite but not by the working V3 siblings (which always compress the query), which makes it the prime suspect for an uninitialized/oversized scratch buffer, an incorrect mask/shape, or a KV-cache slice read before it is written. The MoE gather/expert-selection path (64 experts, 6 active, grouped routing degenerate atn_group=1) is the secondary suspect.Investigation plan
mainfor both backends. This is the metric every subsequent step is measured against, because single runs are unreliable at ~50% flakiness.--temp 0, prefill is nondeterministic and the search stays in prefill / weight-load; if they are identical but later tokens diverge, move the search into the decode step.compute-sanitizer --tool initcheck/memcheckfor the CUDA path; ASan / Valgrind memcheck for the CPU build) to catch uninitialized reads and out-of-bounds directly. This is the highest-leverage step: a memory-safety defect that flips output run to run should surface here.q_lora_rank = nullMLA branch. Diff the attention forward insrc/models/deepseek_v2.rsagainst the working V3 path (deepseek_v3.rs, used by youtu-vl). Look for a scratch/query buffer, a rope/nope split slice, an attention mask, or a KV-cache region that is allocated but read before it is fully written specifically when there is no q_lora projection.Acceptance criteria
--temp 0) decode ofmodels/deepseek-v2-lite-4bitis deterministic across fresh processes and coherent (correct answer to a simple prompt) on both CUDA and CPU, verified over many repetitions.docs/benchmark_results/model_tests_gb10.mdline 21). The false ✅ from a crash-free-but-garbage run is what hid this for two months; the harness should catch degenerate output.docs/benchmark_results/model_tests_gb10.mdrow for deepseek-v2-lite reflects reality once fixed.Refs
Root-cause follow-up to #817 (which holds the full measurement data and the corrected diagnosis). The fused decode-MoE default-on interaction is a separate, scoped degradation noted in #817; this issue targets the underlying nondeterminism.