Skip to content

fix(models): dedicated root-cause debugging of the DeepSeek-V2-Lite greedy nondeterminism (follow-up to #817) #824

Description

@inureyes

UPDATE (root cause pinned): A dedicated debugging pass has pinned this to the MLX 0.32.1 bump 6c53673 (#704, 2026-07-09) via a clean deterministic bisect. Key reframing of #817's conclusion: with MLX_USE_CUDA_GRAPHS=0 the output is deterministic garbage (not a memory-safety nondeterminism), compute-sanitizer initcheck+memcheck are both clean, and the fault is an upstream MLX CUDA kernel change breaking deepseek_v2's naive-MLA shape (qk 192 / v 128); deepseek_v3's absorbed-MLA is unaffected. Full analysis and the concrete next steps (sub-bisect the upstream MLX range, then patch-overlay / upstream-report / absorbed-MLA-rewrite) are in the comment below. No fix merged yet; the investigation plan below still applies for the remaining upstream-narrowing work.


Summary

This is the dedicated root-cause investigation for the DeepSeek-V2-Lite greedy nondeterminism established in #817. #817 determined that the "regressed since the benchmark" framing was wrong and that a mechanical bisect is not viable (there is no coherent-good anchor to bisect from), so a code fix was intentionally not forced. This issue owns the actual debugging: find and close the memory-safety defect that makes deepseek_v2 produce different greedy output across identical fresh-process runs, then restore deterministic, coherent decoding on both backends.

Established facts (from #817)

  • Nondeterministic greedy decode. At --temp 0, repeated fresh-process runs of the same binary on the same prompt produce different outputs, most of them degenerate (a repeated single char like ! or a repeated punctuation gram like }'_ / </:). Greedy decode should be bit-identical run to run; it is not. That is a memory-safety class defect (uninitialized read, out-of-bounds, or a race), not a stable numeric error.
  • Backend-agnostic. Reproduced on both CUDA and CPU (MLXCEL_DEVICE=cpu), so it is not a CUDA-kernel-only fault. The fused decode-MoE default-on (perf(moe): validate fused decode-MoE on M5, then enable MLXCEL_FUSED_MOE by default #282/perf(moe): default-on MLXCEL_FUSED_MOE, validated on M5 (#282) #285) is an additional CUDA-specific degradation stacked on top (it drops coherence further), but it is not the root.
  • Pre-existing. At the 2026-05-19 benchmark commit (331ac6e), coherence is already only ~50% (measured 8/12 and 5/12 coherent over fresh greedy runs with fused MoE off). The benchmark's ✅ means only "generates 100 tokens without crashing", not coherent output, which is why this went unnoticed.
  • Corruption is at the first generated token, so it originates in prefill or weight loading, not late in decode.
  • Checkpoint is intact (shard sha256 match HuggingFace; unmodified since January).
  • Specific to the deepseek_v2 geometry. Sibling models are stable and coherent on the same binary: qwen1.5-moe (switch_layers MoE, standard attention), youtu-vl (DeepSeek-V3-style MLA, with q_lora), unlimited-ocr (DeepSeek v1 MoE decoder). The distinguishing config of deepseek-v2-lite: q_lora_rank = null (no query compression, unlike V3 which always has q_lora), MLA head dims qk 192 = 128 nope + 64 rope / v 128, 64 routed + 2 shared experts, 6 active, n_group = topk_group = 1, first_k_dense_replace = 1.

Leading hypothesis

The q_lora_rank = null branch of the MLA attention in src/models/deepseek_v2.rs is exercised by DeepSeek-V2-Lite but not by the working V3 siblings (which always compress the query), which makes it the prime suspect for an uninitialized/oversized scratch buffer, an incorrect mask/shape, or a KV-cache slice read before it is written. The MoE gather/expert-selection path (64 experts, 6 active, grouped routing degenerate at n_group=1) is the secondary suspect.

Investigation plan

  1. Deterministic repro harness. Reuse the coherence probe from the fix(models): DeepSeek-V2-Lite emits nondeterministic garbage on GB10 CUDA (broken since the 2026-05-19 benchmark) #817 investigation: N fresh greedy runs, classify each as coherent vs degenerate (repeated-char / repeated-gram / low-word-count), report a coherence rate. Establish a stable baseline rate on current main for both backends. This is the metric every subsequent step is measured against, because single runs are unreliable at ~50% flakiness.
  2. Confirm the corruption is in prefill. Dump the first-token logits (or the top-k argmax) across several fresh runs. If they differ at --temp 0, prefill is nondeterministic and the search stays in prefill / weight-load; if they are identical but later tokens diverge, move the search into the decode step.
  3. Sanitizer sweep. The CPU build makes this tractable: run the deepseek-v2-lite prefill under a sanitizer (compute-sanitizer --tool initcheck/memcheck for the CUDA path; ASan / Valgrind memcheck for the CPU build) to catch uninitialized reads and out-of-bounds directly. This is the highest-leverage step: a memory-safety defect that flips output run to run should surface here.
  4. Narrow to the q_lora_rank = null MLA branch. Diff the attention forward in src/models/deepseek_v2.rs against the working V3 path (deepseek_v3.rs, used by youtu-vl). Look for a scratch/query buffer, a rope/nope split slice, an attention mask, or a KV-cache region that is allocated but read before it is fully written specifically when there is no q_lora projection.
  5. Rule the MoE path in or out. Force a single-expert / dense configuration or a reference gather to see whether coherence goes to 100%, isolating the attention branch from the expert-gather branch.
  6. Verify the fix statistically. A candidate fix must drive the probe to ~100% coherent across many reps on BOTH backends, and must make greedy output byte-identical across fresh processes. Do not accept a single coherent run as evidence (see fix(models): DeepSeek-V2-Lite emits nondeterministic garbage on GB10 CUDA (broken since the 2026-05-19 benchmark) #817: two runtime toggles each "fixed" it once by coincidence).

Acceptance criteria

  • Greedy (--temp 0) decode of models/deepseek-v2-lite-4bit is deterministic across fresh processes and coherent (correct answer to a simple prompt) on both CUDA and CPU, verified over many repetitions.
  • A regression test that pins greedy determinism (identical output across two in-process generations with a fixed seed, or a coherence assertion) for a deepseek_v2 model, so this cannot silently regress again.
  • The GB10 benchmark harness is updated to validate output coherence, not just "generates 100 tokens cleanly" (docs/benchmark_results/model_tests_gb10.md line 21). The false ✅ from a crash-free-but-garbage run is what hid this for two months; the harness should catch degenerate output.
  • The corrected docs/benchmark_results/model_tests_gb10.md row for deepseek-v2-lite reflects reality once fixed.

Refs

Root-cause follow-up to #817 (which holds the full measurement data and the corrected diagnosis). The fused decode-MoE default-on interaction is a separate, scoped degradation noted in #817; this issue targets the underlying nondeterminism.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:modelsModel architectures, weights, loading, metadatapriority:highHigh prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions