fix: 512px VAE decode OOM - effective half-size tiling + unconditional retry - #1932
Merged
Merged
Conversation
…l retry - backend_fit: prepare_vae_decode_retry_tiling now defaults rel_size to 0.5 when enabling spatial tiling. get_tile_sizes() defaulted to rel_size=1.0 (factor branch wins) which produced a full-latent tile - tiling was a no-op and 512px decode could exceed device buffer limits (Adreno 740 ~1.94GB -> ~416MB with half tiles). - stable-diffusion: retry VAE decode with tiling on decode failure (empty result, typically OOM) regardless of --auto-fit, so small-buffer mobile GPUs recover automatically. Fires only on failure; happy path unchanged. Evidence (Pocket Chick, K Pad Mali-G925 / Adreno 740): - 512px decode: 1.94GB -> 416MB buffer, renders green channel correctly (prior NaN/white-image on SD3.5 OpenCL); 10-step ~45.8s (0.4% accuracy loss vs non-tiled on K90), 13 tiles; Z-Image K90 39.7s.
danielhanchen
added a commit
to unslothai/stable-diffusion.cpp
that referenced
this pull request
Sep 21, 2026
* fix: preserve "token_refiner" token for MiniMax H3 LoRAs (leejet#1864) * fix: fail with a message when MiniMax-H3 is run in img_gen mode (leejet#1863) * feat: support INT8 ConvRot safetensors (leejet#1857) * fix: replace free_compute_buffer with runner_done in vae (leejet#1872) * sync: update ggml (leejet#1873) * fix(ci): trigger builds for ggml updates * feat: add taeh3 support (leejet#1874) * fix: prevent gallocr hash overflow in tiny graph-cut segments (leejet#1880) * fix: re-clamp streaming VRAM budget to currently free memory (leejet#1878) * fix: mark graph cuts with both a prefix and a suffix (leejet#1883) * fix: make max_order of lms sampler configurable (leejet#1885) * fix: guard against missing sampler/scheduler names (leejet#1887) * chore: format code * fix: use sd_get_preview_interval() (leejet#1907) * feat: configurable image / video compression (leejet#1909) * feat: support standard Qwen3-VL weights for MiniMax-H3 (leejet#1910) * feat: load scaled FP8 weights without upfront conversion (leejet#1913) * fix: match exact weights in LLM config detection (leejet#1923) * feat: add LTX-2.5 support (leejet#1893) Co-authored-by: leejet <leejet714@gmail.com> * fix: correct MiniMax H3 reference audio encoding (leejet#1886) * fix: correct MiniMax H3 audio Euler steps (leejet#1908) * feat: use backend-native FP8 matmul when supported (leejet#1916) * sync: update ggml * feat: additional `--preview-interval` values (leejet#1915) Co-authored-by: leejet <leejet714@gmail.com> * feat: support numbering for preview images (leejet#1895) * fix: use carrier sampling for MiniMax H3 audio (leejet#1924) * feat: generalize temporal tiling across video VAEs (leejet#1926) * feat: prefetch streamed layers during compute (leejet#1905) Co-authored-by: leejet <leejet714@gmail.com> * refactor: unify runner lifecycles and weight residency (leejet#1940) * feat: add verbose logging and log-level selection (leejet#1941) * feat: enable single-GPU auto-fit with tiered parameter placement (leejet#1942) * fix: reuse graph cut plans across CFG passes (leejet#1943) * refactor: split ggml extensions and move implementations to cpp files (leejet#1945) * fix: preserve K-quantized embedding weights (leejet#1936) * fix: correct SDXL embeddings loading (leejet#1939) * refactor: unify model source and weight lifecycle management (leejet#1956) * docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_extend assert) (leejet#1950) * refactor: split generation pipeline out of stable-diffusion.cpp (leejet#1957) * fix: enable VAE decode tiling fallback without auto-fit (leejet#1932) * fix: preserve BF16 embedding weights for get_rows (leejet#1959) * fix: handle invalid option numbers (leejet#1961) * feat: expose the loaded model version name through the public API (leejet#1962) * feat: add SenseNova U1.5 support (leejet#1935) * fix: reuse graph plans when scale parameters change (leejet#1963) * feat: add linear and attention scale overrides (leejet#1964) * feat: preserve explicit backend assignments during auto-fit (leejet#1967) * fix: guard GPU memory capacity and propagate encoding failures (leejet#1958) * fix: bound plain-text runs in parse_prompt_attention regex (leejet#1919) * feat: add Wan2.2 S2V (audio+img-to-video) support (leejet#1925) Co-authored-by: leejet <leejet714@gmail.com> * fix: validate vision projector output dim against LLM hidden size (leejet#1918) * fix: resolve MSVC narrowing conversion warnings (leejet#1969) * feat: Add generation parameters into video metadata (leejet#1901) Co-authored-by: leejet <leejet714@gmail.com> * feat: support external Hugging Face tokenizer JSON files (leejet#1973) * refactor: require external Gemma 2 and GPT-OSS tokenizers (leejet#1974) * feat: support Brownian tree noise in all noise injection samplers (leejet#1899) * fix: use tokenizer-specific pre-tokenization rules (leejet#1975) * fix: remove vision_model. from ununsed tensors (leejet#1983) * perf: eliminate temporary allocations in Philox rounds (leejet#1982) * fix: honor flash attention flag in LLM text encoder attention (leejet#1987) * refactor: remove obsolete unused tensor filtering (leejet#1984) * perf: pad small attention heads to 64 for MMA Flash Attention (leejet#1992) * perf: update ggml for faster direct convolutions (leejet#1993) * perf: accelerate VAE direct 3D convolutions (leejet#1996) * fix: propagate CUDA driver dependency to shared library consumers * fix: prevent clip_preprocess center crop from exceeding the resized image (leejet#1995) * perf: reduce CPU overhead in graph execution and sampling (leejet#1997) * perf: parallelize host tensor elementwise and broadcast ops (leejet#1998) * feat: support building with upstream ggml (leejet#1999) * feat: add Qwen Image 2.1 support (leejet#1994) * feat: restore legacy fp8 handling when building with upstream ggml (leejet#2001) * fix: avoid passing ggml logs as format strings (leejet#2002) * feat: add LLaDA-Image support (leejet#1968) Co-authored-by: leejet <leejet714@gmail.com> * feat: add native CUDA SageAttention support (leejet#2005) * fix: avoid narrowing conversion in SigVQ patch embedding and format code * docs: update CONTRIBUTING.md --------- Co-authored-by: stduhpf <stephduh@live.fr> Co-authored-by: leejet <leejet714@gmail.com> Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com> Co-authored-by: fszontagh <51741446+fszontagh@users.noreply.github.com> Co-authored-by: Wagner Bruna <wbruna@users.noreply.github.com> Co-authored-by: vmobilis <75476228+vmobilis@users.noreply.github.com> Co-authored-by: Piotr Wilkin (ilintar) <ilintar@gmail.com> Co-authored-by: jk212h20 <101200018+jk212h20@users.noreply.github.com> Co-authored-by: assouan <750048+assouan@users.noreply.github.com> Co-authored-by: nan <zjn32202153@gmail.com> Co-authored-by: Hmission <62598659+Hmission@users.noreply.github.com> Co-authored-by: LED-M <105789115+xledx@users.noreply.github.com> Co-authored-by: Maphist0 <28743569+Maphist0@users.noreply.github.com> Co-authored-by: George <35490284+noctrex@users.noreply.github.com> Co-authored-by: Санька Четвёртый <CAHbKA-IV@mail.ru> Co-authored-by: Lin Xuhao <linxuhao84@gmail.com> Co-authored-by: Fabrice Aneche <akhenakh@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two small fixes so 512px VAE decode fits small mobile GPU buffers without resorting to full-image fallbacks:
effective tiling (
src/core/backend_fit.cpp):prepare_vae_decode_retry_tiling()enables spatial tiling but did not setrel_size_x/y, andget_tile_sizes()defaults torel_size=1.0when the factor branch wins - so the "tile" was the full latent (tiling a no-op). Defaultrel_sizeto 0.5 when enabling spatial tiling: 512px decode goes from ~1.94GB to ~416MB of buffers on Adreno 740, with standard VAE tiling overlap (no quality impact, measured 0.4% accuracy delta vs non-tiled on K90 10-step).unconditional retry (
src/stable-diffusion.cpp): retry VAE decode with tiling whendecoded.empty()(typically OOM) regardless of--auto-fit, which defaults to off. Without this, small-buffer devices just get a black/empty image unless the user happens to pass--auto-fit. The retry only fires after a failed decode - the happy path is unchanged, andprepare_vae_decode_retry_tilingstill returns false on second failure so the loop terminates.Motivation
Pocket Chick (React Native app, forked from PocketPal) runs SD3.5 / Z-Image-Turbo on Android phones (Mali-G925, Adreno 740). At 512px the OpenCL decode graph previously exceeded device buffer limits (1.94GB requested) and produced blank/NaN output on some devices. These two fixes are the minimal upstream-friendly version of what we validated on device.
Evidence (on-device, K Pad Mali-G925 / Adreno 740)
Scope
--auto-fitbehavior intentionally broadened (see summary 2).