Skip to content

feat: support Brownian tree noise in all noise injection samplers - #1899

Merged
leejet merged 3 commits into
leejet:masterfrom
wbruna:sd_brownian_tree
Sep 14, 2026
Merged

leejet merged 3 commits into
leejet:masterfrom
wbruna:sd_brownian_tree

Conversation

@wbruna

@wbruna wbruna commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds a new extra sampling parameter noise_sampler=brownian_tree to optionally use Brownian tree noise in LCM and all ancestral samplers.

Related Issue / Discussion

#1743

Checklist

@vmobilis

vmobilis commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

@wbruna, maybe it's possible to make it more generic – as an option for --sampler-rng or even --rng?

Because it looks like noise_sampler is duplicating their functionality.

Probably is not so good idea, because it does not substitute, but rather extends the rng.

@vmobilis

vmobilis commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

@wbruna, hello, I digged a little more, currently the BrownianTreeNoiseSampler() explicitly uses std::make_shared<STDDefaultRNG>():

https://github.com/wbruna/stable-diffusion.cpp/blob/sd_brownian_tree/src/runtime/denoiser.hpp#L2724

auto rng = std::make_shared<STDDefaultRNG>();

(Which is the std option of --sampler-rng):
https://github.com/wbruna/stable-diffusion.cpp/blob/sd_brownian_tree/src/stable-diffusion.cpp#L659

if (rng_type == STD_DEFAULT_RNG) {
    return std::make_shared<STDDefaultRNG>();
}

Can you pass the application-wide rng, as in IIDGaussianNoiseSampler()?

https://github.com/wbruna/stable-diffusion.cpp/blob/sd_brownian_tree/src/runtime/denoiser.hpp#L2702

return sd::Tensor<float>::randn(shape, rng);

Then it should allow to use --sampler-rng choice of std, cpu and cuda in Brownian extension.

@wbruna

wbruna commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Sounds sensible to me. But maybe @fszontagh adopted it to match results from ComfyUI? If so, we may want to keep it as the default for dpm++2m_sde_bt? (but then the rng choice would need to be moved from model loading time to an inference parameter - which also sounds sensible to me, but maybe a bit too much for this PR 🙂)

@vmobilis

vmobilis commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

@wbruna, the std option should pass exactly the same <STDDefaultRNG> as currently is in BrownianTreeNoiseSampler(), so with --sampler-rng std the noise_sampler=brownian_tree should behave exactly as before, if I'm not mistaken.

And probably there is no need to move anything, because you are already using the app-wide rng in IIDGaussianNoiseSampler() and only need to pass it to BrownianTreeNoiseSampler() as well.

@fszontagh

Copy link
Copy Markdown
Contributor

@wbruna Not a ComfyUI parity thing - mine is a hand-written bridge, ComfyUI uses torchsde.BrownianTree, so the noise never matched anyway. No need to keep STDDefaultRNG as the default for dpm++2m_sde_bt.

One caveat: bridge() re-seeds per node, so passing the shared app RNG would re-seed it mid-sampling. Better to clone it. The hardcoded std also keeps the tree device-independent, which is worth preserving.

@vmobilis

vmobilis commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

@wbruna, hello, maybe something like wbruna#3?

I didn't test all variants yet, but so far iid and brownian_tree look same and others produce different images.

The changes are minimal, the rng in BrownianTreeNoiseSampler() is copied, as @fszontagh suggested.

P.s. If you don't need it, just ignore, I've simply shown what I meant. 😉

@leejet

leejet commented Sep 14, 2026

Copy link
Copy Markdown
Owner

Thanks for the discussion. I’ve added brownian_tree_rng with these options:

  • cpu (default)
  • cuda
  • std_default
  • sampler_rng

Each tree owns an independent RNG. sampler_rng clones the selected sampler RNG once, so reseeding tree nodes cannot reset the shared RNG.

The default is now cpu, using the same host-side implementation regardless of the inference backend. This changes existing Brownian outputs; std_default preserves the previous node-noise generation. The Brownian bridge and tree-seed derivation remain unchanged.

I also made tree initialization lazy, so samplers that never request noise no longer consume RNG state unnecessarily.

@leejet
leejet merged commit 07a85c7 into leejet:master Sep 14, 2026
9 checks passed
danielhanchen added a commit to unslothai/stable-diffusion.cpp that referenced this pull request Sep 21, 2026
* fix: preserve "token_refiner" token for MiniMax H3 LoRAs (leejet#1864)

* fix: fail with a message when MiniMax-H3 is run in img_gen mode (leejet#1863)

* feat: support INT8 ConvRot safetensors (leejet#1857)

* fix: replace free_compute_buffer with runner_done in vae (leejet#1872)

* sync: update ggml (leejet#1873)

* fix(ci): trigger builds for ggml updates

* feat: add taeh3 support (leejet#1874)

* fix: prevent gallocr hash overflow in tiny graph-cut segments (leejet#1880)

* fix: re-clamp streaming VRAM budget to currently free memory (leejet#1878)

* fix: mark graph cuts with both a prefix and a suffix (leejet#1883)

* fix: make max_order of lms sampler configurable (leejet#1885)

* fix: guard against missing sampler/scheduler names (leejet#1887)

* chore: format code

* fix: use sd_get_preview_interval() (leejet#1907)

* feat: configurable image / video compression (leejet#1909)

* feat: support standard Qwen3-VL weights for MiniMax-H3 (leejet#1910)

* feat: load scaled FP8 weights without upfront conversion (leejet#1913)

* fix: match exact weights in LLM config detection (leejet#1923)

* feat: add LTX-2.5 support (leejet#1893)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: correct MiniMax H3 reference audio encoding (leejet#1886)

* fix: correct MiniMax H3 audio Euler steps (leejet#1908)

* feat: use backend-native FP8 matmul when supported (leejet#1916)

* sync: update ggml

* feat: additional `--preview-interval` values (leejet#1915)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support numbering for preview images (leejet#1895)

* fix: use carrier sampling for MiniMax H3 audio (leejet#1924)

* feat: generalize temporal tiling across video VAEs (leejet#1926)

* feat: prefetch streamed layers during compute (leejet#1905)

Co-authored-by: leejet <leejet714@gmail.com>

* refactor: unify runner lifecycles and weight residency (leejet#1940)

* feat: add verbose logging and log-level selection (leejet#1941)

* feat: enable single-GPU auto-fit with tiered parameter placement (leejet#1942)

* fix: reuse graph cut plans across CFG passes (leejet#1943)

* refactor: split ggml extensions and move implementations to cpp files (leejet#1945)

* fix: preserve K-quantized embedding weights (leejet#1936)

* fix: correct SDXL embeddings loading (leejet#1939)

* refactor: unify model source and weight lifecycle management (leejet#1956)

* docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_extend assert) (leejet#1950)

* refactor: split generation pipeline out of stable-diffusion.cpp (leejet#1957)

* fix: enable VAE decode tiling fallback without auto-fit (leejet#1932)

* fix: preserve BF16 embedding weights for get_rows (leejet#1959)

* fix: handle invalid option numbers (leejet#1961)

* feat: expose the loaded model version name through the public API (leejet#1962)

* feat: add SenseNova U1.5 support (leejet#1935)

* fix: reuse graph plans when scale parameters change (leejet#1963)

* feat: add linear and attention scale overrides (leejet#1964)

* feat: preserve explicit backend assignments during auto-fit (leejet#1967)

* fix: guard GPU memory capacity and propagate encoding failures (leejet#1958)

* fix: bound plain-text runs in parse_prompt_attention regex (leejet#1919)

* feat: add Wan2.2 S2V (audio+img-to-video) support (leejet#1925)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: validate vision projector output dim against LLM hidden size (leejet#1918)

* fix: resolve MSVC narrowing conversion warnings (leejet#1969)

* feat: Add generation parameters into video metadata (leejet#1901)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support external Hugging Face tokenizer JSON files (leejet#1973)

* refactor: require external Gemma 2 and GPT-OSS tokenizers (leejet#1974)

* feat: support Brownian tree noise in all noise injection samplers (leejet#1899)

* fix: use tokenizer-specific pre-tokenization rules (leejet#1975)

* fix: remove vision_model. from ununsed tensors (leejet#1983)

* perf: eliminate temporary allocations in Philox rounds (leejet#1982)

* fix: honor flash attention flag in LLM text encoder attention (leejet#1987)

* refactor: remove obsolete unused tensor filtering (leejet#1984)

* perf: pad small attention heads to 64 for MMA Flash Attention (leejet#1992)

* perf: update ggml for faster direct convolutions (leejet#1993)

* perf: accelerate VAE direct 3D convolutions (leejet#1996)

* fix: propagate CUDA driver dependency to shared library consumers

* fix: prevent clip_preprocess center crop from exceeding the resized image (leejet#1995)

* perf: reduce CPU overhead in graph execution and sampling (leejet#1997)

* perf: parallelize host tensor elementwise and broadcast ops (leejet#1998)

* feat: support building with upstream ggml (leejet#1999)

* feat: add Qwen Image 2.1 support (leejet#1994)

* feat: restore legacy fp8 handling when building with upstream ggml (leejet#2001)

* fix: avoid passing ggml logs as format strings (leejet#2002)

* feat: add LLaDA-Image support (leejet#1968)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: add native CUDA SageAttention support (leejet#2005)

* fix: avoid narrowing conversion in SigVQ patch embedding and format code

* docs: update CONTRIBUTING.md

---------

Co-authored-by: stduhpf <stephduh@live.fr>
Co-authored-by: leejet <leejet714@gmail.com>
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
Co-authored-by: fszontagh <51741446+fszontagh@users.noreply.github.com>
Co-authored-by: Wagner Bruna <wbruna@users.noreply.github.com>
Co-authored-by: vmobilis <75476228+vmobilis@users.noreply.github.com>
Co-authored-by: Piotr Wilkin (ilintar) <ilintar@gmail.com>
Co-authored-by: jk212h20 <101200018+jk212h20@users.noreply.github.com>
Co-authored-by: assouan <750048+assouan@users.noreply.github.com>
Co-authored-by: nan <zjn32202153@gmail.com>
Co-authored-by: Hmission <62598659+Hmission@users.noreply.github.com>
Co-authored-by: LED-M <105789115+xledx@users.noreply.github.com>
Co-authored-by: Maphist0 <28743569+Maphist0@users.noreply.github.com>
Co-authored-by: George <35490284+noctrex@users.noreply.github.com>
Co-authored-by: Санька Четвёртый <CAHbKA-IV@mail.ru>
Co-authored-by: Lin Xuhao <linxuhao84@gmail.com>
Co-authored-by: Fabrice Aneche <akhenakh@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants