Skip to content

fix: bound plain-text runs in parse_prompt_attention regex - #1919

Merged
leejet merged 7 commits into
leejet:masterfrom
fszontagh:fix/prompt-attention-regex-stack
Sep 13, 2026
Merged

leejet merged 7 commits into
leejet:masterfrom
fszontagh:fix/prompt-attention-regex-stack

Conversation

@fszontagh

Copy link
Copy Markdown
Contributor

Summary

parse_prompt_attention lexes the prompt with an unbounded plain-text alternative. libstdc++'s std::regex recurses once per matched character, so a prompt of a few tens of kilobytes overflows the stack and segfaults before tokenization.

Bound the run length. Splitting a long run is safe because the function's final pass merges adjacent segments of equal weight, and B is excluded from the character class, so a chunk boundary can never fall inside a BREAK.

Related Issue / Discussion

None.

Additional Information

A 45KB prompt exits 139 (SIGSEGV) during tokenization before the change and completes normally after it.

Parsing results are unchanged: comparing old and new output across weighted parentheses, nested ((...)), [...], escaped \(, multiple BREAKs, BREAKING and bare B, bare colons, the empty string, multi-byte UTF-8, and a 21KB run both bare and inside (...:1.3), every case is byte-identical.

Checklist

@leejet leejet left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The plain-text repetition is bounded now, but the same stack-overflow path still exists in the weight alternative: :([+-]?[.\d]+)\) contains another unbounded +.

With libstdc++ (g++ 14.2, default 8 MiB stack), the PR still segfaults for e.g.:

":" + std::string(30000, '1')

A closing ) is not required: the regex recursively consumes the digit run while attempting the weight alternative before falling back to the literal : alternative.

Could we bound the numeric run as well (to a realistic float length), or avoid unbounded std::regex repetitions in this lexer entirely? We should also make sure oversized/invalid weights cannot terminate through std::stof.

@fszontagh

Copy link
Copy Markdown
Contributor Author

Good catch, thanks. Bounded the weight run as well and switched the parse to strtof with a finite check, so an oversized or invalid weight can no longer throw or poison the multipliers.

I had missed it because my equivalence harness ran at -O2, where the smaller frames hide the overflow. At -O0 your ":" + std::string(30000, '1') reproduces exactly as you describe, and is fine after the change.

Parsing is unchanged - old vs new output is byte-identical across weighted parens, nested brackets, escapes, BREAK, bare B, colons, empty input, UTF-8, and long runs both bare and inside (...:1.3).

@fszontagh
fszontagh force-pushed the fix/prompt-attention-regex-stack branch from 1dcaebe to 4066ca0 Compare September 7, 2026 06:11

@leejet leejet left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

max_weight_chars = 32 introduces a parsing regression: valid weights longer than 32 characters are treated as normal prompt text instead of weights. Also, please validate the strtof end pointer, otherwise invalid values like . or +. are silently accepted as 0.0.

@fszontagh

Copy link
Copy Markdown
Contributor Author

Both fixed. Dropped the weight alternative from the regex and lex the number by hand after a : instead. strtof loops instead of recursing, so there is no stack to blow and no length to cap - max_weight_chars is gone rather than retuned.

The end pointer is checked now, so . and +. fall back to literal text instead of 0.0. Same for 1.2.3, which used to parse as 1.2 through the std::stof prefix parse - flagging it since you did not name that one.

Diffed against the previous revision over 29 targeted cases and ~290k fuzzed prompts: those four are the only differences, the rest is byte-identical. ":" + std::string(30000, '1') survives at -O0 with an 8 MiB stack.

@leejet
leejet merged commit 4a7da26 into leejet:master Sep 13, 2026
10 checks passed
@fszontagh
fszontagh deleted the fix/prompt-attention-regex-stack branch September 15, 2026 20:25
danielhanchen added a commit to unslothai/stable-diffusion.cpp that referenced this pull request Sep 21, 2026
* fix: preserve "token_refiner" token for MiniMax H3 LoRAs (leejet#1864)

* fix: fail with a message when MiniMax-H3 is run in img_gen mode (leejet#1863)

* feat: support INT8 ConvRot safetensors (leejet#1857)

* fix: replace free_compute_buffer with runner_done in vae (leejet#1872)

* sync: update ggml (leejet#1873)

* fix(ci): trigger builds for ggml updates

* feat: add taeh3 support (leejet#1874)

* fix: prevent gallocr hash overflow in tiny graph-cut segments (leejet#1880)

* fix: re-clamp streaming VRAM budget to currently free memory (leejet#1878)

* fix: mark graph cuts with both a prefix and a suffix (leejet#1883)

* fix: make max_order of lms sampler configurable (leejet#1885)

* fix: guard against missing sampler/scheduler names (leejet#1887)

* chore: format code

* fix: use sd_get_preview_interval() (leejet#1907)

* feat: configurable image / video compression (leejet#1909)

* feat: support standard Qwen3-VL weights for MiniMax-H3 (leejet#1910)

* feat: load scaled FP8 weights without upfront conversion (leejet#1913)

* fix: match exact weights in LLM config detection (leejet#1923)

* feat: add LTX-2.5 support (leejet#1893)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: correct MiniMax H3 reference audio encoding (leejet#1886)

* fix: correct MiniMax H3 audio Euler steps (leejet#1908)

* feat: use backend-native FP8 matmul when supported (leejet#1916)

* sync: update ggml

* feat: additional `--preview-interval` values (leejet#1915)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support numbering for preview images (leejet#1895)

* fix: use carrier sampling for MiniMax H3 audio (leejet#1924)

* feat: generalize temporal tiling across video VAEs (leejet#1926)

* feat: prefetch streamed layers during compute (leejet#1905)

Co-authored-by: leejet <leejet714@gmail.com>

* refactor: unify runner lifecycles and weight residency (leejet#1940)

* feat: add verbose logging and log-level selection (leejet#1941)

* feat: enable single-GPU auto-fit with tiered parameter placement (leejet#1942)

* fix: reuse graph cut plans across CFG passes (leejet#1943)

* refactor: split ggml extensions and move implementations to cpp files (leejet#1945)

* fix: preserve K-quantized embedding weights (leejet#1936)

* fix: correct SDXL embeddings loading (leejet#1939)

* refactor: unify model source and weight lifecycle management (leejet#1956)

* docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_extend assert) (leejet#1950)

* refactor: split generation pipeline out of stable-diffusion.cpp (leejet#1957)

* fix: enable VAE decode tiling fallback without auto-fit (leejet#1932)

* fix: preserve BF16 embedding weights for get_rows (leejet#1959)

* fix: handle invalid option numbers (leejet#1961)

* feat: expose the loaded model version name through the public API (leejet#1962)

* feat: add SenseNova U1.5 support (leejet#1935)

* fix: reuse graph plans when scale parameters change (leejet#1963)

* feat: add linear and attention scale overrides (leejet#1964)

* feat: preserve explicit backend assignments during auto-fit (leejet#1967)

* fix: guard GPU memory capacity and propagate encoding failures (leejet#1958)

* fix: bound plain-text runs in parse_prompt_attention regex (leejet#1919)

* feat: add Wan2.2 S2V (audio+img-to-video) support (leejet#1925)

Co-authored-by: leejet <leejet714@gmail.com>

* fix: validate vision projector output dim against LLM hidden size (leejet#1918)

* fix: resolve MSVC narrowing conversion warnings (leejet#1969)

* feat: Add generation parameters into video metadata (leejet#1901)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: support external Hugging Face tokenizer JSON files (leejet#1973)

* refactor: require external Gemma 2 and GPT-OSS tokenizers (leejet#1974)

* feat: support Brownian tree noise in all noise injection samplers (leejet#1899)

* fix: use tokenizer-specific pre-tokenization rules (leejet#1975)

* fix: remove vision_model. from ununsed tensors (leejet#1983)

* perf: eliminate temporary allocations in Philox rounds (leejet#1982)

* fix: honor flash attention flag in LLM text encoder attention (leejet#1987)

* refactor: remove obsolete unused tensor filtering (leejet#1984)

* perf: pad small attention heads to 64 for MMA Flash Attention (leejet#1992)

* perf: update ggml for faster direct convolutions (leejet#1993)

* perf: accelerate VAE direct 3D convolutions (leejet#1996)

* fix: propagate CUDA driver dependency to shared library consumers

* fix: prevent clip_preprocess center crop from exceeding the resized image (leejet#1995)

* perf: reduce CPU overhead in graph execution and sampling (leejet#1997)

* perf: parallelize host tensor elementwise and broadcast ops (leejet#1998)

* feat: support building with upstream ggml (leejet#1999)

* feat: add Qwen Image 2.1 support (leejet#1994)

* feat: restore legacy fp8 handling when building with upstream ggml (leejet#2001)

* fix: avoid passing ggml logs as format strings (leejet#2002)

* feat: add LLaDA-Image support (leejet#1968)

Co-authored-by: leejet <leejet714@gmail.com>

* feat: add native CUDA SageAttention support (leejet#2005)

* fix: avoid narrowing conversion in SigVQ patch embedding and format code

* docs: update CONTRIBUTING.md

---------

Co-authored-by: stduhpf <stephduh@live.fr>
Co-authored-by: leejet <leejet714@gmail.com>
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
Co-authored-by: fszontagh <51741446+fszontagh@users.noreply.github.com>
Co-authored-by: Wagner Bruna <wbruna@users.noreply.github.com>
Co-authored-by: vmobilis <75476228+vmobilis@users.noreply.github.com>
Co-authored-by: Piotr Wilkin (ilintar) <ilintar@gmail.com>
Co-authored-by: jk212h20 <101200018+jk212h20@users.noreply.github.com>
Co-authored-by: assouan <750048+assouan@users.noreply.github.com>
Co-authored-by: nan <zjn32202153@gmail.com>
Co-authored-by: Hmission <62598659+Hmission@users.noreply.github.com>
Co-authored-by: LED-M <105789115+xledx@users.noreply.github.com>
Co-authored-by: Maphist0 <28743569+Maphist0@users.noreply.github.com>
Co-authored-by: George <35490284+noctrex@users.noreply.github.com>
Co-authored-by: Санька Четвёртый <CAHbKA-IV@mail.ru>
Co-authored-by: Lin Xuhao <linxuhao84@gmail.com>
Co-authored-by: Fabrice Aneche <akhenakh@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants