fix: bound plain-text runs in parse_prompt_attention regex - #1919
Conversation
leejet
left a comment
There was a problem hiding this comment.
The plain-text repetition is bounded now, but the same stack-overflow path still exists in the weight alternative: :([+-]?[.\d]+)\) contains another unbounded +.
With libstdc++ (g++ 14.2, default 8 MiB stack), the PR still segfaults for e.g.:
":" + std::string(30000, '1')A closing ) is not required: the regex recursively consumes the digit run while attempting the weight alternative before falling back to the literal : alternative.
Could we bound the numeric run as well (to a realistic float length), or avoid unbounded std::regex repetitions in this lexer entirely? We should also make sure oversized/invalid weights cannot terminate through std::stof.
|
Good catch, thanks. Bounded the weight run as well and switched the parse to I had missed it because my equivalence harness ran at Parsing is unchanged - old vs new output is byte-identical across weighted parens, nested brackets, escapes, |
1dcaebe to
4066ca0
Compare
leejet
left a comment
There was a problem hiding this comment.
max_weight_chars = 32 introduces a parsing regression: valid weights longer than 32 characters are treated as normal prompt text instead of weights. Also, please validate the strtof end pointer, otherwise invalid values like . or +. are silently accepted as 0.0.
|
Both fixed. Dropped the weight alternative from the regex and lex the number by hand after a The end pointer is checked now, so Diffed against the previous revision over 29 targeted cases and ~290k fuzzed prompts: those four are the only differences, the rest is byte-identical. |
* fix: preserve "token_refiner" token for MiniMax H3 LoRAs (leejet#1864) * fix: fail with a message when MiniMax-H3 is run in img_gen mode (leejet#1863) * feat: support INT8 ConvRot safetensors (leejet#1857) * fix: replace free_compute_buffer with runner_done in vae (leejet#1872) * sync: update ggml (leejet#1873) * fix(ci): trigger builds for ggml updates * feat: add taeh3 support (leejet#1874) * fix: prevent gallocr hash overflow in tiny graph-cut segments (leejet#1880) * fix: re-clamp streaming VRAM budget to currently free memory (leejet#1878) * fix: mark graph cuts with both a prefix and a suffix (leejet#1883) * fix: make max_order of lms sampler configurable (leejet#1885) * fix: guard against missing sampler/scheduler names (leejet#1887) * chore: format code * fix: use sd_get_preview_interval() (leejet#1907) * feat: configurable image / video compression (leejet#1909) * feat: support standard Qwen3-VL weights for MiniMax-H3 (leejet#1910) * feat: load scaled FP8 weights without upfront conversion (leejet#1913) * fix: match exact weights in LLM config detection (leejet#1923) * feat: add LTX-2.5 support (leejet#1893) Co-authored-by: leejet <leejet714@gmail.com> * fix: correct MiniMax H3 reference audio encoding (leejet#1886) * fix: correct MiniMax H3 audio Euler steps (leejet#1908) * feat: use backend-native FP8 matmul when supported (leejet#1916) * sync: update ggml * feat: additional `--preview-interval` values (leejet#1915) Co-authored-by: leejet <leejet714@gmail.com> * feat: support numbering for preview images (leejet#1895) * fix: use carrier sampling for MiniMax H3 audio (leejet#1924) * feat: generalize temporal tiling across video VAEs (leejet#1926) * feat: prefetch streamed layers during compute (leejet#1905) Co-authored-by: leejet <leejet714@gmail.com> * refactor: unify runner lifecycles and weight residency (leejet#1940) * feat: add verbose logging and log-level selection (leejet#1941) * feat: enable single-GPU auto-fit with tiered parameter placement (leejet#1942) * fix: reuse graph cut plans across CFG passes (leejet#1943) * refactor: split ggml extensions and move implementations to cpp files (leejet#1945) * fix: preserve K-quantized embedding weights (leejet#1936) * fix: correct SDXL embeddings loading (leejet#1939) * refactor: unify model source and weight lifecycle management (leejet#1956) * docs: reflect GGML_MAX_NAME value change in rpc docs (and in ggml_extend assert) (leejet#1950) * refactor: split generation pipeline out of stable-diffusion.cpp (leejet#1957) * fix: enable VAE decode tiling fallback without auto-fit (leejet#1932) * fix: preserve BF16 embedding weights for get_rows (leejet#1959) * fix: handle invalid option numbers (leejet#1961) * feat: expose the loaded model version name through the public API (leejet#1962) * feat: add SenseNova U1.5 support (leejet#1935) * fix: reuse graph plans when scale parameters change (leejet#1963) * feat: add linear and attention scale overrides (leejet#1964) * feat: preserve explicit backend assignments during auto-fit (leejet#1967) * fix: guard GPU memory capacity and propagate encoding failures (leejet#1958) * fix: bound plain-text runs in parse_prompt_attention regex (leejet#1919) * feat: add Wan2.2 S2V (audio+img-to-video) support (leejet#1925) Co-authored-by: leejet <leejet714@gmail.com> * fix: validate vision projector output dim against LLM hidden size (leejet#1918) * fix: resolve MSVC narrowing conversion warnings (leejet#1969) * feat: Add generation parameters into video metadata (leejet#1901) Co-authored-by: leejet <leejet714@gmail.com> * feat: support external Hugging Face tokenizer JSON files (leejet#1973) * refactor: require external Gemma 2 and GPT-OSS tokenizers (leejet#1974) * feat: support Brownian tree noise in all noise injection samplers (leejet#1899) * fix: use tokenizer-specific pre-tokenization rules (leejet#1975) * fix: remove vision_model. from ununsed tensors (leejet#1983) * perf: eliminate temporary allocations in Philox rounds (leejet#1982) * fix: honor flash attention flag in LLM text encoder attention (leejet#1987) * refactor: remove obsolete unused tensor filtering (leejet#1984) * perf: pad small attention heads to 64 for MMA Flash Attention (leejet#1992) * perf: update ggml for faster direct convolutions (leejet#1993) * perf: accelerate VAE direct 3D convolutions (leejet#1996) * fix: propagate CUDA driver dependency to shared library consumers * fix: prevent clip_preprocess center crop from exceeding the resized image (leejet#1995) * perf: reduce CPU overhead in graph execution and sampling (leejet#1997) * perf: parallelize host tensor elementwise and broadcast ops (leejet#1998) * feat: support building with upstream ggml (leejet#1999) * feat: add Qwen Image 2.1 support (leejet#1994) * feat: restore legacy fp8 handling when building with upstream ggml (leejet#2001) * fix: avoid passing ggml logs as format strings (leejet#2002) * feat: add LLaDA-Image support (leejet#1968) Co-authored-by: leejet <leejet714@gmail.com> * feat: add native CUDA SageAttention support (leejet#2005) * fix: avoid narrowing conversion in SigVQ patch embedding and format code * docs: update CONTRIBUTING.md --------- Co-authored-by: stduhpf <stephduh@live.fr> Co-authored-by: leejet <leejet714@gmail.com> Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com> Co-authored-by: fszontagh <51741446+fszontagh@users.noreply.github.com> Co-authored-by: Wagner Bruna <wbruna@users.noreply.github.com> Co-authored-by: vmobilis <75476228+vmobilis@users.noreply.github.com> Co-authored-by: Piotr Wilkin (ilintar) <ilintar@gmail.com> Co-authored-by: jk212h20 <101200018+jk212h20@users.noreply.github.com> Co-authored-by: assouan <750048+assouan@users.noreply.github.com> Co-authored-by: nan <zjn32202153@gmail.com> Co-authored-by: Hmission <62598659+Hmission@users.noreply.github.com> Co-authored-by: LED-M <105789115+xledx@users.noreply.github.com> Co-authored-by: Maphist0 <28743569+Maphist0@users.noreply.github.com> Co-authored-by: George <35490284+noctrex@users.noreply.github.com> Co-authored-by: Санька Четвёртый <CAHbKA-IV@mail.ru> Co-authored-by: Lin Xuhao <linxuhao84@gmail.com> Co-authored-by: Fabrice Aneche <akhenakh@users.noreply.github.com>
Summary
parse_prompt_attentionlexes the prompt with an unbounded plain-text alternative. libstdc++'sstd::regexrecurses once per matched character, so a prompt of a few tens of kilobytes overflows the stack and segfaults before tokenization.Bound the run length. Splitting a long run is safe because the function's final pass merges adjacent segments of equal weight, and
Bis excluded from the character class, so a chunk boundary can never fall inside aBREAK.Related Issue / Discussion
None.
Additional Information
A 45KB prompt exits 139 (SIGSEGV) during tokenization before the change and completes normally after it.
Parsing results are unchanged: comparing old and new output across weighted parentheses, nested
((...)),[...], escaped\(, multipleBREAKs,BREAKINGand bareB, bare colons, the empty string, multi-byte UTF-8, and a 21KB run both bare and inside(...:1.3), every case is byte-identical.Checklist