Summary
nlp.exe hangs indefinitely (no termination, no error, must be killed) when the input text contains a Unicode em-dash (—, U+2014) or en-dash (–, U+2013). The ASCII hyphen-minus (-, U+002D) is processed normally. The hang occurs very early — only the first-pass tree (ana001.tree) is emitted, then the process spins forever — which points at tokenization / early input processing rather than any specific analyzer pass.
This makes any text using em-dashes for dialogue (very common in Portuguese, Spanish, French prose) impossible to analyze.
Minimal reproduction
Create a one-line input file and run any analyzer over it:
# hangs (em-dash U+2014):
printf '\xe2\x80\x94 Bom dia.' > emdash.txt
# hangs (en-dash U+2013):
printf '\xe2\x80\x93 Bom dia.' > endash.txt
# works fine (control, no dash):
printf 'Bom dia.' > ctl.txt
nlp.exe -ANA <analyzer-dir> -WORK <engine-dir> emdash.txt -DEV
| Input |
First char |
Result |
Bom dia. |
— |
completes normally |
— Bom dia. |
U+2014 em-dash |
hangs forever |
– Bom dia. |
U+2013 en-dash |
hangs forever |
- Bom dia. |
U+002D hyphen |
completes normally |
The em-dash can be anywhere in the text, not just the first character.
Expected behavior
The Unicode dash should be tokenized as a punctuation/dash token (like the ASCII hyphen) and analysis should complete.
Actual behavior
The process never terminates and consumes CPU; only ana001.tree is written to the _log directory. It has to be killed (e.g. timeout/taskkill). Observed processes hung for 11+ hours.
Discovered via
A Portuguese dialogue test whose sentences use travessões (em-dashes):
— Bom dia! — disse o velho padeiro. — Você tem pão fresco? ...
Environment
- Engine build: 3.1.25 (shipped with the VS Code NLP++ extension
dehilster.nlp-3.1.25); also reproduces with the nlp-engine-windows build.
- ICU: 78 (
icudt78 / icuin78 / icuuc78).
- OS: Windows 11 Pro (10.0.26200).
Suspected area
UTF-8 decoding of multi-byte dash code points in the tokenizer — likely a loop that fails to advance the byte/char cursor for these specific code points (cf. related Unicode handling in #500 "Reading currency like control characters in dictionary" and #488 "daccent function not working correctly").
Summary
nlp.exehangs indefinitely (no termination, no error, must be killed) when the input text contains a Unicode em-dash (—, U+2014) or en-dash (–, U+2013). The ASCII hyphen-minus (-, U+002D) is processed normally. The hang occurs very early — only the first-pass tree (ana001.tree) is emitted, then the process spins forever — which points at tokenization / early input processing rather than any specific analyzer pass.This makes any text using em-dashes for dialogue (very common in Portuguese, Spanish, French prose) impossible to analyze.
Minimal reproduction
Create a one-line input file and run any analyzer over it:
Bom dia.— Bom dia.– Bom dia.- Bom dia.The em-dash can be anywhere in the text, not just the first character.
Expected behavior
The Unicode dash should be tokenized as a punctuation/dash token (like the ASCII hyphen) and analysis should complete.
Actual behavior
The process never terminates and consumes CPU; only
ana001.treeis written to the_logdirectory. It has to be killed (e.g.timeout/taskkill). Observed processes hung for 11+ hours.Discovered via
A Portuguese dialogue test whose sentences use travessões (em-dashes):
Environment
dehilster.nlp-3.1.25); also reproduces with thenlp-engine-windowsbuild.icudt78/icuin78/icuuc78).Suspected area
UTF-8 decoding of multi-byte dash code points in the tokenizer — likely a loop that fails to advance the byte/char cursor for these specific code points (cf. related Unicode handling in #500 "Reading currency like control characters in dictionary" and #488 "daccent function not working correctly").