Skip to content

Engine hangs indefinitely on Unicode dash characters (em-dash U+2014 / en-dash U+2013) in input #669

Description

@ddehilster

Summary

nlp.exe hangs indefinitely (no termination, no error, must be killed) when the input text contains a Unicode em-dash (—, U+2014) or en-dash (–, U+2013). The ASCII hyphen-minus (-, U+002D) is processed normally. The hang occurs very early — only the first-pass tree (ana001.tree) is emitted, then the process spins forever — which points at tokenization / early input processing rather than any specific analyzer pass.

This makes any text using em-dashes for dialogue (very common in Portuguese, Spanish, French prose) impossible to analyze.

Minimal reproduction

Create a one-line input file and run any analyzer over it:

# hangs (em-dash U+2014):
printf '\xe2\x80\x94 Bom dia.' > emdash.txt

# hangs (en-dash U+2013):
printf '\xe2\x80\x93 Bom dia.' > endash.txt

# works fine (control, no dash):
printf 'Bom dia.' > ctl.txt

nlp.exe -ANA <analyzer-dir> -WORK <engine-dir> emdash.txt -DEV
Input First char Result
Bom dia. — completes normally
— Bom dia. U+2014 em-dash hangs forever
– Bom dia. U+2013 en-dash hangs forever
- Bom dia. U+002D hyphen completes normally

The em-dash can be anywhere in the text, not just the first character.

Expected behavior

The Unicode dash should be tokenized as a punctuation/dash token (like the ASCII hyphen) and analysis should complete.

Actual behavior

The process never terminates and consumes CPU; only ana001.tree is written to the _log directory. It has to be killed (e.g. timeout/taskkill). Observed processes hung for 11+ hours.

Discovered via

A Portuguese dialogue test whose sentences use travessões (em-dashes):

— Bom dia! — disse o velho padeiro. — Você tem pão fresco? ...

Environment

  • Engine build: 3.1.25 (shipped with the VS Code NLP++ extension dehilster.nlp-3.1.25); also reproduces with the nlp-engine-windows build.
  • ICU: 78 (icudt78 / icuin78 / icuuc78).
  • OS: Windows 11 Pro (10.0.26200).

Suspected area

UTF-8 decoding of multi-byte dash code points in the tokenizer — likely a loop that fails to advance the byte/char cursor for these specific code points (cf. related Unicode handling in #500 "Reading currency like control characters in dictionary" and #488 "daccent function not working correctly").

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions