You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add a first-class python pass type to analyzer.seq, so a pass can run a Python script as part of the pipeline — with two flavors:
a normal Python pass that runs at its position in the sequence (after tokenization, with access to the current text-tree), and
a pre-tokenization Python pass that runs before the tokenizer, receiving the raw input text and able to update the KB/dictionary before tokenization/segmentation happens.
Today the only way to reach external code from a pass is system() (shell escape), which is unstructured, blocking, and flaky (see #546). A native pass type would make Python integration a clean, reusable building block.
Motivation / driving use case
We've built a family of NLP++ phrase analyzers that share one template across six languages — parse-pt-br, parse-es-es, parse-it-it, parse-ro-ro, parse-fr-fr, and parse-zh-cn. Every one of them has the same recurring need: the dictionary is always missing words (proper nouns, neologisms, domain terms, OOV).
We want a language-agnostic dictionary gap-filler / enricher:
detect out-of-vocabulary words during analysis (uniform test across all languages: if (!dictfindword(word)) …),
hand them to Python to be defined/enriched (e.g., POS + lemma/gender/number for Romance languages; POS + pinyin/sense for Chinese — optionally via a local LLM such as Qwen2.5 through ollama, or jieba for Chinese segmentation/POS),
feed the results back into the dictionary so subsequent analysis improves.
This is fundamentally a Python integration problem, and a native pass type is the clean way to express it in the sequence. It also unlocks many other Python-shaped passes: embeddings, statistical NER, ML classifiers, external tokenizers/segmenters, normalization, etc.
Current state & limitations
Supported pass types today (from shipped analyzer.seq files): tokenize / dicttokz / dicttok (tokenizer), nlp (RUG/NLP++ rule file), rec (recursive nlp), and stub … end grouping markers. There is no Python pass type, and no pass runs before the tokenizer.
Working around it with system() has real problems:
unstructured data exchange (no clean way to hand the tree to Python or get structured data back),
and it still can't run before tokenization, which is exactly where dictionary updates must happen for dictionary-driven segmentation (e.g., Chinese).
Proposed design
Sequence syntax
python <script> # normal pass: runs at this position, tree available
pre <script> # pre-tokenization pass: runs before the tokenizer on raw text
(Exact keywords open to bikeshedding — e.g., pythonpre, or a modifier on python.)
Data exchange (phased)
MVP (no binding required). The engine invokes the script (python <script>) with context provided via environment variables / args, e.g. analyzer dir ($apppath), input file path, current pass name/number, and the path of the current tree dump (the engine already writes ana###.tree). Convention-based return: after the pass, the engine auto-take()s a conventionally named KB file the script may have written (e.g., <app>/kb/_python_out.kb). This alone enables the gap-filler: the Python pass updates the dictionary KB, and later passes see the new words via dictfindword.
Phase 2 (structured tree access). Expose read/write tree access to the Python pass via the existing engine bindings (there is already py-package-nlpengine / npm bindings) — pass a handle so the script can read nodes/vars and add/rename nodes, returning a structured result. This makes Python passes full peers of nlp passes.
Pre-tokenization hook
A pre pass receives the raw input text (and can rewrite it and/or update the KB) before the tokenizer runs. Rationale:
Dictionary-driven segmentation (Chinese): to recognize a newly discovered multi-character word in the same run, the dictionary must be updated before tokenization/segmentation — which is impossible today since every pass runs after the tokenizer.
pre prefill # python: read raw text, fill dict with OOV defs (LLM/jieba), update KB
tokenize nil
nlp ... # normal analyzer passes; dictfindword now sees the new words
nlp gaps # nlp pass: append still-unknown words to gaps.tsv for offline review
nlp output
Acceptance criteria (MVP)
python <script> and a pre-tokenization variant are recognized in analyzer.seq and run at the correct point.
The script receives analyzer dir, input path, and pass identity (env or args).
A non-zero/failed script is reported as a pass error (not a silent skip).
A conventionally named KB file produced by the script is loaded so subsequent passes can use it.
Works headless (CLI nlp.exe) and from the VS Code extension.
Backward compatibility
Purely additive — new pass-type keywords; existing sequences (nlp/rec/tokenize/dicttokz/stub/end) are unaffected.
Fully out-of-band Python (run analyzer → post-process → re-run) — robust and we use it, but it can't influence the current run and isn't expressed in the sequence.
Summary
Add a first-class
pythonpass type toanalyzer.seq, so a pass can run a Python script as part of the pipeline — with two flavors:Today the only way to reach external code from a pass is
system()(shell escape), which is unstructured, blocking, and flaky (see #546). A native pass type would make Python integration a clean, reusable building block.Motivation / driving use case
We've built a family of NLP++ phrase analyzers that share one template across six languages —
parse-pt-br,parse-es-es,parse-it-it,parse-ro-ro,parse-fr-fr, andparse-zh-cn. Every one of them has the same recurring need: the dictionary is always missing words (proper nouns, neologisms, domain terms, OOV).We want a language-agnostic dictionary gap-filler / enricher:
if (!dictfindword(word)) …),This is fundamentally a Python integration problem, and a native pass type is the clean way to express it in the sequence. It also unlocks many other Python-shaped passes: embeddings, statistical NER, ML classifiers, external tokenizers/segmenters, normalization, etc.
Current state & limitations
Supported pass types today (from shipped
analyzer.seqfiles):tokenize/dicttokz/dicttok(tokenizer),nlp(RUG/NLP++ rule file),rec(recursive nlp), andstub…endgrouping markers. There is no Python pass type, and no pass runs before the tokenizer.Working around it with
system()has real problems:Proposed design
Sequence syntax
(Exact keywords open to bikeshedding — e.g.,
pythonpre, or a modifier onpython.)Data exchange (phased)
MVP (no binding required). The engine invokes the script (
python <script>) with context provided via environment variables / args, e.g. analyzer dir ($apppath), input file path, current pass name/number, and the path of the current tree dump (the engine already writesana###.tree). Convention-based return: after the pass, the engine auto-take()s a conventionally named KB file the script may have written (e.g.,<app>/kb/_python_out.kb). This alone enables the gap-filler: the Python pass updates the dictionary KB, and later passes see the new words viadictfindword.Phase 2 (structured tree access). Expose read/write tree access to the Python pass via the existing engine bindings (there is already
py-package-nlpengine/ npm bindings) — pass a handle so the script can read nodes/vars and add/rename nodes, returning a structured result. This makes Python passes full peers ofnlppasses.Pre-tokenization hook
A
prepass receives the raw input text (and can rewrite it and/or update the KB) before the tokenizer runs. Rationale:Concrete example (gap-filler)
Acceptance criteria (MVP)
python <script>and a pre-tokenization variant are recognized inanalyzer.seqand run at the correct point.nlp.exe) and from the VS Code extension.Backward compatibility
Purely additive — new pass-type keywords; existing sequences (
nlp/rec/tokenize/dicttokz/stub/end) are unaffected.Alternatives considered
system()from annlppass — works partially but is unstructured, blocking, flaky (Built-in function system() not working #546), and cannot run before tokenization.