Skip to content

Feature: native python pass type in the analyzer sequence (+ pre-tokenization hook) #672

Description

@ddehilster

Summary

Add a first-class python pass type to analyzer.seq, so a pass can run a Python script as part of the pipeline — with two flavors:

  1. a normal Python pass that runs at its position in the sequence (after tokenization, with access to the current text-tree), and
  2. a pre-tokenization Python pass that runs before the tokenizer, receiving the raw input text and able to update the KB/dictionary before tokenization/segmentation happens.

Today the only way to reach external code from a pass is system() (shell escape), which is unstructured, blocking, and flaky (see #546). A native pass type would make Python integration a clean, reusable building block.

Motivation / driving use case

We've built a family of NLP++ phrase analyzers that share one template across six languages — parse-pt-br, parse-es-es, parse-it-it, parse-ro-ro, parse-fr-fr, and parse-zh-cn. Every one of them has the same recurring need: the dictionary is always missing words (proper nouns, neologisms, domain terms, OOV).

We want a language-agnostic dictionary gap-filler / enricher:

  • detect out-of-vocabulary words during analysis (uniform test across all languages: if (!dictfindword(word)) …),
  • hand them to Python to be defined/enriched (e.g., POS + lemma/gender/number for Romance languages; POS + pinyin/sense for Chinese — optionally via a local LLM such as Qwen2.5 through ollama, or jieba for Chinese segmentation/POS),
  • feed the results back into the dictionary so subsequent analysis improves.

This is fundamentally a Python integration problem, and a native pass type is the clean way to express it in the sequence. It also unlocks many other Python-shaped passes: embeddings, statistical NER, ML classifiers, external tokenizers/segmenters, normalization, etc.

Current state & limitations

Supported pass types today (from shipped analyzer.seq files): tokenize / dicttokz / dicttok (tokenizer), nlp (RUG/NLP++ rule file), rec (recursive nlp), and stub … end grouping markers. There is no Python pass type, and no pass runs before the tokenizer.

Working around it with system() has real problems:

  • unstructured data exchange (no clean way to hand the tree to Python or get structured data back),
  • it blocks the parse on the child process,
  • reliability issues (Built-in function system() not working #546),
  • and it still can't run before tokenization, which is exactly where dictionary updates must happen for dictionary-driven segmentation (e.g., Chinese).

Proposed design

Sequence syntax

python   <script>        # normal pass: runs at this position, tree available
pre      <script>        # pre-tokenization pass: runs before the tokenizer on raw text

(Exact keywords open to bikeshedding — e.g., pythonpre, or a modifier on python.)

Data exchange (phased)

MVP (no binding required). The engine invokes the script (python <script>) with context provided via environment variables / args, e.g. analyzer dir ($apppath), input file path, current pass name/number, and the path of the current tree dump (the engine already writes ana###.tree). Convention-based return: after the pass, the engine auto-take()s a conventionally named KB file the script may have written (e.g., <app>/kb/_python_out.kb). This alone enables the gap-filler: the Python pass updates the dictionary KB, and later passes see the new words via dictfindword.

Phase 2 (structured tree access). Expose read/write tree access to the Python pass via the existing engine bindings (there is already py-package-nlpengine / npm bindings) — pass a handle so the script can read nodes/vars and add/rename nodes, returning a structured result. This makes Python passes full peers of nlp passes.

Pre-tokenization hook

A pre pass receives the raw input text (and can rewrite it and/or update the KB) before the tokenizer runs. Rationale:

  • Dictionary-driven segmentation (Chinese): to recognize a newly discovered multi-character word in the same run, the dictionary must be updated before tokenization/segmentation — which is impossible today since every pass runs after the tokenizer.
  • Text normalization: clean problematic input before tokenization (e.g., normalize the Unicode em-dash that currently hangs the tokenizer, Engine hangs indefinitely on Unicode dash characters (em-dash U+2014 / en-dash U+2013) in input #669; normalize full-width punctuation; strip control chars).

Concrete example (gap-filler)

pre        prefill        # python: read raw text, fill dict with OOV defs (LLM/jieba), update KB
tokenize   nil
nlp        ...            # normal analyzer passes; dictfindword now sees the new words
nlp        gaps           # nlp pass: append still-unknown words to gaps.tsv for offline review
nlp        output

Acceptance criteria (MVP)

  • python <script> and a pre-tokenization variant are recognized in analyzer.seq and run at the correct point.
  • The script receives analyzer dir, input path, and pass identity (env or args).
  • A non-zero/failed script is reported as a pass error (not a silent skip).
  • A conventionally named KB file produced by the script is loaded so subsequent passes can use it.
  • Works headless (CLI nlp.exe) and from the VS Code extension.

Backward compatibility

Purely additive — new pass-type keywords; existing sequences (nlp/rec/tokenize/dicttokz/stub/end) are unaffected.

Alternatives considered

  • system() from an nlp pass — works partially but is unstructured, blocking, flaky (Built-in function system() not working #546), and cannot run before tokenization.
  • Fully out-of-band Python (run analyzer → post-process → re-run) — robust and we use it, but it can't influence the current run and isn't expressed in the sequence.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions