Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
60 changes: 60 additions & 0 deletions Help/markdown/$dev.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
[← Help Contents](index.md) | [📘 NLP++ Textbook](NLP++_Textbook.md)

# $dev

## Purpose

Check whether the engine was run in development mode (`nlp.exe -DEV`). Returns 1 under `-DEV`, else 0.

Use it to gate an analyzer's own diagnostic passes on what the caller actually asked for, instead of hardcoding a flag in the analyzer.

## Syntax

```
returnedBoolean = variableType("$dev")
```

```
returnedBoolean - type: int

variableType - type: G
```

## Returns

Returns 1 if the engine was run with `-DEV`, 0 otherwise.

## Remarks

Before `$dev`, NLP++ had no way to ask which mode the engine was running in, so analyzers had to hardcode their debug flags and no engine switch could reach them. Gating a diagnostic flag on `$dev` lets the same analyzer run fast in production and verbose during development, with no edit between the two.

`-DEV` also makes the engine itself far more expensive: it writes a full parse-tree dump per pass. On a 9 KB input against a 136-pass analyzer that was measured at 137 files / 44 MB and roughly 3.6x total run time. Do not use `-DEV` when measuring processing speed.

`-SILENT` is not the opposite of `-DEV`. It selects the engine's quietest configuration, which suppresses the analyzer's **own** output files as well, so a `-SILENT` run produces no `out.txt` at all. For "run fast but still write my results", use the default mode (neither switch) together with `$dev`-gated diagnostics.

**NEW 3.8.4**

## Example

```
@CODE
# Diagnostic passes run only when the caller asked for them.
G("verbose") = G("$dev");

if (G("verbose"))
"trace.txt" << "pass " << str(G("$passnum")) << " reached\n";
@@CODE
```

A pass whose only job is diagnostics can bail out immediately:

```
@CODE
if (!G("$dev"))
exitpass();
@@CODE
```

## See Also

[$silent]($silent.md), [$passnum]($passnum.md), [$isdirrun]($isdirrun.md), [Special Variables](NLP_PP_Stuff/Special_Variables.md#table)
49 changes: 49 additions & 0 deletions Help/markdown/$silent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
[← Help Contents](index.md) | [📘 NLP++ Textbook](NLP++_Textbook.md)

# $silent

## Purpose

Check whether the engine was run in silent mode (`nlp.exe -SILENT`). Returns 1 under `-SILENT`, else 0.

## Syntax

```
returnedBoolean = variableType("$silent")
```

```
returnedBoolean - type: int

variableType - type: G
```

## Returns

Returns 1 if the engine was run with `-SILENT`, 0 otherwise.

## Remarks

`-SILENT` selects the engine's quietest configuration. It suppresses logs and dump files — and also the analyzer's **own** output files, so a `-SILENT` run writes no `out.txt` even if the analyzer asks for one. Anything an analyzer does under `-SILENT` other than build the parse tree is therefore invisible.

Because of that, `$silent` is mostly useful for skipping work whose only product is a file that will not be written anyway. If what you want is "run fast but still produce my results", do not use `-SILENT`: run in the default mode (neither `-DEV` nor `-SILENT`) and gate diagnostics on [$dev]($dev.md) instead.

`-SILENT` is not a speed switch. Measured against the default mode on a 9 KB input it made no difference (8.39 sec vs 8.41 sec) — the default already writes no dump files. The switch that costs real time is `-DEV`.

**NEW 3.8.4**

## Example

```
@CODE
# Nothing this pass produces can be written under -SILENT, so skip it.
if (G("$silent"))
exitpass();

"report.txt" << "clauses: " << str(G("clause count")) << "\n";
@@CODE
```

## See Also

[$dev]($dev.md), [$isdirrun]($isdirrun.md), [Special Variables](NLP_PP_Stuff/Special_Variables.md#table)
2 changes: 2 additions & 0 deletions Help/markdown/NLP_PP_Stuff/Special_Variables.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,8 @@ A consolidated table of NLP++ special variables is provided below. Each item in
| [**$inputtail**](../$inputtail.md) | **G** | Get input file tail. E.g., "txt" |
| **$passnum** | G | Get current pass number. **NEW 2.0.2.4** |
| **$rulenum** | G | Get current rule number within current pass. **NEW 2.0.2.4** |
| [**$dev**](../$dev.md) | G | Check if the engine was run with **-DEV**. Returns 1 under -DEV, else 0. Use it to gate an analyzer's own diagnostic passes on what the caller asked for. **NEW 3.8.4** |
| [**$silent**](../$silent.md) | G | Check if the engine was run with **-SILENT**. Returns 1 under -SILENT, else 0. Note -SILENT also suppresses the analyzer's own output files. **NEW 3.8.4** |
| [**$isdirrun**](../$isdirrun.md) | G | Check if the engine is processing a directory of files instead of an individual file. Returns 1 if it is analyzing a directory, else 0. **NEW 1.26.0** |
| [**$isfirstfile**](../$isfirstfile.md) | G | Check if is first file processed in a directory. Returns 1 if it is the first file, else 0. **NEW 1.24.0** |
| [**$islastfile**](../$islastfile.md) | G | Check if is last file processed in a directory. Returns 1 if it is the last file, else 0. **NEW 1.24.0** |
Expand Down
7 changes: 7 additions & 0 deletions Help/markdown/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,9 @@ to a topic.
- **Parser Variables**
- [$passnum]($passnum.md)
- [$rulenum]($rulenum.md)
- **Run Mode Variables**
- [$dev]($dev.md)
- [$silent]($silent.md)
- **Test Variables**
- [$end]($end.md)
- [$exists]($exists.md)
Expand Down Expand Up @@ -596,6 +599,10 @@ above. They are grouped here so that every page remains reachable from this inde

[$anaspath]($anaspath.md) &nbsp;|&nbsp; [$inputparent]($inputparent.md) &nbsp;|&nbsp; [$isdirrun]($isdirrun.md) &nbsp;|&nbsp; [$isfirstfile]($isfirstfile.md) &nbsp;|&nbsp; [$islastfile]($islastfile.md) &nbsp;|&nbsp; [$kbpath]($kbpath.md)

### Run Mode Variables

[$dev]($dev.md) &nbsp;|&nbsp; [$silent]($silent.md)

### Control Flow Keywords

[else](else.md) &nbsp;|&nbsp; [if](if.md) &nbsp;|&nbsp; [while](while.md)
Expand Down
4 changes: 3 additions & 1 deletion Help/markdown/vscode/prompts/00-prime-claude.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
Prime Claude for NLP++
<!-- desc: Paste this into a fresh Claude session before asking for any NLP++ help — it is everything Claude needs to know to work on your analyzers: the installed engine, example, and template paths, the conventions that keep it on the rails (use the Knowledge Base template as intended, build results into a KB and emit JSON with JsonKB, build on the library dictionaries and knowledge bases instead of hand-typed word lists, reach for the knowledge base rather than a dictionary when only one rule needs the fact, and run with -WORK pointing at the engine), and the engine facts that trip up newcomers. No task to fill in — once Claude has read this, ask it for whatever you need. -->
<!-- desc: Paste this into a fresh Claude session before asking for any NLP++ help — it is everything Claude needs to know to work on your analyzers: the installed engine, example, and template paths, the conventions that keep it on the rails (use the Knowledge Base template as intended, build results into a KB and emit JSON with JsonKB, build on the library dictionaries and knowledge bases instead of hand-typed word lists, write vocabularies that anticipate the variations the test inputs do not show and keep a set of inputs you never tune on, reach for the knowledge base rather than a dictionary when only one rule needs the fact, and run with -WORK pointing at the engine), and the engine facts that trip up newcomers. No task to fill in — once Claude has read this, ask it for whatever you need. -->
This is everything you need to know to help me work on NLP++ analyzers. NLP++ is a rule-based programming language for natural language processing, run by the NLP engine. Read this in full before we start — afterward I will ask you to build a new analyzer, extend an existing one, write dictionaries and knowledge bases, generate test inputs, or debug why a rule fires. Everything you need is already installed on this machine at the paths below:

- NLP engine executable (run analyzers with this): {{engineExe}}
Expand Down Expand Up @@ -30,6 +30,8 @@ Build the analyzer by copying an appropriate template from the templates directo
Rules of thumb — build on the libraries, never on hand-typed word lists:

- **Look in the libraries before you write a vocabulary.** {{languagesDir}} already ships, per language, dictionaries and matching knowledge bases for countries, nationalities/demonyms, languages, months and days, first names and surnames, professions, US states and street suffixes, numbers, prepositions, determiners, pronouns, conjunctions and stop words — plus a full lexicon (en-full.dict / en-full.kbb) carrying part of speech, root, and verb/noun features. {{miscDir}} adds currencies, ISO country codes, telephone country codes, timezones, emojis, roman numerals and URL domains. A pass built on the twenty words you typed by hand is brittle: it handles the test file and nothing else. The library file handles the whole vocabulary.
- **Write the vocabulary for the text you have not seen yet.** The files in input/ are a handful of the ways people write; the analyzer will meet all the others. Whenever you add an entry to a word list, a dictionary or a knowledge base, add the family a person writing the rules would anticipate: abbreviations and acronyms ("BC" for birth certificate, "SSN", "DOB"), synonyms and near-synonyms ("ID card" beside "identification card", "legal name" beside "full name"), spelling, spacing and hyphenation variants ("e-mail" / "email", "zip code" / "zipcode"), singular and plural, and qualified forms that mean something different ("nursing license" vs "foreign nursing license"). Where the variation is productive, write a pattern instead of a list of strings — a rule on the parts of "city", "state", "country" + "of birth" covers "city and state of birth", which no single entry does. Fixing each miss as it turns up in the sample only fits the sample: an analyzer tuned that way to 100% on the 40 messages it was built from found about 80% of what the next 40 messages asked.
- **Keep a set of inputs you never tune on.** Build and fix on one sample; check on a second whose expected results you write down before you run the analyzer on it, and report that second score as the honest one. The moment you change a rule or an entry because of a miss in the second set, it has become part of the sample — draw a fresh set for the next honest score.
- **Wire one in by copying the .dict and its matching .kbb into the analyzer's kb/user/.** The tokenizer auto-loads every .dict there and tags matching tokens. Keep the pair together — a .kbb with the same stem supplies the concept hierarchy and the ambiguity readings for its .dict. Files named `…full.dict` / `…full.kbb` lazy-load word by word, so even a large lexicon costs little at startup. For a heavy domain KB that only some inputs need, load it on demand with loaddict / loadkbb (a file name in kb/user, not a path), guarded by a global so it loads once.
- **A dictionary changes what every rule sees; a knowledge base answers when asked.** A .dict in kb/user tags matching tokens for the whole run, in every pass — that is a blast radius, and it is the right one when the analyzer is built on that vocabulary. When one rule needs one fact, the KB is the instrument instead. Match on *shape* and validate in code: the rule is permissive, the loop is strict. A pass that picks US state codes out of a list matches a run of two-letter uppercase tokens separated by commas, then loops calling a KB lookup (`StateName()`) on each candidate — two characters, uppercase, and known to the knowledge base. Words like "multiple", "locations" and "in" fail at least one of those three tests and drop out on their own, without the rule having to name them. Loading a .dict to buy that one lookup would tag those codes in every other pass too, and a dictionary attribute like `s=` even changes tokenization — dicttok builds a typed node for those words, so plain _xALPHA rules stop matching them.
- **Match on the attributes, not on the words.** Dictionary attributes ride on the node, so `@PRE <1,1> var("nationality");` selects every entry in the file, and `N("country",1)`, `N("iso3",1)` read the data straight off the node — no dictfindword call needed. Rules written this way keep working as the library file grows.
Expand Down
Loading