diff --git a/Help/markdown/$dev.md b/Help/markdown/$dev.md new file mode 100644 index 0000000..44b7d02 --- /dev/null +++ b/Help/markdown/$dev.md @@ -0,0 +1,60 @@ +[← Help Contents](index.md) | [📘 NLP++ Textbook](NLP++_Textbook.md) + +# $dev + +## Purpose + +Check whether the engine was run in development mode (`nlp.exe -DEV`). Returns 1 under `-DEV`, else 0. + +Use it to gate an analyzer's own diagnostic passes on what the caller actually asked for, instead of hardcoding a flag in the analyzer. + +## Syntax + +``` +returnedBoolean = variableType("$dev") +``` + +``` +returnedBoolean - type: int + +variableType - type: G +``` + +## Returns + +Returns 1 if the engine was run with `-DEV`, 0 otherwise. + +## Remarks + +Before `$dev`, NLP++ had no way to ask which mode the engine was running in, so analyzers had to hardcode their debug flags and no engine switch could reach them. Gating a diagnostic flag on `$dev` lets the same analyzer run fast in production and verbose during development, with no edit between the two. + +`-DEV` also makes the engine itself far more expensive: it writes a full parse-tree dump per pass. On a 9 KB input against a 136-pass analyzer that was measured at 137 files / 44 MB and roughly 3.6x total run time. Do not use `-DEV` when measuring processing speed. + +`-SILENT` is not the opposite of `-DEV`. It selects the engine's quietest configuration, which suppresses the analyzer's **own** output files as well, so a `-SILENT` run produces no `out.txt` at all. For "run fast but still write my results", use the default mode (neither switch) together with `$dev`-gated diagnostics. + +**NEW 3.8.4** + +## Example + +``` +@CODE + # Diagnostic passes run only when the caller asked for them. + G("verbose") = G("$dev"); + + if (G("verbose")) + "trace.txt" << "pass " << str(G("$passnum")) << " reached\n"; +@@CODE +``` + +A pass whose only job is diagnostics can bail out immediately: + +``` +@CODE + if (!G("$dev")) + exitpass(); +@@CODE +``` + +## See Also + +[$silent]($silent.md), [$passnum]($passnum.md), [$isdirrun]($isdirrun.md), [Special Variables](NLP_PP_Stuff/Special_Variables.md#table) diff --git a/Help/markdown/$silent.md b/Help/markdown/$silent.md new file mode 100644 index 0000000..a49a85c --- /dev/null +++ b/Help/markdown/$silent.md @@ -0,0 +1,49 @@ +[← Help Contents](index.md) | [📘 NLP++ Textbook](NLP++_Textbook.md) + +# $silent + +## Purpose + +Check whether the engine was run in silent mode (`nlp.exe -SILENT`). Returns 1 under `-SILENT`, else 0. + +## Syntax + +``` +returnedBoolean = variableType("$silent") +``` + +``` +returnedBoolean - type: int + +variableType - type: G +``` + +## Returns + +Returns 1 if the engine was run with `-SILENT`, 0 otherwise. + +## Remarks + +`-SILENT` selects the engine's quietest configuration. It suppresses logs and dump files — and also the analyzer's **own** output files, so a `-SILENT` run writes no `out.txt` even if the analyzer asks for one. Anything an analyzer does under `-SILENT` other than build the parse tree is therefore invisible. + +Because of that, `$silent` is mostly useful for skipping work whose only product is a file that will not be written anyway. If what you want is "run fast but still produce my results", do not use `-SILENT`: run in the default mode (neither `-DEV` nor `-SILENT`) and gate diagnostics on [$dev]($dev.md) instead. + +`-SILENT` is not a speed switch. Measured against the default mode on a 9 KB input it made no difference (8.39 sec vs 8.41 sec) — the default already writes no dump files. The switch that costs real time is `-DEV`. + +**NEW 3.8.4** + +## Example + +``` +@CODE + # Nothing this pass produces can be written under -SILENT, so skip it. + if (G("$silent")) + exitpass(); + + "report.txt" << "clauses: " << str(G("clause count")) << "\n"; +@@CODE +``` + +## See Also + +[$dev]($dev.md), [$isdirrun]($isdirrun.md), [Special Variables](NLP_PP_Stuff/Special_Variables.md#table) diff --git a/Help/markdown/NLP_PP_Stuff/Special_Variables.md b/Help/markdown/NLP_PP_Stuff/Special_Variables.md index cace118..d8fc8bc 100644 --- a/Help/markdown/NLP_PP_Stuff/Special_Variables.md +++ b/Help/markdown/NLP_PP_Stuff/Special_Variables.md @@ -49,6 +49,8 @@ A consolidated table of NLP++ special variables is provided below. Each item in | [**$inputtail**](../$inputtail.md) | **G** | Get input file tail. E.g., "txt" | | **$passnum** | G | Get current pass number. **NEW 2.0.2.4** | | **$rulenum** | G | Get current rule number within current pass. **NEW 2.0.2.4** | +| [**$dev**](../$dev.md) | G | Check if the engine was run with **-DEV**. Returns 1 under -DEV, else 0. Use it to gate an analyzer's own diagnostic passes on what the caller asked for. **NEW 3.8.4** | +| [**$silent**](../$silent.md) | G | Check if the engine was run with **-SILENT**. Returns 1 under -SILENT, else 0. Note -SILENT also suppresses the analyzer's own output files. **NEW 3.8.4** | | [**$isdirrun**](../$isdirrun.md) | G | Check if the engine is processing a directory of files instead of an individual file. Returns 1 if it is analyzing a directory, else 0. **NEW 1.26.0** | | [**$isfirstfile**](../$isfirstfile.md) | G | Check if is first file processed in a directory. Returns 1 if it is the first file, else 0. **NEW 1.24.0** | | [**$islastfile**](../$islastfile.md) | G | Check if is last file processed in a directory. Returns 1 if it is the last file, else 0. **NEW 1.24.0** | diff --git a/Help/markdown/index.md b/Help/markdown/index.md index 0f02bc5..6c49e6c 100644 --- a/Help/markdown/index.md +++ b/Help/markdown/index.md @@ -94,6 +94,9 @@ to a topic. - **Parser Variables** - [$passnum]($passnum.md) - [$rulenum]($rulenum.md) + - **Run Mode Variables** + - [$dev]($dev.md) + - [$silent]($silent.md) - **Test Variables** - [$end]($end.md) - [$exists]($exists.md) @@ -596,6 +599,10 @@ above. They are grouped here so that every page remains reachable from this inde [$anaspath]($anaspath.md)  |  [$inputparent]($inputparent.md)  |  [$isdirrun]($isdirrun.md)  |  [$isfirstfile]($isfirstfile.md)  |  [$islastfile]($islastfile.md)  |  [$kbpath]($kbpath.md) +### Run Mode Variables + +[$dev]($dev.md)  |  [$silent]($silent.md) + ### Control Flow Keywords [else](else.md)  |  [if](if.md)  |  [while](while.md) diff --git a/Help/markdown/vscode/prompts/00-prime-claude.md b/Help/markdown/vscode/prompts/00-prime-claude.md index efc42b9..00048e9 100644 --- a/Help/markdown/vscode/prompts/00-prime-claude.md +++ b/Help/markdown/vscode/prompts/00-prime-claude.md @@ -1,5 +1,5 @@ Prime Claude for NLP++ - + This is everything you need to know to help me work on NLP++ analyzers. NLP++ is a rule-based programming language for natural language processing, run by the NLP engine. Read this in full before we start — afterward I will ask you to build a new analyzer, extend an existing one, write dictionaries and knowledge bases, generate test inputs, or debug why a rule fires. Everything you need is already installed on this machine at the paths below: - NLP engine executable (run analyzers with this): {{engineExe}} @@ -30,6 +30,8 @@ Build the analyzer by copying an appropriate template from the templates directo Rules of thumb — build on the libraries, never on hand-typed word lists: - **Look in the libraries before you write a vocabulary.** {{languagesDir}} already ships, per language, dictionaries and matching knowledge bases for countries, nationalities/demonyms, languages, months and days, first names and surnames, professions, US states and street suffixes, numbers, prepositions, determiners, pronouns, conjunctions and stop words — plus a full lexicon (en-full.dict / en-full.kbb) carrying part of speech, root, and verb/noun features. {{miscDir}} adds currencies, ISO country codes, telephone country codes, timezones, emojis, roman numerals and URL domains. A pass built on the twenty words you typed by hand is brittle: it handles the test file and nothing else. The library file handles the whole vocabulary. +- **Write the vocabulary for the text you have not seen yet.** The files in input/ are a handful of the ways people write; the analyzer will meet all the others. Whenever you add an entry to a word list, a dictionary or a knowledge base, add the family a person writing the rules would anticipate: abbreviations and acronyms ("BC" for birth certificate, "SSN", "DOB"), synonyms and near-synonyms ("ID card" beside "identification card", "legal name" beside "full name"), spelling, spacing and hyphenation variants ("e-mail" / "email", "zip code" / "zipcode"), singular and plural, and qualified forms that mean something different ("nursing license" vs "foreign nursing license"). Where the variation is productive, write a pattern instead of a list of strings — a rule on the parts of "city", "state", "country" + "of birth" covers "city and state of birth", which no single entry does. Fixing each miss as it turns up in the sample only fits the sample: an analyzer tuned that way to 100% on the 40 messages it was built from found about 80% of what the next 40 messages asked. +- **Keep a set of inputs you never tune on.** Build and fix on one sample; check on a second whose expected results you write down before you run the analyzer on it, and report that second score as the honest one. The moment you change a rule or an entry because of a miss in the second set, it has become part of the sample — draw a fresh set for the next honest score. - **Wire one in by copying the .dict and its matching .kbb into the analyzer's kb/user/.** The tokenizer auto-loads every .dict there and tags matching tokens. Keep the pair together — a .kbb with the same stem supplies the concept hierarchy and the ambiguity readings for its .dict. Files named `…full.dict` / `…full.kbb` lazy-load word by word, so even a large lexicon costs little at startup. For a heavy domain KB that only some inputs need, load it on demand with loaddict / loadkbb (a file name in kb/user, not a path), guarded by a global so it loads once. - **A dictionary changes what every rule sees; a knowledge base answers when asked.** A .dict in kb/user tags matching tokens for the whole run, in every pass — that is a blast radius, and it is the right one when the analyzer is built on that vocabulary. When one rule needs one fact, the KB is the instrument instead. Match on *shape* and validate in code: the rule is permissive, the loop is strict. A pass that picks US state codes out of a list matches a run of two-letter uppercase tokens separated by commas, then loops calling a KB lookup (`StateName()`) on each candidate — two characters, uppercase, and known to the knowledge base. Words like "multiple", "locations" and "in" fail at least one of those three tests and drop out on their own, without the rule having to name them. Loading a .dict to buy that one lookup would tag those codes in every other pass too, and a dictionary attribute like `s=` even changes tokenization — dicttok builds a typed node for those words, so plain _xALPHA rules stop matching them. - **Match on the attributes, not on the words.** Dictionary attributes ride on the node, so `@PRE <1,1> var("nationality");` selects every entry in the file, and `N("country",1)`, `N("iso3",1)` read the data straight off the node — no dictfindword call needed. Rules written this way keep working as the library file grows.