From 2d0aa6685868f8583b563148dd266445aa6c75ce Mon Sep 17 00:00:00 2001 From: David de Hilster Date: Tue, 22 Sep 2026 08:26:10 -0400 Subject: [PATCH] docs(help): the missing-words prompt named a file the engine does not write 05-missing-words.md tells you to find missing_words.txt in the analyzer's output and to check the output/ folder first. No such file exists anywhere: the engine writes missing-words.log, one per input, in the per-input log directory. Anyone following the prompt looks in the wrong place for a file under the wrong name and concludes the analyzer collected nothing. Co-Authored-By: Claude Opus 5 (1M context) --- Help/markdown/vscode/prompts/05-missing-words.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/Help/markdown/vscode/prompts/05-missing-words.md b/Help/markdown/vscode/prompts/05-missing-words.md index 642489d..ae9d9c6 100644 --- a/Help/markdown/vscode/prompts/05-missing-words.md +++ b/Help/markdown/vscode/prompts/05-missing-words.md @@ -1,12 +1,12 @@ Add missing words to the English dictionary - + I want you to extend the full English dictionary from the missing words an analyzer collected. Everything you need is on this machine at the paths below: - Current analyzer (its output log holds the missing-words list): {{currentAnalyzer}} - Shared language dictionaries / KBs (the English full dictionary lives here): {{languagesDir}} - NLP engine executable (to re-run and verify): {{engineExe}} -The missing-words list: find **missing_words.txt** in the analyzer's output — check the analyzer's `output/` folder first, then the per-input logs under `input/.txt_log/`. It is a plain list of words the analyzer did not find in the dictionary (one word per line; if a line carries a count or other columns, use only the word). De-duplicate and lowercase the words before processing. +The missing-words list: find **missing-words.log** in the analyzer's output — the engine writes one per input, at `input/.txt_log/missing-words.log`. It is a plain list of words the analyzer did not find in the dictionary (one word per line; if a line carries a count or other columns, use only the word). De-duplicate and lowercase the words before processing. The two files to edit (both under `{{languagesDir}}/English/`): @@ -48,11 +48,11 @@ Note: in en-full.dict the verb readings collapse to one line per (word, pos) wit Precision — this dictionary should stay clean: - Skip words that are already present (check both files first). -- Add genuine words only. If the list contains obvious noise — typos, fragments, single letters, run-together tokens, or proper nouns — do NOT pollute the dictionary with them. Collect those in a separate `missing_words_skipped.txt` (with a one-line reason each) so I can review them. +- Add genuine words only. If the list contains obvious noise — typos, fragments, single letters, run-together tokens, or proper nouns — do NOT pollute the dictionary with them. Collect those in a separate `missing-words-skipped.txt` (with a one-line reason each) so I can review them. - If a word is a plural/inflection whose base is not yet in the dictionary, add the base entry too so every `root=` resolves. When done: - Show me a summary: how many words added, broken down by part of speech, and the skipped list. -- Re-run the analyzer over its inputs with the engine executable and confirm the previously missing words are now recognized (the new missing_words.txt should be shorter). +- Re-run the analyzer over its inputs with the engine executable and confirm the previously missing words are now recognized (the new missing-words.log should be shorter). - Note: these two files are generated from `en-full-feat.dict` (the featured source of truth in the same directory). If it is present, add the same featured entry there as well so the additions survive a future regeneration.