Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions Help/markdown/vscode/prompts/05-missing-words.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
Add missing words to the English dictionary
<!-- desc: Reads missing_words.txt from the analyzer's output log and, for each genuine word, adds a properly featured entry (part of speech, root, verb/noun features) to en-full.dict and en-full.kbb β€” keeping both alphabetized with identical headword sets. -->
<!-- desc: Reads missing-words.log from the analyzer's output log and, for each genuine word, adds a properly featured entry (part of speech, root, verb/noun features) to en-full.dict and en-full.kbb β€” keeping both alphabetized with identical headword sets. -->
I want you to extend the full English dictionary from the missing words an analyzer collected. Everything you need is on this machine at the paths below:

- Current analyzer (its output log holds the missing-words list): {{currentAnalyzer}}
- Shared language dictionaries / KBs (the English full dictionary lives here): {{languagesDir}}
- NLP engine executable (to re-run and verify): {{engineExe}}

The missing-words list: find **missing_words.txt** in the analyzer's output β€” check the analyzer's `output/` folder first, then the per-input logs under `input/<file>.txt_log/`. It is a plain list of words the analyzer did not find in the dictionary (one word per line; if a line carries a count or other columns, use only the word). De-duplicate and lowercase the words before processing.
The missing-words list: find **missing-words.log** in the analyzer's output β€” the engine writes one per input, at `input/<file>.txt_log/missing-words.log`. It is a plain list of words the analyzer did not find in the dictionary (one word per line; if a line carries a count or other columns, use only the word). De-duplicate and lowercase the words before processing.

The two files to edit (both under `{{languagesDir}}/English/`):

Expand Down Expand Up @@ -48,11 +48,11 @@ Note: in en-full.dict the verb readings collapse to one line per (word, pos) wit
Precision β€” this dictionary should stay clean:

- Skip words that are already present (check both files first).
- Add genuine words only. If the list contains obvious noise β€” typos, fragments, single letters, run-together tokens, or proper nouns β€” do NOT pollute the dictionary with them. Collect those in a separate `missing_words_skipped.txt` (with a one-line reason each) so I can review them.
- Add genuine words only. If the list contains obvious noise β€” typos, fragments, single letters, run-together tokens, or proper nouns β€” do NOT pollute the dictionary with them. Collect those in a separate `missing-words-skipped.txt` (with a one-line reason each) so I can review them.
- If a word is a plural/inflection whose base is not yet in the dictionary, add the base entry too so every `root=` resolves.

When done:

- Show me a summary: how many words added, broken down by part of speech, and the skipped list.
- Re-run the analyzer over its inputs with the engine executable and confirm the previously missing words are now recognized (the new missing_words.txt should be shorter).
- Re-run the analyzer over its inputs with the engine executable and confirm the previously missing words are now recognized (the new missing-words.log should be shorter).
- Note: these two files are generated from `en-full-feat.dict` (the featured source of truth in the same directory). If it is present, add the same featured entry there as well so the additions survive a future regeneration.
Loading