Skip to content

fix(ncurses): don't map decoded UTF-8 characters to ncurses keycodes - #2311

Open
dnsl48 wants to merge 1 commit into
lem-project:mainfrom
dnsl48:fix/2310/ncurses-input-diacritics
Open

dnsl48 wants to merge 1 commit into
lem-project:mainfrom
dnsl48:fix/2310/ncurses-input-diacritics

Conversation

@dnsl48

@dnsl48 dnsl48 commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Typing a letter with a diacritic in the ncurses frontend could run an unrelated command instead of inserting the character. Esperanto ĉ (U+0109 = 265) was read as F1, ĝ (285) as Shift-F9, Czech č (269) as F5, and š (353) as Shift-Tab.

"lem-ncurses/input:get-key" decoded a multibyte sequence into a character and then handed that character to char-to-key, which looks up (char-code char) in keycode-table. That table is keyed on raw wgetch values -- ASCII plus the ncurses KEY_ constants 258-567 -- so any decoded code point in that range matched a function key. 59 Latin Extended letters collide, which is why some diacritics worked and others did not.

A multibyte sequence is always a character and never a keycode, so that path now builds the key directly with make-key and skips the table. Single-byte input still goes through char-to-key, so real function keys are unaffected: utf8-bytes reports >1 only for lead bytes 0xC2-0xF4, and keycode-table holds no entry in that range, so no legitimate lookup is lost.

The decode path moves verbatim into "lem-ncurses/input:read-multibyte-char", with two changes:

  • octets-to-string passes :encoding :utf-8 explicitly instead of depending on babel:default-character-encoding.
  • the handler widens from invalid-utf8-continuation-byte to its parent character-decoding-error. The narrow handler let overlong-utf8-sequence, character-out-of-range and end-of-input-in-character escape into input-loop, which handles neither, killing the editor on malformed input such as a pasted CESU-8 surrogate.

Addresses #2310

Typing a letter with a diacritic in the ncurses frontend could run an
unrelated command instead of inserting the character. Esperanto ĉ
(U+0109 = 265) was read as F1, ĝ (285) as Shift-F9, Czech č (269) as
F5, and š (353) as Shift-Tab.

"lem-ncurses/input:get-key" decoded a multibyte sequence into a
character and then handed that character to char-to-key,
which looks up (char-code char) in *keycode-table*.
That table is keyed on raw wgetch values -- ASCII plus the
ncurses KEY_ constants 258-567 -- so any decoded code point
in that range matched a function key. 59 Latin Extended letters
collide, which is why some diacritics worked and others did not.

A multibyte sequence is always a character and never a keycode, so
that path now builds the key directly with make-key and skips the
table. Single-byte input still goes through char-to-key, so real
function keys are unaffected: utf8-bytes reports >1 only for lead
bytes 0xC2-0xF4, and *keycode-table* holds no entry in that range, so
no legitimate lookup is lost.

The decode path moves verbatim into "lem-ncurses/input:read-multibyte-char",
with two changes:

- octets-to-string passes :encoding :utf-8 explicitly instead of
  depending on babel:*default-character-encoding*.
- the handler widens from invalid-utf8-continuation-byte to its
  parent character-decoding-error. The narrow handler let
  overlong-utf8-sequence, character-out-of-range and
  end-of-input-in-character escape into input-loop, which handles
  neither, killing the editor on malformed input such as a pasted
  CESU-8 surrogate.
@code-contractor-app

code-contractor-app Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

✅ Code Contractor Validation: PASSED

📌 Result for commit 32e90a8

=== Contract: contract ===

✓ Code Contractor Validation Result
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

📋 Contract Source: Repository

📊 Statistics:
  Files Changed:    1
  Lines Added:      18
  Lines Deleted:    13
  Total Changed:    31
  Delete Ratio:     0.42 (42%)

Status: PASSED ✅

🤖 AI Providers:
  - codex — model: (Codex default)

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🎉 No violations detected. Great job!
📋 Contract Configuration: contract (Source: Repository)
version: 2

trigger:
  paths:
    - "extensions/**"
    - "frontends/**/*.lisp"
    - "src/**"
    - "tests/**"
    - "contrib/**"
    - "**/*.asd"
  head_branches:
    exclude:
      - 'revert-*'

validation:
  limits:
    max_total_changed_lines: 400
    max_delete_ratio: 0.5
    max_files_changed: 10
    severity: warning

  ai:
    system_prompt: |
      You are a senior Common Lisp engineer reviewing code for Lem editor.
      Lem is a text editor with multiple frontends (ncurses, SDL2, webview).
      Focus on maintainability, consistency with existing code, and Lem-specific conventions.
    rules:
      # === File Structure ===
      - name: defpackage_rule
        prompt: |
          First form must be `defpackage` or `uiop:define-package`.
          Package name should match filename (e.g., `foo.lisp` → `:lem-ext/foo` or `:lem-foo`).
          Extensions must use `lem-` prefix (e.g., `:lem-python-mode`).

      - name: file_structure_rule
        prompt: |
          File organization (top to bottom):
          1. defpackage
          2. defvar/defparameter declarations
          3. Key bindings (define-key, define-keys)
          4. Class/struct definitions
          5. Functions and commands

          Only flag violations when elements appear OUT OF ORDER (e.g., key bindings AFTER functions, or defvar AFTER functions).
          Key bindings appearing BEFORE functions is correct and expected.

      # === Style ===
      - name: loop_keywords_rule
        prompt: |
          Loop keywords must use colons: `(loop :for x :in list :do ...)`
          NOT: `(loop for x in list do ...)`

      - name: naming_conventions_rule
        prompt: |
          Naming conventions:
          - Functions/variables: kebab-case (e.g., `find-buffer`)
          - Special variables: *earmuffs* (e.g., `*global-keymap*`)
          - Constants: +plus-signs+ (e.g., `+default-tab-size+`)
          - Predicates: -p suffix for functions (e.g., `buffer-modified-p`)
          - Do NOT use -p suffix for user-configurable variables

      # === Documentation ===
      - name: docstring_rule
        prompt: |
          Required docstrings for:
          - Exported functions, methods, classes
          - `define-command` (explain what the command does)
          - Generic functions (`:documentation` option)
          Important functions should explain "why", not just "what".
        severity: warning

      # === Lem-Specific ===
      - name: internal_symbol_rule
        prompt: |
          Use exported symbols from `lem` or `lem-core` package.
          Avoid `lem::internal-symbol` access.
          If internal access is necessary, document why.

      - name: error_handling_rule
        prompt: |
          - `error`: Internal/programming errors
          - `editor-error`: User-facing errors (displayed in echo area)
          Always use `editor-error` for messages shown to users.

      - name: frontend_interface_rule
        prompt: |
          Frontend-specific code must use `lem-if:*` protocol.
          Do not call frontend implementation directly from core.
        severity: warning

      # === Functional Style ===
      - name: functional_style_rule
        prompt: |
          Prefer explicit function arguments over dynamic variables.
          Avoid using `defvar` for state passed between functions.
          Exception: Well-documented cases like `*current-buffer*`.

      - name: dynamic_symbol_call_rule
        prompt: |
          Avoid `uiop:symbol-call`. Rethink architecture instead.
          If unavoidable, document the reason.

      # === Libraries ===
      - name: alexandria_usage_rule
        prompt: |
          Alexandria utilities allowed: `if-let`, `when-let`, `with-gensyms`, etc.
          Avoid: `alexandria:curry` (use explicit lambdas)
          Avoid: `alexandria-2:*` functions not yet used in codebase

      # === Macros ===
      - name: macro_style_rule
        prompt: |
          Keep macros small. For complex logic, use `call-with-*` pattern:
          ```lisp
          (defmacro with-foo (() &body body)
            `(call-with-foo (lambda () ,@body)))
          ```
          Prefer `list` over backquote outside macros.
📚 About Code Contractor

Declarative Code Standards That Learn and Improve

Define domain-specific validation rules in YAML.
Your contracts document team knowledge and evolve into more accurate AI enforcement.

Want this for your repo?
→ Install Code Contractor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant