Skip to content

Fix ANE-LM review findings after PR #86 - #90

Merged
IchenDEV merged 7 commits into
mainfrom
fix/ane-lm-review-fixes
Aug 31, 2026
Merged

IchenDEV merged 7 commits into
mainfrom
fix/ane-lm-review-fixes

Conversation

@IchenDEV

Copy link
Copy Markdown
Owner

Summary

  • disable Qwen3 thinking in the direct ANE prompt so formatting and structured edit-command requests spend their token budget on the requested output
  • reject ANE requests that exceed the packaged runtime's 2,048-slot KV cache, reserving the slots needed by generated tokens so oversized requests enter the existing MLX fallback path instead of silently dropping the prompt prefix
  • unload the ANE runtime before MLX fallback for both generation and warmup paths
  • preserve the first fallback outcome in the environment-gated repeated-request test

Regression coverage

  • exact non-thinking Qwen3 assistant suffix
  • 2,048-slot boundary and output-headroom cases
  • ANE cleanup occurs before MLX starts, while disabled fallback performs no cleanup
  • real fallback test asserts the ANE engine is unloaded and captures .fallback before later MLX iterations clear task-local state
  • environment-gated real ANE lifecycle test rejects thinking-tag output and requires the first structured command response to parse

Follow-up to #86 and its unresolved review threads.

@IchenDEV
IchenDEV merged commit a41cb4f into main Aug 31, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant