Skip to content

ADFA-5188 | Enable KV cache quantization and flash attention - #76

Open
jatezzz wants to merge 1 commit into
fix/ADFA-5187-dynamic-n-ctxfrom
fix/ADFA-5188-kv-cache-quantization
Open

ADFA-5188 | Enable KV cache quantization and flash attention#76
jatezzz wants to merge 1 commit into
fix/ADFA-5187-dynamic-n-ctxfrom
fix/ADFA-5188-kv-cache-quantization

Conversation

@jatezzz

@jatezzz jatezzz commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

Description

This PR implements KV cache quantization (q8_0) and enables flash attention in the llama.cpp context parameters to drastically improve memory efficiency and generation speed.

  • What: Configures the context parameters type_k and type_v to use GGML_TYPE_Q8_0 and enables flash attention (LLAMA_FLASH_ATTN_TYPE_AUTO).
  • How: These settings are exposed through the configure-before-load pattern. If the native context creation fails (e.g., the model's head width is incompatible with the quantized block size), it gracefully falls back to the previous defaults: f16 KV cache and flash attention disabled.
  • Why: Storing the KV cache as q8_0 halves its byte size, allowing for much longer conversations before context is dropped and drastically reducing mid-generation crashes on 4–6 GB RAM devices. Flash attention mitigates generation latency on longer contexts.

Details

Logs confirming n_ctx initialization, cache type used (q8_0 vs f16), and fallback activations.

Flash Attention Enabled

Screenshot 2026-08-21 at 12 16 48 PM

Q8_0 KV cache

Screenshot 2026-08-21 at 1 01 13 PM

Ticket

ADFA-5188

Observation

This implementation works in tandem with dynamic n_ctx sizing and should be validated alongside ADFA-5187, as the memory measurement relies on both features working concurrently.

#75 Needs to be merged first

The cache was always f16. A load now asks for q8_0 where the model's head widths divide into whole blocks, costing 34 bytes per 32 elements instead of 64 and so buying about 1.88x the context from the same RAM budget. Flash attention is requested as AUTO, since llama.cpp needs it for a quantized value cache; if a context is refused anyway, the native side retries at f16 at a shorter length.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@jatezzz jatezzz changed the title feat(ai-agent-local): store the KV cache as q8_0 to double the context ADFA-5188 | Enable KV cache quantization and flash attention Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant