Skip to content

docs(research): GPTQ max-exactness fix -- partial, closes ~35% of the RTN gap - #554

Merged
gHashTag merged 1 commit into
mainfrom
research/gptq-maxexact-fix-2026-08-11
Oct 2, 2026
Merged

gHashTag merged 1 commit into
mainfrom
research/gptq-maxexact-fix-2026-08-11

Conversation

@gHashTag

Copy link
Copy Markdown
Owner

Runs the fix proposed but not executed in THEOREM_2026-08-09.md: protect each group's per-row maximal element from GPTQ's sequential compensation (equivalent to processing it first, since an exactly-quantised column propagates zero error).

  • max-exactness restored: 95.3% -> 0.4% blocks with an inexact row-max
  • perplexity: GPTQ-original 17.7047 -> GPTQ-fixed 17.5860 (-0.1186, same run)
  • still worse than RTN 4-bit (17.3667) by +0.2194 -- gptq_gate.py's pass condition remains unmet
  • reproduction note: this environment's GPTQ-original (17.7047) differs from THEOREM's (17.7846) by 0.08; all comparisons here are within-run to control for that

Conclusion: max-exactness destruction explains part (~35%) of GPTQ's gap to RTN under block-max scaling, not all of it. Residual mechanism unidentified -- open question, not run further (5-bit / promote-only / cross-model out of scope for this PR).

Adds research/block/gptq_maxexact_fix.py, research/block/GPTQ_MAXEXACT_FIX_2026-08-11.md.

… RTN gap, does not clear the gate

Runs the fix proposed but not executed in THEOREM_2026-08-09.md: protect each group's
per-row maximal element from GPTQ's sequential compensation (equivalent to processing it
first, since an exactly-quantised column propagates zero error).

- max-exactness restored: 95.3% -> 0.4% blocks with an inexact row-max
- perplexity: GPTQ-original 17.7047 -> GPTQ-fixed 17.5860 (-0.1186, same run)
- still worse than RTN 4-bit (17.3667) by +0.2194 -- gptq_gate.py's pass condition remains unmet
- reproduction note: this environment's GPTQ-original (17.7047) differs from THEOREM's
  (17.7846) by 0.08; all comparisons here are within-run to control for that

Conclusion: max-exactness destruction explains part (~35%) of GPTQ's gap to RTN under
block-max scaling, not all of it. Residual mechanism unidentified -- open question.

Adds: research/block/gptq_maxexact_fix.py, research/block/GPTQ_MAXEXACT_FIX_2026-08-11.md
@gHashTag gHashTag added the bee-reviewed A reviewer bee reviewed and verified this PR after its head commit; required to merge label Oct 2, 2026
@gHashTag

gHashTag commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Reviewer bee Z: sound, merging.

Evidence (head c072c38):

  • The PR is additive: 2 new files under research/block/, and the merge into main is clean. Nothing on main covers it: the only max-exactness text is the proposal in THEOREM_2026-08-09.md.
  • gptq_maxexact_fix.py compiles. Its imports are numpy, torch, transformers and pyarrow only, with no network, subprocess or eval calls. The hardcoded weights path follows the same convention as its siblings (e.g. block_tnf.py).
  • The arithmetic in the doc re-checks:
    • 3,162,231/3,317,760 = 95.3%
    • 12,211/3,317,760 = 0.37%
    • gap 17.7047−17.3667 = 0.338 → 0.2193, i.e. −35.1%
  • The conclusion is stated honestly: the fix is partial, the gate is not cleared, and the residual is unexplained.

Nit, not blocking: the script's own repro_ok check (|GPTQ-orig − 17.7846| < 0.02) is false in the environment that produced these numbers (Δ 0.080). The doc discloses this and keeps every comparison within one run, which is the right control.

@gHashTag
gHashTag merged commit 5f64d21 into main Oct 2, 2026
36 of 44 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bee-reviewed A reviewer bee reviewed and verified this PR after its head commit; required to merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant