Repository navigation
Check IQ CUDA encoder ties per block instead of by count - #2708
LinCanNerd wants to merge 1 commit into
Conversation
test_cuda_pack_reconstruction_matches_pytorch_at_scale[iq1_m] fails on Jetson Thor: 2 of 1024 blocks differ from the PyTorch encoder, over the blocks // 1000 = 1 bound. Both are exact ties: the differing encodings reconstruct their block with bit-identical squared error. How often such near-ties round differently depends on how the compiler fuses the CUDA encoder's arithmetic for the target GPU, so the count bound encodes one GPU's rate rather than the property the test is after. The test now requires every differing block to reconstruct as well as the reference, keeps the total-error check, and loosens the count bound to 1%. Flipping a grid-index bit in one block still fails it. Signed-off-by: LinCanNerd <lincanecdl@gmail.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review. 📝 WalkthroughWalkthroughThe CUDA IQ-format parity test now compares reconstruction squared error per 256-value block, allows more differing packed blocks, and checks per-block and total errors against PyTorch. ChangesCUDA IQ reconstruction parity
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix · Severity of issue fixed: Low Suggested reviewers: Merge Risk: ⚪ Minimal · up to This test-only change checks the intended reconstruction-error parity, with no identified merge-blocking risk. 🚥 Pre-merge checks | ✅ 6✅ Passed checks (6 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
What does this PR do?
Type of change: Bug fix (tests)
Fixes #2705.
test_cuda_pack_reconstruction_matches_pytorch_at_scale[iq1_m]fails on Jetson Thor because 2 of 1024 blocks differ from the PyTorch encoder, and the bound isblocks // 1000 = 1. Both blocks are exact ties: the CUDA and PyTorch encodings decode to different values but reconstruct the block with bit-identical squared error. As the docstring oftest_cuda_pack_matches_pytorch_encoder_and_is_decodableexplains, such near-ties round differently because the two encoders fuse the same arithmetic differently. How often that happens depends on how the compiler fuses the CUDA encoder's arithmetic for the target GPU, so the count bound encodes one GPU's tie rate. The property the test is after ("may disagree on a near-tied scale, but not on quality") is not what it checks.The test now checks that property directly:
rtol=1e-5).The neighbouring docstring now says the tie rate depends on the GPU.
Usage
No API change.
Testing
On Jetson Thor (sm_110, JetPack 7.2.1, torch 2.13.0+cu130):
pytest tests/gpu/torch/quantization/test_iq_formats_cuda.py: 61 passed. Before the change, theiq1_mcase failed.iq1_mandiq2_xs) makes the updated test fail in all four cases. It still catches a real quality difference.pre-commit run --files tests/gpu/torch/quantization/test_iq_formats_cuda.py: passed.Before your PR is "Ready for review"
Make sure you read and follow Contributor guidelines and your commits are signed (
git commit -s -S).Make sure you read and follow the Security Best Practices (e.g. avoiding hardcoded
trust_remote_code=True,torch.load(..., weights_only=False),pickle, etc.).CONTRIBUTING.md: N/AAdditional Information
Found while running
tests/gpu/torch/quantizationandtests/gpu/torch/kernelson Jetson Thor. The only other Thor-specific failure is #2706, which has its own PR. #2686 also edits this test file, but only appends tests after line 226, so the two changes don't overlap.Summary by CodeRabbit