Skip to content

Replace Greedy with Tenary Search for SmoothBSE - #2419

Merged
Qubitium merged 2 commits into
ModelCloud:mainfrom
namgyu-youn:smoothmse-search-perf
Feb 15, 2026
Merged

Qubitium merged 2 commits into
ModelCloud:mainfrom
namgyu-youn:smoothmse-search-perf

Conversation

@namgyu-youn

@namgyu-youn namgyu-youn commented Feb 14, 2026 •

Copy link
Copy Markdown
Contributor

Overview:
For faster convergence, replace greedy with ternary search. Ternary search only needs 38 min for Qwen3-8B, wheras 65 min in greedy search, achieving 1.97x speedup.

Benchmark Result:
Fork (this PR, 38 min); https://huggingface.co/namgyu-youn/Qwen3-8B-tenary

Perplexity (ppl; accuracy):

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value |   |Stderr|
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  |0.8535|±  |0.0156|
|     |       |strict-match    |     5|exact_match|↑  |0.6270|±  |0.0214|

Throughput: 2.31 requests/s, 2659.53 total tokens/s, 295.50 output tokens/s

Upstream (65 min; https://huggingface.co/namgyu-youn/Qwen3-8B-greedy

Perplexity (ppl; accuracy):

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value |   |Stderr|
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  |0.8535|±  |0.0156|
|     |       |strict-match    |     5|exact_match|↑  |0.6270|±  |0.0214|

Throughput: 2.31 requests/s, 2659.84 total tokens/s, 295.54 output tokens/s

@namgyu-youn

namgyu-youn commented Feb 14, 2026 •

Copy link
Copy Markdown
Contributor Author

repro:

import torch
import time
from gptqmodel import GPTQModel
from gptqmodel.quantization import (
    QuantizeConfig, FORMAT, METHOD, HessianConfig,
    FailSafe, FailSafeStrategy, SmoothMSE
)
from gptqmodel.quantization.config import VramStrategy
from datasets import load_dataset
from utils import prepare_calibration_data

BASE_MODEL = "Qwen/Qwen3-8B-Base"
HF_REPO_QUANT = "namgyu-youn/Qwen3-8B-tenary"
BITS = 4
GROUP_SIZE = 128


def main():
    start_time = time.time()

    # Calibration
    calibration_data = prepare_calibration_data(num_samples=512)

    quantize_config = QuantizeConfig(
        bits=BITS,
        group_size=GROUP_SIZE,
        quant_method=METHOD.GPTQ,
        format=FORMAT.GPTQ,
        sym=True,
        desc_act=False,
        act_group_aware=True,
        mse=2.4,
        damp_percent=0.01,
        damp_auto_increment=0.005,

        failsafe=FailSafe(
            strategy=FailSafeStrategy.MEDIAN,
            threshold="100%",  # NOTE: Forced trigger for testing
            smooth=SmoothMSE(steps=64, maxshrink=0.75, group_size_threshold=GROUP_SIZE)
        ),

        hessian=HessianConfig(
            chunk_size=None,
            chunk_bytes=512*1024*1024,
            staging_dtype=torch.float16
        ),
        offload_to_disk=False,
        vram_strategy=VramStrategy.EXCLUSIVE,
    )

    # Quantize
    model = GPTQModel.from_pretrained(
        BASE_MODEL,
        quantize_config=quantize_config,
        trust_remote_code=True,
        dtype=torch.float16
    )
    model.quantize(calibration_data, batch_size=1)

    # Save or upload model here
    print(f"Completed in {int((time.time() - start_time) / 60)} minutes")


if __name__ == "__main__":
    main()

@Qubitium
Qubitium merged commit 51daec1 into ModelCloud:main Feb 15, 2026
@Qubitium

Copy link
Copy Markdown
Collaborator

@namgyu-youn LGTM. Thanks for the 2x speedup optimization!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants