Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 14 additions & 6 deletions .agents/skills/deployment/references/trtllm.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,14 +69,22 @@ If you encounter a legacy checkpoint (no `hf_quant_config.json`, has `rank*.safe

## Evaluation with TRT-LLM

```python
# examples/llm_eval/lm_eval_tensorrt_llm.py
# Runs lm_evaluation_harness benchmarks with TRT-LLM
python examples/llm_eval/lm_eval_tensorrt_llm.py \
--model_path <checkpoint_path> \
--tasks gsm8k,mmlu
Runs lm-evaluation-harness benchmarks through lm-eval's built-in `trtllm` backend
(requires `lm_eval>=0.4.12`; ModelOpt's own `lm_eval_tensorrt_llm.py` has been removed).
`lm_eval_trtllm.py` is a thin wrapper that corrects the backend's `prompt_logprobs`
alignment — without it every loglikelihood task raises `KeyError`.

```bash
python examples/llm_eval/lm_eval_trtllm.py \
--model trtllm \
--model_args model=<checkpoint_path>,tokenizer=<checkpoint_path>,tensor_parallel_size=<tp>,max_batch_size=<bs>,max_input_len=4096,max_output_len=512 \
--tasks gsm8k,mmlu \
--batch_size <bs>
```

`max_input_len` defaults to 2048 and longer prompts are silently truncated, so set it
explicitly for few-shot tasks.

## Common Issues

| Issue | Fix |
Expand Down
2 changes: 2 additions & 0 deletions CHANGELOG.rst
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,8 @@ Changelog

**Deprecations**

- Remove ``examples/llm_eval/lm_eval_tensorrt_llm.py`` (the ``trt-llm`` model). lm-evaluation-harness 0.4.12 ships its own TensorRT-LLM backend (``lm_eval.models.trtllm_causallms``, registered as ``trtllm``), which also implements ``loglikelihood_rolling`` and pipeline parallelism, so the example uses that instead and pins ``lm_eval>=0.4.12,<0.5``. Replace ``python lm_eval_tensorrt_llm.py --model trt-llm --model_args tokenizer=<tok>,checkpoint_dir=<ckpt>`` with ``python lm_eval_trtllm.py --model trtllm --model_args model=<ckpt>,tokenizer=<tok>``, and set ``tensor_parallel_size`` and ``max_input_len`` explicitly — they default to 1 and 2048, and longer prompts are silently truncated. ``lm_eval_trtllm.py`` is a thin entry point that overrides the backend's ``_parse_logprobs``, which is off by one against TensorRT-LLM's next-token-aligned ``prompt_logprobs`` and otherwise raises ``KeyError`` on every loglikelihood task; use it rather than the plain ``lm_eval`` CLI until that is fixed upstream. Loglikelihood tasks additionally require **TensorRT-LLM >= 1.3.0rc11**, the release that started passing prompt token ids into ``compute_logprobs`` so the scored token is present in every entry; older releases return only the top-1 token per position and the run aborts with an explicit error instead of a confusing lookup failure. Generative tasks are unaffected. ``examples/hf_ptq/scripts/huggingface_example.sh`` gains ``--input`` (``BUILD_MAX_INPUT_LEN``, default 4096) to size the evaluation engine's context, and honours a preset ``LM_EVAL_TP`` to override the tensor-parallel size.

**Bug Fixes**

0.46 (2026-08-xx)
Expand Down
17 changes: 12 additions & 5 deletions examples/hf_ptq/scripts/huggingface_example.sh
Original file line number Diff line number Diff line change
Expand Up @@ -297,11 +297,18 @@ if [[ $TASKS =~ "lm_eval" ]]; then

pip install -r requirements.txt

echo "Using the following config: max output $BUILD_MAX_OUTPUT_LEN max batch $BUILD_MAX_BATCH_SIZE"

python lm_eval_tensorrt_llm.py \
--model trt-llm \
--model_args tokenizer=$MODEL_PATH,checkpoint_dir=$SAVE_PATH,max_gen_toks=$BUILD_MAX_OUTPUT_LEN \
# lm-eval's `trtllm` backend defaults to 1 GPU; shard over every visible one instead.
# Override LM_EVAL_TP to lower it -- TRT-LLM enables expert parallelism at higher TP,
# which fails in DeepEP kernels for MoE checkpoints on some GPUs (e.g. SM 12.0).
LM_EVAL_TP=${LM_EVAL_TP:-$(python -c "import torch; print(max(torch.cuda.device_count(), 1))")}

echo "Using the following config: max input $BUILD_MAX_INPUT_LEN max output $BUILD_MAX_OUTPUT_LEN max batch $BUILD_MAX_BATCH_SIZE tp $LM_EVAL_TP"

# max_input_len defaults to 2048, which silently truncates 5-shot prompts, so pass it
# explicitly; the engine's max_seq_len is max_input_len + max_output_len.
python lm_eval_trtllm.py \
--model trtllm \
--model_args "model=$SAVE_PATH,tokenizer=$MODEL_ABS_PATH,tensor_parallel_size=$LM_EVAL_TP,max_batch_size=$BUILD_MAX_BATCH_SIZE,max_gen_toks=$BUILD_MAX_OUTPUT_LEN,max_input_len=$BUILD_MAX_INPUT_LEN,max_output_len=$BUILD_MAX_OUTPUT_LEN" \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot comment.

max_gen_toks=$BUILD_MAX_OUTPUT_LEN is carried over from the old backend, but the PR body states the upstream TRTLLM.__init__ only forwards a fixed set of kwargs and silently drops the rest. Is max_gen_toks in that set? If it is dropped, generation length is no longer capped by --output, which is exactly what test_qwen3_eval_fp8 relies on (output=128 "Cap generation length: gsm8k/humaneval otherwise generate up to 1024 tokens/sample") — the test would get slower and the updated comment would be misleading. Worth confirming, and dropping the arg if it's a no-op.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checked — max_gen_toks is honored, not dropped, so --output still caps generation and test_qwen3_eval_fp8's comment stays accurate.

It is an explicit named parameter of TRTLLM.__init__, not part of the **kwargs that get discarded (lm_eval/models/trtllm_causallms.py):

:60   max_gen_toks: int = 256,           # named __init__ parameter
:86   self._max_gen_toks = max_gen_toks
:575  def max_gen_toks(self) -> int: return self._max_gen_toks
:411  kwargs, until, max_gen_toks = self.modify_gen_kwargs(
:412      gen_kwargs, eos=eos, default_max_gen_toks=self.max_gen_toks)
:427  SamplingParams(max_tokens=max_gen_toks, stop=until, **kwargs)

The PR body caused the confusion and I have fixed it: the "only forwards a fixed set" claim is about the kwargs handed to the TensorRT-LLM LLM API (tensor_parallel_size, max_input_len, kv_cache_config, ...), where engine-level extras really are silently dropped. lm-eval's own named parameters are consumed by the backend normally. Keeping the argument.

--tasks $LM_EVAL_TASKS \
--batch_size $BUILD_MAX_BATCH_SIZE $lm_eval_flags | tee $LM_EVAL_RESULT

Expand Down
7 changes: 6 additions & 1 deletion examples/hf_ptq/scripts/parser.sh
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ parse_options() {
CALIB_WITH_IMAGES=false

# Parse command-line options
ARGS=$(getopt -o "" -l "model:,quant:,recipe:,kv_cache_quant:,tp:,pp:,sparsity:,awq_block_size:,calib:,calib_batch_size:,output:,batch:,tasks:,lm_eval_tasks:,lm_eval_limit:,simple_eval_tasks:,simple_eval_limit:,mmlu_limit:,trust_remote_code,use_seq_device_map,gpu_max_mem_percentage:,kv_cache_free_gpu_memory_fraction:,low_memory_mode,no-verbose,calib_dataset:,calib_seq:,auto_quantize_checkpoint:,auto_quantize_bits:,auto_quantize_method:,auto_quantize_score_size:,auto_quantize_cost_model:,auto_quantize_active_moe_expert_ratio:,moe_calib_experts_ratio:,cast_mxfp4_to_nvfp4,vlm,calib_with_images" -n "$0" -- "$@")
ARGS=$(getopt -o "" -l "model:,quant:,recipe:,kv_cache_quant:,tp:,pp:,sparsity:,awq_block_size:,calib:,calib_batch_size:,input:,output:,batch:,tasks:,lm_eval_tasks:,lm_eval_limit:,simple_eval_tasks:,simple_eval_limit:,mmlu_limit:,trust_remote_code,use_seq_device_map,gpu_max_mem_percentage:,kv_cache_free_gpu_memory_fraction:,low_memory_mode,no-verbose,calib_dataset:,calib_seq:,auto_quantize_checkpoint:,auto_quantize_bits:,auto_quantize_method:,auto_quantize_score_size:,auto_quantize_cost_model:,auto_quantize_active_moe_expert_ratio:,moe_calib_experts_ratio:,cast_mxfp4_to_nvfp4,vlm,calib_with_images" -n "$0" -- "$@")

eval set -- "$ARGS"
while true; do
Expand All @@ -56,6 +56,7 @@ parse_options() {
--awq_block_size ) AWQ_BLOCK_SIZE="$2"; shift 2;;
--calib ) CALIB_SIZE="$2"; shift 2;;
--calib_batch_size ) CALIB_BATCH_SIZE="$2"; shift 2;;
--input ) BUILD_MAX_INPUT_LEN="$2"; shift 2;;
--output ) BUILD_MAX_OUTPUT_LEN="$2"; shift 2;;
--batch ) BUILD_MAX_BATCH_SIZE="$2"; shift 2;;
--tasks ) TASKS="$2"; shift 2;;
Expand Down Expand Up @@ -90,6 +91,7 @@ parse_options() {
DEFAULT_CALIB_SIZE=512
DEFAULT_CALIB_SEQ=512
DEFAULT_CALIB_BATCH_SIZE=0
DEFAULT_BUILD_MAX_INPUT_LEN=4096
DEFAULT_BUILD_MAX_OUTPUT_LEN=1024
DEFAULT_BUILD_MAX_BATCH_SIZE=2

Expand All @@ -102,6 +104,9 @@ parse_options() {
if [ -z "$CALIB_BATCH_SIZE" ]; then
CALIB_BATCH_SIZE=$DEFAULT_CALIB_BATCH_SIZE
fi
if [ -z "$BUILD_MAX_INPUT_LEN" ]; then
BUILD_MAX_INPUT_LEN=$DEFAULT_BUILD_MAX_INPUT_LEN
fi
if [ -z "$BUILD_MAX_OUTPUT_LEN" ]; then
BUILD_MAX_OUTPUT_LEN=$DEFAULT_BUILD_MAX_OUTPUT_LEN
fi
Expand Down
34 changes: 33 additions & 1 deletion examples/llm_eval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,10 +109,42 @@ If `trust_remote_code` needs to be true, please append the command with the `--t

### TensorRT-LLM

Uses the `trtllm` backend built into lm-eval (>= 0.4.12), which loads the quantized
checkpoint directly with the TensorRT-LLM LLM API.

```sh
python lm_eval_tensorrt_llm.py --model trt-llm --model_args tokenizer=<HF model folder>,checkpoint_dir=<Quantized checkpoint dir> --tasks <comma separated tasks> --batch_size <max batch size>
python lm_eval_trtllm.py --model trtllm \
--model_args model=<Quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<max batch size>,max_input_len=4096,max_output_len=512 \
--tasks <comma separated tasks> \
--batch_size <max batch size>
```

> **_NOTE:_** Loglikelihood tasks (mmlu, hellaswag, arc, ...) need **TensorRT-LLM >=
> 1.3.0rc11**, which is when the engine started returning the requested token in every
> `prompt_logprobs` entry. Earlier releases return only the top-1 token per position, so a
> continuation token's logprob cannot be recovered and the run aborts with a clear error.
> Generative tasks (gsm8k, ifeval) are unaffected.

> **_NOTE:_** Set `max_input_len` and `max_output_len` explicitly. They default to 2048 and
> 512, and prompts longer than `max_input_len` are silently truncated — 5-shot MMLU or
> gsm8k prompts exceed 2048 tokens. `max_seq_len` of the engine is their sum.

> **_NOTE:_** `tensor_parallel_size` defaults to 1; set it to the number of GPUs the
> checkpoint needs. `pipeline_parallel_size` is also supported.

> **_NOTE:_** Use `lm_eval_trtllm.py` rather than the plain `lm_eval` CLI. lm-eval 0.4.12's
> `trtllm` backend misaligns TensorRT-LLM's `prompt_logprobs` by one position, so every
> loglikelihood task (hellaswag, mmlu, arc, ...) fails with a `KeyError`;
> `lm_eval_trtllm.py` overrides the alignment. It goes away once the fix lands upstream.

> **_NOTE:_** The backend forwards only a fixed set of arguments to TensorRT-LLM, so the
> tuning the old `lm_eval_tensorrt_llm.py` applied is not reachable: expert parallelism is
> left at the TensorRT-LLM default (MoE checkpoints can fail in DeepEP kernels on some
> GPUs, e.g. SM 12.0) and the KV cache uses 90% of free GPU memory rather than 70%. Lower
> `tensor_parallel_size` if you hit either.

`lm_eval_tensorrt_llm.py` (`--model trt-llm`) has been removed; use the command above.

## MMLU

[Massive Multitask Language Understanding](https://arxiv.org/abs/2009.03300). A score (0-1, higher is better) will be printed at the end of the benchmark.
Expand Down
5 changes: 3 additions & 2 deletions examples/llm_eval/lm_eval_hf.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,8 +48,9 @@
from lm_eval import utils
from packaging.version import Version

if Version(version("lm_eval")) < Version("0.4.10"):
raise ImportError(f"lm_eval_hf.py requires lm-eval >= 0.4.10; found {version('lm_eval')}.")
if Version(version("lm_eval")) < Version("0.4.12"):
# Matches the floor pinned in requirements.txt.
raise ImportError(f"lm_eval_hf.py requires lm-eval >= 0.4.12; found {version('lm_eval')}.")

from lm_eval._cli import HarnessCLI
from lm_eval.api.model import T
Expand Down
213 changes: 0 additions & 213 deletions examples/llm_eval/lm_eval_tensorrt_llm.py

This file was deleted.

Loading
Loading