Repository navigation
Use lm-eval 0.4.12's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py #2066
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
7 commits
Select commit
Hold shift + click to select a range
b68114b
Use lm-eval's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py
cjluo-nv dcedd37
Move the trtllm logprob fix out of lm_eval_hf.py, drop the legacy shim
cjluo-nv 622b97c
Unit-test the trtllm logprob override and guard against upstream drift
cjluo-nv 196c326
Pin --trust_remote_code propagation for the trtllm entry point
cjluo-nv d5f868c
Skip the trtllm tests on the backend module, not just lm_eval
cjluo-nv 6a54747
Require TensorRT-LLM >= 1.3.0rc11 and fix the tests on datasets 4.x
cjluo-nv 482a84e
Document the TensorRT-LLM version requirement only in llm_eval
cjluo-nv File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file was deleted.
Oops, something went wrong.
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
max_gen_toks=$BUILD_MAX_OUTPUT_LENis carried over from the old backend, but the PR body states the upstreamTRTLLM.__init__only forwards a fixed set of kwargs and silently drops the rest. Ismax_gen_toksin that set? If it is dropped, generation length is no longer capped by--output, which is exactly whattest_qwen3_eval_fp8relies on (output=128"Cap generation length: gsm8k/humaneval otherwise generate up to 1024 tokens/sample") — the test would get slower and the updated comment would be misleading. Worth confirming, and dropping the arg if it's a no-op.There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Checked —
max_gen_toksis honored, not dropped, so--outputstill caps generation andtest_qwen3_eval_fp8's comment stays accurate.It is an explicit named parameter of
TRTLLM.__init__, not part of the**kwargsthat get discarded (lm_eval/models/trtllm_causallms.py):The PR body caused the confusion and I have fixed it: the "only forwards a fixed set" claim is about the kwargs handed to the TensorRT-LLM
LLMAPI (tensor_parallel_size,max_input_len,kv_cache_config, ...), where engine-level extras really are silently dropped. lm-eval's own named parameters are consumed by the backend normally. Keeping the argument.