[None][feat] Support the DMD2-distilled Cosmos3-Super-Text2Image-4Step checkpoint - #16563
Conversation
|
Could a maintainer please add the |
WalkthroughAdds checkpoint-aware Cosmos3 sampling for distilled fixed-step text-to-image generation, integrates the new checkpoint and deployment configuration, adds transformer compatibility defaults, updates scheduler denoising support, and expands unit and integration coverage. ChangesCosmos3 distilled text-to-image
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant User
participant Cosmos3OmniMoTPipeline
participant Cosmos3SamplingPolicy
participant BasePipeline
participant FlowMatchEulerDiscreteScheduler
User->>Cosmos3OmniMoTPipeline: request distilled text-to-image
Cosmos3OmniMoTPipeline->>Cosmos3SamplingPolicy: validate fixed steps and guidance
Cosmos3OmniMoTPipeline->>Cosmos3SamplingPolicy: configure timesteps and step kwargs
Cosmos3OmniMoTPipeline->>BasePipeline: start denoising
BasePipeline->>FlowMatchEulerDiscreteScheduler: step with generator
FlowMatchEulerDiscreteScheduler-->>Cosmos3OmniMoTPipeline: updated latents
Cosmos3OmniMoTPipeline-->>User: output image
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
tensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.py (1)
189-247: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winDefault distilled requests and warmup still select video mode.
For
Cosmos3-Super-Text2Image-4Step, omittedoutput_typeresolves to"video", while warmup remains 720×1280×189 and also invokes the video path. Default distilled requests should select"image", and warmup should use the T2I resolution with one frame; explicitly reject video mode if this checkpoint does not support it.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.py` around lines 189 - 247, Update Cosmos3-Super-Text2Image-4Step defaults so omitted output_type resolves to "image" rather than "video". Adjust default_warmup_resolutions and default_warmup_num_frames, and _run_warmup, to use the T2I resolution with one frame. In infer, explicitly reject video mode for this checkpoint while preserving supported image generation behavior.
🧹 Nitpick comments (1)
tensorrt_llm/_torch/visual_gen/models/cosmos3/transformer_cosmos3.py (1)
46-51: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAdd parameter and return annotations to the compatibility helper.
Use the concrete pretrained-config type accepted by
DiffusionModelConfig.As per coding guidelines, “Annotate every function.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tensorrt_llm/_torch/visual_gen/models/cosmos3/transformer_cosmos3.py` around lines 46 - 51, Annotate apply_pretrained_config_compat_defaults with the concrete pretrained-config type accepted by DiffusionModelConfig and its return type, reflecting that it mutates and returns the same configuration object. Preserve the existing idempotent default-filling behavior.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/integration/defs/examples/visual_gen/test_visual_gen.py`:
- Around line 1868-1900: Update the test around the output_path assertion in the
Cosmos3 visual generation case to remove any existing PNG before invoking
venv_check_call, then verify the newly generated artifact exists and is
non-empty. Keep the existing output path and invocation unchanged, and state
that coverage is sufficient only when the fresh file check passes.
In `@tests/unittest/_torch/visual_gen/conftest.py`:
- Around line 19-33: Annotate every new callable with complete, precise types
and avoid Any: in tests/unittest/_torch/visual_gen/conftest.py lines 19-33,
update disable_cosmos3_guardrails with a concrete iterator/generator return
type; in tests/unittest/_torch/visual_gen/test_cosmos3_distilled.py lines
57-552, annotate all helper parameters and returns and add -> None to test
methods; in tests/integration/defs/examples/visual_gen/test_visual_gen.py lines
1857-1900, annotate fixture parameters and add -> None; and in
tests/unittest/_torch/visual_gen/test_cosmos3_transformer.py lines 475-495, add
-> None to each new test method.
In `@tests/unittest/_torch/visual_gen/test_cosmos3_transformer.py`:
- Around line 481-488: Extend test_old_schema_untouched in
apply_pretrained_config_compat_defaults coverage so every explicitly provided
field uses a non-default sentinel and has its own preservation assertion,
including position_embedding_type and temporal_compression_factor_sound
alongside max_position_embeddings. Confirm TensorRT-LLM coverage is not relying
on only one sentinel, and add equivalent assertions there if its compatibility
tests cover this helper.
---
Outside diff comments:
In `@tensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.py`:
- Around line 189-247: Update Cosmos3-Super-Text2Image-4Step defaults so omitted
output_type resolves to "image" rather than "video". Adjust
default_warmup_resolutions and default_warmup_num_frames, and _run_warmup, to
use the T2I resolution with one frame. In infer, explicitly reject video mode
for this checkpoint while preserving supported image generation behavior.
---
Nitpick comments:
In `@tensorrt_llm/_torch/visual_gen/models/cosmos3/transformer_cosmos3.py`:
- Around line 46-51: Annotate apply_pretrained_config_compat_defaults with the
concrete pretrained-config type accepted by DiffusionModelConfig and its return
type, reflecting that it mutates and returns the same configuration object.
Preserve the existing idempotent default-filling behavior.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 938dea2c-547e-404e-92bc-4778696ac546
📒 Files selected for processing (17)
.gitignoredocs/source/models/supported-models.mddocs/source/models/visual-generation.mdexamples/visual_gen/configs/cosmos3-t2i-1gpu.yamlexamples/visual_gen/models/cosmos3/README.mdrequirements.txttensorrt_llm/_torch/visual_gen/models/cosmos3/defaults.pytensorrt_llm/_torch/visual_gen/models/cosmos3/pipeline_cosmos3.pytensorrt_llm/_torch/visual_gen/models/cosmos3/sampling.pytensorrt_llm/_torch/visual_gen/models/cosmos3/transformer_cosmos3.pytensorrt_llm/_torch/visual_gen/pipeline.pytests/integration/defs/examples/visual_gen/test_visual_gen.pytests/integration/test_lists/test-db/l0_b200.ymltests/unittest/_torch/visual_gen/conftest.pytests/unittest/_torch/visual_gen/test_cosmos3_distilled.pytests/unittest/_torch/visual_gen/test_cosmos3_pipeline.pytests/unittest/_torch/visual_gen/test_cosmos3_transformer.py
|
/bot run --disable-fail-fast |
|
PR_Github #60087 [ run ] triggered by Bot. Commit: |
|
Responses to the two CodeRabbit items without inline threads: Nitpick — annotate Outside-diff — make omitted |
|
PR_Github #60087 [ run ] completed with state
|
BowenFu
left a comment
There was a problem hiding this comment.
LGTM — additive DMD2-distilled Cosmos3 support; the shared pipeline.py change adds an optional scheduler_step_kwargs (default None→{}, so existing models are unchanged) and there's no public tensorrt_llm/visual_gen API change. Only cross-cutting item: the diffusers floor bump 0.37.1→0.39.0 raises the minimum for all diffusers-based VisualGen models — CI-covered, worth a heads-up.
e3ec19c to
d30be43
Compare
|
/bot run --disable-fail-fast |
Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
…xample test main replaced the eager venv_check_call import with a lazy wrapper for multiprocessing safety; the semantic merge left the new test calling the now-undefined name (F821). Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
A FlowMatchEuler scheduler with a fixed t_list but stochastic_sampling disabled previously loaded as distilled and would silently run the wrong ODE recipe with guidance forced to 1.0. The distilled combination now requires stochastic_sampling, and a declared fixed_step_sampler_config.sample_type must be 'sde'. Also documents the default-constructed policy as the explicit pre-load placeholder, and carries the test for the next commit's image-conditioning rejection alongside the new malformed-recipe tests. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
…eckpoints The stochastic distilled scheduler re-noises the conditioned frame at every step; this pipeline only restores it once before decoding, which silently produces incorrect output. Reject the request until per-step re-anchoring lands (implemented in the follow-up I2V-4Step work). Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
…contract The docstring promised all defaults resolved, but pipelines with mode-dependent defaults (Cosmos3: text-to-image and video requests use different resolutions/steps/guidance) deliberately leave those fields None until the output mode is known per request. Document that None means the mode's default rather than unset; runtime behavior is unchanged. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
The distilled e2e run takes up to 30 minutes; keep pre-merge lean and run it as a post-merge B200 canary instead. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
requirements.txt now requires diffusers>=0.39.0 (seeded FlowMatchEuler stochastic sampling); the ==0.38.0 dev pin from the LPIPS stabilization made the combined resolve unsatisfiable. The LPIPS pipelines pass no generator to scheduler.step, so the huggingface/diffusers#13678 behavior change does not reach their outputs. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
The shared conftest set TLLM_DISABLE_MPI=1 for the whole visual_gen unit test directory, but Mapping.__new__ dispatches on that variable: set means DeviceMeshTopology, unset means MpiTopology. test_flux_attention.py builds Mapping(world_size=2, tp_size=2) directly and needs the MPI topology, so the directory-wide set flipped its fused QK-norm + RoPE gate and failed test_fused_qk_norm_rope_enabled_only_for_tp1 in CI. The Cosmos3 unit tests never spawn a VisualGen executor and pass without the variable, so nothing gains it here. Keep it only where it was before this PR: module-level in the two pre-existing Cosmos3 test modules, with their teardown restored. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
14eff83 to
a9a7dac
Compare
|
Rebased onto current main ( Also worth noting from the last run: /bot run --disable-fail-fast |
|
/bot run --disable-fail-fast |
|
PR_Github #62229 [ run ] triggered by Bot. Commit: |
|
PR_Github #62229 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #62576 [ run ] triggered by Bot. Commit: |
|
/bot run --disable-fail-fast |
|
PR_Github #62694 [ run ] triggered by Bot. Commit: |
|
PR_Github #62576 [ run ] completed with state |
|
PR_Github #62694 [ run ] completed with state |
main's Cosmos3-Super-Text2Image-4Step work (NVIDIA#16563) replaced the scheduler machinery this branch extended: Cosmos3SamplingPolicy now owns scheduler construction, so the branch's _set_flow_shift / _scheduler_use_karras_sigmas helpers are gone. Port V2V onto that policy. set_flow_shift() grows an optional use_karras_sigmas so the policy stays the single owner of scheduler rebuilds -- V2V needs flow_shift=10.0 with the uniform sigma schedule, which the policy could not previously express. Both knobs default to 'whatever the checkpoint shipped'. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
NVIDIA#16563 made Cosmos3 declare None for height/width/num_inference_steps/ guidance_scale so the executor's merge could not stamp video values onto a request whose output mode was not yet known. That silenced the public VisualGen.default_params: callers saw None where Nano/Super previously reported 720p, and arithmetic on the result raised TypeError. The merge never needed the None. Pydantic already records whether a field was set by the caller, so the executor now fills its defaults and un-marks them, and Cosmos3's infer() re-resolves anything not caller-assigned against the request's own mode table. default_params reports the checkpoint's video-mode values again (Nano/Super 1280x720, Edge 832x480), and a params object whose output_type is switched to image still resolves to the text-to-image defaults instead of inheriting video ones. Distilled checkpoints keep their fixed steps and guidance in every path: the sampling policy's overrides are layered ahead of the mode table. The default_params docstring returns to its pre-NVIDIA#16563 wording, since the behavior it described is exactly what this restores. Unit coverage drives the real path end to end - default_params, _merge_defaults, deep copy, pickle, infer - so neither un-marking site can be dropped silently. Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
## Summary - add a TensorRT-LLM walkthrough for the two published DMD2-distilled Cosmos3-Super students, `nvidia/Cosmos3-Super-Text2Image-4Step` (text-to-image) and `nvidia/Cosmos3-Super-Image2Video-4Step` (image-to-video), each on a single GPU - document the distilled request contract: the checkpoint's scheduler config locks the schedule to four stochastic (SDE) steps and classifier-free guidance is baked into the weights, so requests omit `num_inference_steps` and `guidance_scale` entirely and the server fills both in; a conflicting value is rejected, not clamped - leave `use_system_prompt` unset in every request so the image-to-video student's checkpoint-declared `default_use_system_prompt: true` applies - add both `trtllm-serve` launch commands to the shared setup guide, and index the notebook from the root README - **separately**: reword two guardrail blocklist false positives in the shared audiovisual negative prompts, without which no image-to-video request can run at all (see below) This is the follow-up #310 deferred ("keep the four-step distilled T2I/I2V examples out of this PR pending separate output-quality follow-up"). The examples match the assets every other backend already uses for its distilled sections — `robot_draping` for text-to-image and `car_driving` for image-to-video — so the cookbook surface is consistent across Cosmos Framework, Diffusers, vLLM-Omni, SGLang, and now TensorRT-LLM. The shipped scope is text-to-image and image-to-video only. These students are task-specialized and no other backend advertises text-to-video, video-to-video, or synchronized audio for them either. ## The guardrail fix, and why it is in this PR The shared audiovisual negative prompts contained two substrings that the `cosmos_guardrail` 0.3.0 `Blocklist` censors. Any request carrying them fails with HTTP 500 and `Text guardrail blocked prompt` **before any denoising runs**: | Text | Censored word | Replacement | |---|---|---| | `color bleeding between elements` | `bleeding` | `color smearing between elements` | | `The scene feels lifeless and sterile` | `lifeless` | `The scene feels inert and sterile` | Neither describes anything the blocklist is meant to catch; both are ordinary image-artifact vocabulary. The replacement wording is not invented here. TensorRT-LLM already ships a corrected copy of this exact prompt as `examples/visual_gen/models/cosmos3/cosmos3_negative_prompt.json`, and that is the default every Cosmos3 request falls back to when it sends no negative prompt of its own. Upstream's copy and the cookbook's differ in **exactly four leaves, and all four are these two substrings** — so the cookbook asset is simply stale. Both files are now identical to the TensorRT-LLM default. Kept as separate commits so they can be split out if preferred. Text-to-image was unaffected because that path sends no negative prompt — which is why only the image-to-video example hit it. Three related issues found while diagnosing this, **not changed here**: 1. The two files are shared by the Diffusers, vLLM-Omni, SGLang and base TensorRT-LLM notebooks, so this also unblocks their image-to-video / text-to-video / video-to-video examples. 2. `assets/prompts/image2video/humanoid_robot.json` is blocked by a *different* rule (partial match, not the word blocklist). That is the prompt the merged `run_with_trt_llm.ipynb` uses for `i2v_nano` and `i2v_super`. Left alone deliberately: rewording a *positive* prompt changes generated content. 3. `transfer/assets/negative_prompt.json` carries the same two substrings. That file belongs to #310, which now rewords both, so it is covered there rather than duplicated here. ## Validation ### GPU inference Both notebook sections were executed **as written** with `jupyter nbconvert --to notebook --execute` against real `trtllm-serve` VisualGen servers — the shipped notebook's own cells, not a re-implementation. TensorRT-LLM at `505a02ce04`, whose `tensorrt_llm/_torch/visual_gen`, `tensorrt_llm/visual_gen` and `examples/visual_gen` trees are byte-identical to `main` (`2dbab44d4f`); the only delta is test files. Hardware: 1 x NVIDIA B300. **Text-to-image** — `Cosmos3-Super-Text2Image-4Step` with `configs/cosmos3-t2i-1gpu.yaml`: ``` Distilled Cosmos3 checkpoint: fixed 4-step schedule [1.0, 0.9375, 0.8333333333333334, 0.625], classifier-free guidance baked in. Running warmup for Cosmos3OmniMoTPipeline: 1 shape(s) [1024x1024x1], 4 steps Denoising: 4 steps, guidance=1.0 POST /v1/videos/generations HTTP/1.1" 200 OK ``` Artifact: 57,617 bytes, decoded with PyAV as H.264 1024x1024, 1 frame. **Image-to-video** — `Cosmos3-Super-Image2Video-4Step`, no config file: ``` Running warmup for Cosmos3OmniMoTPipeline: 1 shape(s) [720x1280x189], 4 steps Denoising: 4 steps, guidance=1.0 POST /v1/videos/generations HTTP/1.1" 200 OK ``` Artifact: 4,890,310 bytes, decoded as H.264 1280x720, 24 fps, 189 frames. The warmup shape confirms the claim that this checkpoint deploys at the default omni shape and needs no config file. **Negative control** (both servers): a request that sets `guidance_scale=6.0` is rejected rather than clamped — ``` HTTP 400: This is a distilled Cosmos3 checkpoint; classifier-free guidance is baked into the weights. guidance_scale must be 1.0 or left unset (got 6.0). ``` **Guardrail fix verified directly** against `cosmos_guardrail` 0.3.0's `Blocklist`: both reworded negative prompts and both positive prompts used by the notebook now report `safe=True`. The trigger words were located by diffing the censored output against the original, not guessed. Re-verified after adopting upstream's wording: `bleeding` and `lifeless` are present in NVIDIA's `blocklist`/`exact_match` word sets (383 and 1430 entries) while `smearing` and `inert` are not. **Visual review**: the text-to-image frame renders the prompt's robot-arm-draping-satin scene coherently. The image-to-video clip follows its `temporal_caption` — coastal dashcam view, rockfall beginning a few seconds in, boulders and dust building until the road is partially blocked — with no flicker or geometry snapping across the 189 frames. ### Local validation - parsed the new notebook as JSON - validated it against the nbformat schema with `nbformat==5.11.1` (the CI gate) - linted it with `ruff==0.16.4 check --select=E9,F63,F7,F82` (the CI gate) - verified unique cell IDs, empty outputs, null execution counts, and that every code cell parses - re-validated both edited JSON assets with `json.tool` - ran `git diff --check` ### Source contract audit The request contract is read from TensorRT-LLM `main` (`tensorrt_llm/_torch/visual_gen/models/cosmos3/sampling.py`): - `load_scheduler` selects `FlowMatchEulerDiscreteScheduler` for distilled checkpoints and `UniPCMultistepScheduler` for base ones; an unknown declaration is a load-time error - `Cosmos3SamplingPolicy.validate_request` rejects any `num_inference_steps` other than `len(t_list)` and any `guidance_scale` other than `1.0` - `generation_default_overrides` supplies both values when the request leaves them unset - `set_flow_shift` is a structural no-op for distilled checkpoints, so `flow_shift` does not apply - both checkpoints declare `fixed_step_sampler_config.t_list = [1.0, 0.9375, 0.8333333333333334, 0.625]` with `stochastic_sampling: true` ## Dependencies and review state - TensorRT-LLM support is already merged: NVIDIA/TensorRT-LLM#16563 (distilled text-to-image) and NVIDIA/TensorRT-LLM#16690 (distilled image-to-video). - No dependency on #310; the two PRs touch different notebooks. If #310 merges first there is no expected conflict. --------- Signed-off-by: Igor Shovkun <ishovkun@nvidia.com>
Summary by CodeRabbit
New Features
Documentation
Bug Fixes
Tests
Description
Adds support for
nvidia/Cosmos3-Super-Text2Image-4Step, a DMD2-distilled text-to-image Cosmos3 checkpoint. Unlike the base checkpoints, it samples withFlowMatchEulerDiscreteScheduleron a fixed 4-sigma stochastic (SDE) schedule declared in the checkpoint's scheduler config, with classifier-free guidance baked into the weights (one forward per step). It also ships a newer diffusers conversion of the transformer config that omits a few schema fields older conversions carried.What changed:
sampling.py): the pipeline instantiates the scheduler class the checkpoint declares — UniPC for base checkpoints (a missing declaration also resolves to UniPC, preserving existing behavior), FlowMatchEuler for distilled ones. An explicitly unknown declaration is a load-time error rather than a silent UniPC substitution.Cosmos3SamplingPolicy(sampling.py): an immutable value object holding the checkpoint's sampling facts (fixed sigmas, distilled detection, UniPC base config for flow-shift rebuilds). Only two recipes are valid — UniPC without fixed sigmas (base) and FlowMatchEuler with fixed sigmas plusstochastic_sampling=true(distilled); malformed combinations, including non-stochastic or non-SDE declarations, fail at load. Requests that conflict with a distilled checkpoint's fixed steps/guidance are rejected with a clear error, as are image-conditioned requests (correct distilled conditioning needs per-step re-anchoring, which lands with I2V-4Step support).default_generation_paramsreports the checkpoint's true steps/guidance (4 / 1.0 for distilled). Mode-dependent fields (height,width,num_inference_steps,guidance_scale) stayNoneuntilinfer()resolves the request mode (video vs. image) exactly once; explicit request values pass through unchanged.pipeline.py): the shared denoise loop now threadsscheduler_step_kwargsinto everyscheduler.step()call so the request-seededtorch.Generatordrives the stochastic step's noise. Requiresdiffusers>=0.39.0— earlier versions ignore a caller-supplied generator in the stochastic branch (Fix ignored generator in FlowMatchEulerDiscreteScheduler huggingface/diffusers#13678). Same-seed runs produce bit-identical images.transformer_cosmos3.py): newer conversions omitposition_embedding_type,max_position_embeddings, andtemporal_compression_factor_sound; these are filled with their historical values at model construction (idempotent).examples/visual_gen/configs/cosmos3-t2i-1gpu.yamlwarms up the deployed 1024×1024 single-frame shape (warmup follows the workflow, not the checkpoint name); README and model docs updated with the exact invocation.Behavior for base checkpoints (Cosmos3-Nano, Cosmos3-Super) is unchanged: same UniPC scheduler, same defaults, no new step kwargs.
Test Coverage
tests/unittest/_torch/visual_gen/test_cosmos3_distilled.py(new, 69 tests): scheduler loading from real config files, recipe-matrix validation (malformed combinations fail at load), distilled request validation, flow-shift rebuild/restore semantics, fixed-sigma timestep programming, SDE seed determinism (same seed reproduces, different seeds diverge), generation defaults,infer()mode resolution, guidance-1.0 denoise-loop contract (single forward per step, step kwargs reach everyscheduler.step), registry dispatch. Includes a canary pinning that diffusers retains unknown scheduler-config keys, which distilled detection depends on.tests/unittest/_torch/visual_gen/test_cosmos3_transformer.py: config schema compat-default tests.tests/integration/defs/examples/visual_gen/test_visual_gen.py::test_cosmos3_t2i_4step_example(inl0_b200.yml, B200 post-merge): runs the documented example invocation against the real checkpoint and asserts an image is produced.PR Checklist
Please review the following before submitting your PR:
PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.
PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.
Test cases are provided for new code paths (see test instructions)
If PR introduces API changes, an appropriate PR label is added - either
api-compatibleorapi-breaking. Forapi-breaking, includeBREAKINGin the PR title.Any new dependencies have been scanned for license and vulnerabilities
CODEOWNERS updated if ownership changes
Documentation updated as needed
Update tava architecture diagram if there is a significant design change in PR.
The reviewers assigned automatically/manually are appropriate for the PR.
Please check this after reviewing the above items as appropriate for this PR.