Repository navigation
🚨 Exporters Runtimes : Export and run any Transformers model (Dynamo/OpenVINO/OnnxRuntime/ExecuTorch) - #48986
Draft
IlyasMoutawwakil wants to merge 115 commits into
Draft
IlyasMoutawwakil wants to merge 115 commits into
IlyasMoutawwakil wants to merge 115 commits into
Conversation
Every M-RoPE model carried its own ~150-line `get_rope_index`, copied and lightly edited per family. Move the layouts into `modeling_rope_utils.get_mrope_index`, dispatched on a new `config.mrope_layout`, and reduce each model's `get_rope_index` to a one-line wrapper over it. Four layouts cover every M-RoPE family: - `interleaved_runs` — the shared Qwen2-VL-style layout (15 families). Runs come from `mm_token_type_ids`; the per-family differences are now config attributes (`spatial_merge_size`, `tokens_per_second`, `temporal_merge_size`, and a new `timestamped_video_frames` for the processors that split a video into per-frame spans). - `indexed_images` — HunYuanVL: a 1D text baseline overwritten per image span with `(width, height, image_index)` on the last three of `mrope_section`'s axes. - `audio_chunked` — Qwen2.5-Omni: spans located by scanning `input_ids` for placeholder tokens, audio laid out along the temporal axis, and audio-in-video interleaved in `seconds_per_chunk` time chunks. - `audio_merged` — Qwen3-Omni: same shape, but scanning for the opening tokens, merging audio and video position by position, and returning float positions since its temporal axis is fractional. The layouts are pure functions of their inputs, so they can be driven and tested without a model instance, and `get_mrope_index` is the single place a config is read. An unknown `mrope_layout` raises `NotImplementedError` rather than silently producing wrong positions. Also adds `get_mrope_text_positions` (the shared no-vision fallback, which is what the Omni families use for audio-only prompts), declares `vision_start_token_id` on the Omni thinker configs (it was read but never declared), and drops the now-absorbed `get_llm_pos_ids_for_vision` / `get_chunked_index` helpers — which in turn removes Qwen3-Omni's Talker delegating overrides. GLM-Image keeps its bespoke implementation: its `get_rope_index` caches decode-stage positions on the module, so it is not a pure function. Tests: `tests/utils/test_modeling_rope_utils.py::MRopeLayoutTest` covers all four layouts against positions produced by the implementations they were ported from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Multimodal models re-declare the same handful of methods almost verbatim. Counting identical bodies across the 71 modeling files that define `get_placeholder_mask`: `get_vision_position_ids` was 17 copies of one function, `compute_3d_position_ids` 14 of 17, `_prepare_position_ids_for_ generation` 14 of 16, `_get_image_nums_and_video_nums` 10 of 17. Add `modeling_multimodal_utils.py` with two mixins, one per attachment point, since these helpers live on two different classes: - `MultiModalPreTrainedModelMixin` on the base `<X>Model`: `get_vision_position_ids` (now a thin method form of `modeling_rope_utils.get_mrope_vision_positions`, which is the same function it had been copied from), `get_rope_index`, and `compute_3d_position_ids`. - `MultiModalGenerationMixin` on `<X>ForConditionalGeneration`, mixed in before `GenerationMixin` so its overrides win while `super()` still reaches the text-only behaviour: `_prepare_position_ids_for_generation` and `_get_image_nums_and_video_nums`. Mixing them into the family roots and regenerating cascades to every child, so this deletes the copies rather than relocating them. Families that genuinely differ keep their overrides: qwen2_5_vl (threads `second_per_grid_ts`), hunyuan_vl (image-only layout), the two Omni models (audio signatures), glm_image (caches decode positions on the module), exaone4_5 (BC shim), video_llama_3, and the glm4v family for `_get_image_nums_and_video_nums` (different token ids). `get_rope_index` on the mixin swallows `**kwargs` without forwarding them, matching the per-model wrappers it replaces — generation passes the whole model-kwargs dict through it. Two things a mixin cannot do that a parent class body can, both of which bit during this refactor and are handled at the mixin level rather than per model: - It cannot be *deleted* by modular's `raise AttributeError` idiom. Four models used that to drop these helpers (video_llama_3, exaone4_5, kimi_k25, muse_glimmer) and the stubs started being emitted into generated code, where `generate` would have called them. Those four are exactly the models whose configs declare no M-RoPE, so the mixin now gates on `uses_mrope(self.config)`: `compute_3d_position_ids` returns `None` and `_prepare_position_ids_for_generation` returns the text positions, which is what an absent method used to give them. The stubs are gone. - It changes what `super()` means in existing overrides. cosmos3_edge and hunyuan_vl start from the plain 2D text positions, and their `super()` call silently began returning the mixin's four-axis result instead (caught by five cosmos3_edge generate tests). Rather than pin those two call sites, `_prepare_position_ids_for_generation` is now a thin hook that hands `text_positions` to an overridable `_prepare_mrope_position_ids_for_generation`, so no override has to reason about the MRO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…to unified-mrope
The vision helpers in `vision_utils.py` take their settings as plain arguments, but several of those arguments were only available on the *module*: hardcoded in `__init__` (`self.interpolation_mode = "bilinear"`), derived there (`self.num_grid_per_side = int(config.num_position_embeddings ** 0.5)`), or passed as a literal at the call site (`merge_temporal=True`). That makes the helpers uncallable from a saved config alone — you need a built module to know how to call them. Declare all of them on the vision configs instead, and have the modules read them back, so config is the single source of truth: - fields for the hardcoded values: `interpolation_mode`, `interpolation_align_corners`, `interpolation_padding`, `merge_temporal_attention` (was a literal `True` in `get_vision_cu_seqlens`), `include_temporal_position_ids` (likewise in `get_vision_position_ids`), and `resample_before_merge`, which records that a family resamples its position grid at patch resolution — the reason kimi_k25 / muse_glimmer / paddleocr_vl pass a merge size of 1 while the qwen3_vl family passes its real one. - properties for the derived ones: `num_grid_per_side`, muse_glimmer's `window_size`, and `spatial_merge_size` on kimi_k25 / muse_glimmer, which spell it `merge_kernel_size` / `merge_size` — so every vision config now exposes that name. Five config classes cover all nine vision modules, since six of them inherit `Qwen3VLVisionConfig`. minicpmv4_6 is left alone: its `get_vision_window_index(..., spatial_merge_size=1, patch_size=1)` call passes an already-merged grid, which is a property of that call site rather than a config setting. Verified by resolving every helper argument off a bare `CONFIG_MAPPING[...]()` and running the helpers with no model or module in play, for all ten families. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ormers into unified-mrope
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`configuration_utils` imports `modeling_rope_utils` for `RotaryEmbeddingConfigMixin`, so that module has to stay importable in a torch-free install — which is why every torch annotation in it is deferred. Hosting 800+ lines of tensor code there was a standing trap: CI caught it once already when a `dtype: torch.dtype = torch.long` default was evaluated at import time. The layouts belong with the mixins that call them anyway: `modeling_multimodal_utils` imports torch unconditionally and is only reached from modeling code. Move `get_mrope_index`, the four layouts and their helpers, `get_mrope_vision_positions` / `get_mrope_image_positions` / `get_mrope_text_positions` and `uses_mrope` there, leaving `modeling_rope_utils` as the frequency machinery it was. Splits the tests the same way: `tests/utils/test_modeling_multimodal_utils.py` holds the layout tests, `test_modeling_rope_utils.py` the frequency ones. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Now that every M-RoPE layout and vision knob is declared on the config, the exporter no longer needs a model (or module introspection) to know how to call the shared helpers: - `precompute_export_inputs(config, inputs) -> dict` is config-only and no longer mutates its argument. It took a model solely for `model.get_rope_index`, which is now a thin wrapper over `get_mrope_index(config, ...)`, so it calls the layout directly. A component without a config (`embed_tokens`) simply has nothing to precompute, which the caller decides rather than the function silently accepting a module. - The grid_thw preparer reads `merge_temporal_attention` off the config instead of inferring clip-level attention from `hasattr(module, "get_vision_frame_index")`, and derives the resample merge size the same way the modules do (`1 if resample_before_merge else spatial_merge_size`). - Drop the `lm_head` component: nothing consumed it (the runtime never built a runner for it), so it was an exported graph holding a single matmul. Rename `language_model` -> `text_decoder`, which is what `get_decoder()` returns and no longer reads like the `decode` phase. Adds `test_precomputed_inputs_match_eager`: the other export tests feed the precomputed tensors to both the eager reference and the exported program, so a wrong-but-consistent value cancels out and passes. This compares the model run with precomputed inputs against the same model run without them, at exact tolerance — verified to fail when the preparer derives them at the wrong merge size. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`materialize_cache_layers` derived the KV cache geometry as `num_key_value_heads x head_dim`, which is wrong for multi-head latent attention: deepseek_v3 and kimi_k25 cache a single compressed latent head of `kv_lora_rank`. The runtime therefore handed the decode graph a `[batch, num_key_value_heads, 0, head_dim]` cache while the graph had been traced on `[batch, 1, seq, kv_lora_rank]`, failing its own guard (`past_key_values[0].size()[1] == 1`) — the multi-token decode path that multimodal models use, since they export no standalone prefill graph. Branch on `kv_lora_rank`, the config-level marker for MLA, so both the capture and the runtime derive the same shape from the same place. Also casts the generate inputs to the model dtype in `_assert_generate_matches_eager`, which the half-precision models need (kimi's tower is bf16 while the regenerated inputs were fp32). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Cuts the torch.export sweep from ~171 to ~55 real failures, and removes a CUDA cascade that was reporting 504 phantom failures on top of that. Runtime / decomposition: - granite_speech names its audio tower `encoder`, which the canonical `get_encoder(modality="audio")` lookup does not search, so the model exported as a single graph with a baked audio-feature length. Its masked_scatter then overran on CUDA, poisoning two xdist workers. Answer the accessor and it decomposes properly (inherited by granite_speech_plus). - Anyres vision towers (llava_next, llava_next_video, llava_onevision, granite4_vision) export as tower + projector with the packing done in the runtime, so `image_sizes` is no longer baked into the graph. granite4_vision gets per-layer deepstack support; the grid side comes off the projected tensor, not the vision config, since a deepstack projector downsamples. - Route each modality's own generate kwarg onto its graph's feature input, filling only what nothing else named. A video getter that declares the generic `pixel_values` was being handed the images. - Keep the 2D padding mask when the decode graph was traced on one, and pad it to the static cache width. `prepare_inputs_for_generation` upgrades it to 4D for any compileable cache, which is wrong for a graph traced on the 2D form. - Rebuild an `EncoderDecoderCache`'s cross half with full attention: generate builds both halves from the same decoder config, sliding layer_types and all. - Move every tensor a growing cache layer already holds onto the target device; a sliding layer builds `_sliding_window_tensor` in `__init__` and relies on its own lazy init to place it. - Choose growing vs fixed-size by the layer's kind, not `get_max_length()`. A `DynamicSlidingWindowLayer` grows while reporting its window as a max length, so it took the static path and got rank-1 empties the graph cannot index. - Read a growing sliding layer's length off its keys tensor while tracing, and normalise the step counter in the pytree context, so a graph traced at one step accepts a cache at another. - Concatenate `token_type_ids` when merging decode steps; left at one step it specialised the merged graph's query axis back to 1. - Drop kwargs that are `None` before tracing when the parameter defaults to `None`. torch.export records them as placeholders carrying no value, so the graph declares an input the runtime then has to fill for no benefit. Masking: - `prepare_padding_mask` pads by `sym_max(...)` instead of branching on a width that can be unbacked (opt sizes its own mask from the cache length). - `_ignore_causal_mask_sdpa` returns before reading that width while tracing; it already returned False there, so the comparisons were only a guard source. Models: - internvl / qianfan_ocr: `pixel_shuffle` compared a dimension against a float `scale_factor`, which is both vacuous (any integer % 0.5 == 0) and asks the exporter for a value range over a float. Check divisibility by the reciprocal through `torch_compilable_check`. - ernie4_5_vl_moe: hoist the temporal even/odd gather behind `get_vision_temporal_slice_index`, following the existing `get_vision_*_index` convention, and forward kwargs to the resampler so the precompute reaches it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ExportedGenerator` carried a text-generation loop and a vision/audio one in the same class: three methods plus two module-level helpers (~200 lines) that a text-only runtime never reaches, and a fork in `forward` between feeding token ids and running them through an embedding graph. Introduce `_text_feed` as the seam — the base hands the ids straight to the graph — and move the rest to `ExportedMultimodalGenerator`: the modality merge, each modality's feed, the M-RoPE positions, and the modality half of `_validate_model_kwargs`. `from_runners` already branched on whether the decode graph takes embeddings, so it picks the class there and stays the single entry point. Cache construction, masks and the decode loop are untouched and shared. Also give `OnnxModelRunner` and `ExecutorchModelRunner` the `mask_rank` the dynamo runner had, so the "feed the mask kind this graph was traced with" rule is not dynamo-only: ORT reports input shapes on the session, ExecuTorch through `input_tensor_meta`. Declared on `ModelRunner` so the contract is documented once and the call sites can read the attribute directly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- `ModelRunner.cache_input_shapes` has no reader anywhere: it predates `kv_geometry`, which is what actually sizes the cache the runtime feeds back. Remove the attribute, the helper behind it, and all three backends' computations of it. - `_fresh_recurrent_cache` is referenced only by its own definition. - `_traced_mask_dict_ranks` and `_traced_mask_rank` ran the same spec check and the same placeholder scan, the single rank being the first element of the list the other built. One `_traced_mask_ranks` returns both, and exactly one of the two is set — a graph either took a plain mask or a per-type dict. 1264 -> 1193 lines, no behaviour change: dynamo 54 passed (including nemotron_h for the per-type mask path), and ONNX/ExecuTorch unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Runtime: - `forward` fed the graph-declared kwargs twice. The second pass selects the same set as the slicing loop above it (`name in input_names and name in kwargs and name not in feed`), so it can never add anything — instrumented across llava_next / bloom / csm / paligemma and it fired zero times. - Declare `mask_dict_ranks` on `ModelRunner` beside `mask_rank`, so the ExecuTorch runner has it and `_mask_feed` reads the attribute rather than probing with `getattr`. Both are the same question — which mask kind the graph took — and exactly one of them is ever set. Export: - `decompose_multimodal` had the body of `_modality_owner` copied out inline; call the helper, which leaves its `base` local unused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_ORT_TO_TORCH_DTYPE` was seven entries maintained by hand, and `.get()` on it turned anything absent into `None`. ORT spells its element types in ONNX's own vocabulary, so go through ONNX's enum and torch's own mapping (`JitScalarType.from_onnx_type`): every dtype ONNX has is covered, including the fp8 types the table omitted, and a non-tensor type (`seq(tensor(float))`) still resolves to `None` because it names no tensor. Lives next to `_session_kv_geometry`, above the runner that uses it, the way the dynamo and ExecuTorch helpers sit with theirs. `onnx` is an optional dependency, so the import is local — only an ONNX runner ever asks. bloom/opt ONNX unchanged at 10 failed / 6 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Runtime: beam-expanded cache batch, graph-derived cacheless generation (openai-gpt), pytree kwargs matched by leaf names on every backend (encoder_outputs, mask dicts, the ONNX input. prefix for mutated inputs), prompt ids for modality graphs that declare them, dict-keyed cache write-back, encoder-side feature scatter for multimodal encoder-decoders (florence2). Exporters: value_head_dim on Cache.early_initialization (MLA caches build both buffers key-wide otherwise, silent OOB on CUDA), scalar counters normalized before the pytree walk (orphaned SymInt leaf broke every ET sliding-window model), output names baked into the .pte from the program's out_spec, best-effort assert-feeder erase (fx SystemError on audio towers), aten.mul.Scalar / aten.rsub.Scalar ONNX translations. Models: qwen2_audio get_audio_features, omni rope-delta branch on the embeds path, openai-gpt consuming generation-loop kwargs, MiniMaxCache init sync. Skips recorded with evidence: csm, xlnet, xlm, vibevoice_asr, kyutai, deepseek_v4, minimax, xlstm, blt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…time Exporters / runtime: - Keep the captured prefill as its own component when the decode cache carries recurrent state (conv / linear-attention / lightning layers) or is the only writer of cross-attention K/V: a graph traced mid-generation bakes its step-regime branches (conv mask width, lightning chunk boundary, warm-cache cross reads), which a prompt on a fresh cache cannot satisfy. Unblocks multi-token decode for lfm2_vl, inkling, minicpmv4_6, qwen3_5(-moe), minimax_m3_vl and idefics. - Materialize the sparse-indexer cache slot before the `is_initialized` skip: `early_initialization` marks the layer done while the indexer buffer and its cpu counter are untouched (deepseek_v32 / axk2 / glm_moe_dsa fed the graph a cpu scalar for a cuda input). - Normalize `cumulative_length_int` out of the pytree context, as the growing counter already was — a static sliding layer otherwise pins the graph to the step it was traced at (exaone4_5, muse_glimmer, t5gemma, t5gemma2). - Restore `forward` by deleting the capture wrapper instead of assigning the bound method back: a later `copy.copy` of the module carried a `forward` bound to the original, so patches and getter captures on the copy silently never fired (empty modality capture on the omni thinkers). - Modality specs carry their aux kwargs (`*_position_mask`, `*_attention_mask`, `*_sizes`) so they are neither trimmed as per-step tensors nor mistaken for a grid by the precompute; scatter features by an explicit position mask for a model with no placeholder token id. - Audio: bridge the getter's `feature_attention_mask` seam to the packed `feature_lens` preparer, and derive `audio_seqlens` for the audio-merged M-RoPE layout at both the export precompute and the runtime rebuild. - ONNX: translate `higher_order.associative_scan` to a dynamic-trip-count `Loop`, keyed on the higher-order op itself. - ExecuTorch: turn the mixers' `use_associative_scan` off for the trace — no lowering for the op, and no loop primitive to run it as. - Build the skip-scope facets once and prefix each with the backend, so `<backend>.<facet>` keys work without mirroring the list by hand. Models: - mamba / falcon_mamba / jamba: pass the cached SSM state into the selective scan (`initial_states`, the contract mamba2 and the deltanet models already implement). A multi-token continuation over a warm cache restarted from zero state — prefill 6 then continue 6 diverged 2e-2 from the full forward, now bit-exact. Covered by `test_multi_token_continuation`; the generic continuation test skips every `cache_params` model, hence unnoticed. - idefics: cache the gated cross-attention K/V (a `cross_attention` cache layer, immune to text rollbacks) instead of raising NotImplementedError and re-feeding image features every step. `layer_types` names the gated layers' slots; decode steps carry text and mask only, and assisted generation now runs. - kosmos2_5: add `get_image_features` (vision tower + image-to-text projection) so the decomposition has a modality seam to split out. - minicpmv4_6: forward `**kwargs` through `get_image_features` / `get_video_features` so the precomputed vision tensors reach the tower. - qwen2_5_omni: `torch_compilable_check` in place of `sum(...tolist())`. Skips recorded with evidence: rwkv (state outside the cache API), prophetnet (its own single-token assertion), idefics cross-attention routing, perception_lm (videos through a second image getter), sam2 dynamic (sympy), zamba (per-head mixer), and the ExecuTorch-only SSM multi-token limitation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four ONNX-side bugs, found by the full sweep (91 failures, ~67 of them these): - `patch_model_outputs` captured the graph's input names *after* running the traced forward, and appended on every trace. A forward mutates its pytree kwargs in place, so a cache whose states the model creates on that very call contributed leaf names the graph never took as inputs — shifting every later name onto the wrong tensor: a recurrent model's `attention_mask` was declared as `cache_params.layers.0.conv_states.0` (rank-2 int64), and feeding it displaced the real input. Names are now taken before the forward and replace rather than accumulate. - `_ort_to_torch_dtype` went through `JitScalarType`, which aliases ONNX's plain integers onto torch's *quantized* dtypes (`int32` -> `qint32`, `int8`/`uint8` -> `qint8`/`quint8`). Casting a real tensor to one raises `empty_strided not supported on quantized tensors`, which every grid VLM hit on its int32 `cu_seqlens` (31 tests). Resolved by name against `torch` instead, covering fp8 and complex as well. - The ONNX runner zipped the graph's declared cache inputs against the cache's tensor leaves *positionally*. A recurrent layer's states are `None` until a step creates them, so the leaves are a subset of what the graph declares and in a different order: the feed silently lost inputs (`Required inputs (['input.cache_params.layers.0.conv_states.0']) are missing`) and could pair a name with the wrong tensor. Now keyed by each input's dotted leaf path, with a `None` slot fed the zeros the layer's own lazy initialization would have made, sized from the declared shape. - Guard `_prepare_masked_omni_audio_inputs` on the model actually having the chunked-audio helpers: unlike `feature_lens`, the `feature_attention_mask` marker is not omni-specific, so it fired on qwen2_audio and asked its module for `chunk_and_pad_features`. Also: keep a *cpu* ONNX export on the sequential SSM scan — the associative scan's `generic` combine mode (what a cpu tensor selects) lowers through `vmap`, which `run_decompositions` cannot take apart. Recorded with evidence: llava_next_video hits the same onnxscript `SplitToSequence` constant-folding crash as prophetnet/zoedepth (optimizer off under dynamic shapes); higgs_audio_v2 cannot be driven by the generic generate loop — its `prepare_inputs_for_generation` masks the audio ids the cache already holds and in decode drops `input_ids` to pass only the last codebook row (export and per-component parity still run). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two torch-export fixes found by the full sweep: - `_materialize_layers` moved a layer's own `__init__` tensors to the target device only on the growing branch, and below the `is_initialized` skip that `early_initialization` has already tripped. A sparse-index layer's counter (minimax_m3_vl's `idx_cumulative_length`; `StaticLayer.lazy_initialization` moves `cumulative_length` and nothing else) was therefore left on cpu among cuda leaves, and the graph's input spec mismatched on device. The move now runs first, for every layer kind. - Don't drive the runtime for a static-shape multi-modal export: those models embed their text in a graph of its own, captured on the prompt, so under static shapes it is specialized to the prompt's length and cannot serve the loop's 1-token steps (`Guard failed: input_ids.size()[1] == 39`). The static-shape exports themselves are still asserted; driving them needs a length-generic embedder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_patch_reshape` clones a non-contiguous input before `reshape`/`view`, and asked `input.is_contiguous()` to decide. Deciding contiguity compares the strides against the products of the sizes, so on a tensor sized by *unbacked* symbols — a NaViT packer's data-dependent patch count — the question has no answer and plain `is_contiguous()` raises `GuardOnDataDependentSymNode`, failing the whole export on a check meant only to pick a fast path. Nothing in the model was at fault: the same components export fine through `torch.export` (and through `run_decompositions` with the ONNX table) once this patch is out of the way, which is why only the ONNX backend saw it. `torch._prims_common.is_contiguous_or_false` answers "not provably contiguous", which lands on the copy — always correct, and the copy is what makes the view legal in the first place. Fixes the data-dependent guard failures for lfm2_vl, minicpmv4_6, paddleocr_vl and modernvbert (20 tests); wavlm / wav2vec2_bert, the models the patch was written for, keep passing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…er-less VLM The branch that keeps a captured prefill as its own component was keyed on "a multi-modal model with no modality getter", which was meant for idefics (whose gated layers cache the image K/V the prefill writes). It also caught mllama: that model's cross-attention reads a per-step `cross_attention_mask` no generic loop grows, so driving it indexed past the mask — an out-of-bounds device-side assert that poisons the whole xdist worker, turning one real failure into ~148 cascaded ones. Its multi-token variant is skipped for exactly this reason. Keyed on the layer kind now (`DynamicCrossAttentionLayer`), which is what the branch was always about. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r-tie ids - `_modality_feed` matched a graph's declared inputs by exact name, so a *list-valued* input never reached it: ernie4_5_vl_moe's `temporal_slice_index` (the even/odd gather pair) is declared as the flattened leaves `temporal_slice_index.0` / `.1`, and ORT rejected the feed for missing them. It now asks `_declares`, which already knows a graph naming only the leaves still takes the pytree — the same rule the decode feed uses. - Compare generated ids step by step, and only while eager's own top-2 gap says the choice is not a coin flip. A tiny random model routinely has gaps of a few 1e-3, and how far a backend's kernels drift is model-specific (an ONNX chameleon drifts past 1e-3 where kosmos2_5 stays at 6e-5, dynamo lands at ~1e-9), so the argmax flips on that drift while saying nothing about correctness — which is what made chameleon and kosmos2_5 fail only in some orderings. Numeric fidelity is still asserted per component above; half precision, whose rounding scale is uniform, still compares scores too. - Skip git's multi-token capture with the measured evidence: its forward corrects for the image tokens only on a single-token step (`position_ids` offset by the past length, `attention_mask` widened by the cached image tokens, both gated on `seq_len == 1`), so a capture traced at 2 tokens bakes those branches off and every later 1-token step runs without them — the second step's scores drift ~1e-2 from eager. The single-token capture keeps the branches and is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`MllamaModel.forward` indexes `cross_attention_mask[:, :, arange(seq_len) + past_seen_tokens]`, which only stays in bounds because the model's own `_update_model_kwargs_for_generation` appends a row per decode step. Any other caller — an exported runtime, a hand-written loop — indexes past the end, and an out-of-bounds gather is a device-side assert, not an exception: it kills the CUDA context, so in a parallel test run one such call turned into ~148 cascaded failures on the poisoned workers. The appended row is a copy of the last one, so clamping to the mask's own last row is exactly what that growth would have produced, and cannot run off the end. mllama's eager suite is unchanged (the 4 pre-existing static-cache / multi-gpu failures, one test more passing). Its export skip now records what this unblocks and what still holds it: the out-of-bounds is gone and the remaining blocker is routing `pixel_values` through the runtime, for which the idefics recipe (cross slots declared in `layer_types`) is verified to work — held back only by mllama's static-cache compile path, since that layer type is necessarily dynamic. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ExecutorchModelRunner` never set `dtype`, so it inherited the base class's fp32. The generation layer sizes the cache it feeds from that, so every half-precision export — a grouped-mm MoE, a varlen-attention VLM — was handed fp32 cache leaves for a program declaring bf16 ones, and the method refused them while binding its inputs. That is the whole `set_inputs ... error 0x12` cluster: 109 failures across ~40 MoE models in the ExecuTorch sweep, all from one missing attribute. The precision is in the `.pte` metadata already; read it off the first floating-point input. `set_inputs` failures no longer count as a backend limitation, anywhere. Binding inputs is our side of the contract — it fails when the runtime hands the method something it never declared — and treating it as ExecuTorch's ceiling is what hid this for every MoE model at once. Only `execute()` keeps that treatment (a portable kernel refusing the step's shapes mid-run), and the generate drive now absorbs it there too, the way the component checks already did: deepseek_v3's merged decode reaches exactly that ceiling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rwritten Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # src/transformers/cache_utils.py # src/transformers/generation/utils.py # src/transformers/masking_utils.py
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…NO runner Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ences Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… export tests Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cross-attention) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…fig fields Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fresh process had none registered, so torch.export.load could not rebuild the input spec. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The trace specializes an int logits_to_keep, so as an input it was a port nothing read. It now joins the output flags prepare_for_export pops, and the traced forward receives the flags the config doesn't declare as fixed values. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…th ONNX Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… declares Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ength, shorter metadata accessors Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
executorch 1.4.1 prunes empty cat inputs with guard_or_true and 1.5.0 writes non-finite constants for flatc itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eed; honour fft norm on OpenVINO Each cited model exports and matches eager without them. The OpenVINO fft patch ignored norm, giving wrong values for ortho/forward. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…thod, one ONNX/OpenVINO compare Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
CI recapDashboard: View test results in Grafana |
Saving goes through each backend's `save_artifact` (and `save_pretrained`), so `OnnxConfig.output_path`, `external_data` and `export_params` and `OpenVINOConfig.output_path` were a second, partial way to save; the ONNX ones were also passed to `torch.onnx.export`, where they do nothing without `f`. `compress_to_fp16` stays: it compresses the converted model, not the saved file. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Now we can deprecate optimum-onnx/optimum-executorch
Code Agent Policy
The Transformers repo is currently being overwhelmed by a large number of PRs and issue comments written by
code agents. These often are low-quality, or fix extremely minor issues that occur rarely or never in practice.
As a result, we're instituting a rule that first-time contributors should not use code agents to submit PRs or issues.
We'd also ask autonomous "OpenClaw"-like agents not to open any PRs or issues.
Issues/PRs from first-time contributors that violate this rule will probably just be closed without review, and we
might block you, especially if you open more than one or appear to be deliberately ignoring this. We especially do not
want new contributors to jump in on random issues to contribute an agent-written fix. This creates lots of noise
for reviewers and other users and will almost certainly get you blocked.
For more information, please read
CONTRIBUTING.md.Before submitting
Pull Request checks?
to it if that's the case.
Who can review?
Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.