Skip to content

🚨 Exporters Runtimes : Export and run any Transformers model (Dynamo/OpenVINO/OnnxRuntime/ExecuTorch) - #48986

Draft
IlyasMoutawwakil wants to merge 115 commits into
mainfrom
exporters-runtime
Draft

IlyasMoutawwakil wants to merge 115 commits into
mainfrom
exporters-runtime

Conversation

@IlyasMoutawwakil

@IlyasMoutawwakil IlyasMoutawwakil commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

CPU CI GPU run-slow

What does this PR do?

Now we can deprecate optimum-onnx/optimum-executorch

Code Agent Policy

The Transformers repo is currently being overwhelmed by a large number of PRs and issue comments written by
code agents. These often are low-quality, or fix extremely minor issues that occur rarely or never in practice.
As a result, we're instituting a rule that first-time contributors should not use code agents to submit PRs or issues.
We'd also ask autonomous "OpenClaw"-like agents not to open any PRs or issues.

Issues/PRs from first-time contributors that violate this rule will probably just be closed without review, and we
might block you, especially if you open more than one or appear to be deliberately ignoring this. We especially do not
want new contributors to jump in on random issues to contribute an agent-written fix. This creates lots of noise
for reviewers and other users and will almost certainly get you blocked.

For more information, please read CONTRIBUTING.md.

  • (First-time contributors only): I confirm that this PR description and code is not written by an LLM or code agent

Before submitting

  • This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case).
  • Did you read the contributor guideline and the
    Pull Request checks?
  • Was this discussed/approved via a Github issue or the forum? Please add a link
    to it if that's the case.
  • Did you make sure to update the documentation with your changes according to the guidelines?
  • Did you write any new necessary tests?

Who can review?

Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.

IlyasMoutawwakil and others added 30 commits August 14, 2026 15:14
Every M-RoPE model carried its own ~150-line `get_rope_index`, copied and
lightly edited per family. Move the layouts into
`modeling_rope_utils.get_mrope_index`, dispatched on a new
`config.mrope_layout`, and reduce each model's `get_rope_index` to a
one-line wrapper over it.

Four layouts cover every M-RoPE family:

- `interleaved_runs` — the shared Qwen2-VL-style layout (15 families).
  Runs come from `mm_token_type_ids`; the per-family differences are now
  config attributes (`spatial_merge_size`, `tokens_per_second`,
  `temporal_merge_size`, and a new `timestamped_video_frames` for the
  processors that split a video into per-frame spans).
- `indexed_images` — HunYuanVL: a 1D text baseline overwritten per image
  span with `(width, height, image_index)` on the last three of
  `mrope_section`'s axes.
- `audio_chunked` — Qwen2.5-Omni: spans located by scanning `input_ids`
  for placeholder tokens, audio laid out along the temporal axis, and
  audio-in-video interleaved in `seconds_per_chunk` time chunks.
- `audio_merged` — Qwen3-Omni: same shape, but scanning for the opening
  tokens, merging audio and video position by position, and returning
  float positions since its temporal axis is fractional.

The layouts are pure functions of their inputs, so they can be driven and
tested without a model instance, and `get_mrope_index` is the single place
a config is read. An unknown `mrope_layout` raises `NotImplementedError`
rather than silently producing wrong positions.

Also adds `get_mrope_text_positions` (the shared no-vision fallback, which
is what the Omni families use for audio-only prompts), declares
`vision_start_token_id` on the Omni thinker configs (it was read but never
declared), and drops the now-absorbed `get_llm_pos_ids_for_vision` /
`get_chunked_index` helpers — which in turn removes Qwen3-Omni's Talker
delegating overrides.

GLM-Image keeps its bespoke implementation: its `get_rope_index` caches
decode-stage positions on the module, so it is not a pure function.

Tests: `tests/utils/test_modeling_rope_utils.py::MRopeLayoutTest` covers
all four layouts against positions produced by the implementations they
were ported from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Multimodal models re-declare the same handful of methods almost verbatim.
Counting identical bodies across the 71 modeling files that define
`get_placeholder_mask`: `get_vision_position_ids` was 17 copies of one
function, `compute_3d_position_ids` 14 of 17, `_prepare_position_ids_for_
generation` 14 of 16, `_get_image_nums_and_video_nums` 10 of 17.

Add `modeling_multimodal_utils.py` with two mixins, one per attachment
point, since these helpers live on two different classes:

- `MultiModalPreTrainedModelMixin` on the base `<X>Model`:
  `get_vision_position_ids` (now a thin method form of
  `modeling_rope_utils.get_mrope_vision_positions`, which is the same
  function it had been copied from), `get_rope_index`, and
  `compute_3d_position_ids`.
- `MultiModalGenerationMixin` on `<X>ForConditionalGeneration`, mixed in
  before `GenerationMixin` so its overrides win while `super()` still
  reaches the text-only behaviour: `_prepare_position_ids_for_generation`
  and `_get_image_nums_and_video_nums`.

Mixing them into the family roots and regenerating cascades to every
child, so this deletes the copies rather than relocating them. Families
that genuinely differ keep their overrides: qwen2_5_vl (threads
`second_per_grid_ts`), hunyuan_vl (image-only layout), the two Omni
models (audio signatures), glm_image (caches decode positions on the
module), exaone4_5 (BC shim), video_llama_3, and the glm4v family for
`_get_image_nums_and_video_nums` (different token ids).

`get_rope_index` on the mixin swallows `**kwargs` without forwarding
them, matching the per-model wrappers it replaces — generation passes the
whole model-kwargs dict through it.

Two things a mixin cannot do that a parent class body can, both of which
bit during this refactor and are handled at the mixin level rather than
per model:

- It cannot be *deleted* by modular's `raise AttributeError` idiom. Four
  models used that to drop these helpers (video_llama_3, exaone4_5,
  kimi_k25, muse_glimmer) and the stubs started being emitted into
  generated code, where `generate` would have called them. Those four are
  exactly the models whose configs declare no M-RoPE, so the mixin now
  gates on `uses_mrope(self.config)`: `compute_3d_position_ids` returns
  `None` and `_prepare_position_ids_for_generation` returns the text
  positions, which is what an absent method used to give them. The stubs
  are gone.
- It changes what `super()` means in existing overrides. cosmos3_edge and
  hunyuan_vl start from the plain 2D text positions, and their `super()`
  call silently began returning the mixin's four-axis result instead
  (caught by five cosmos3_edge generate tests). Rather than pin those two
  call sites, `_prepare_position_ids_for_generation` is now a thin hook
  that hands `text_positions` to an overridable
  `_prepare_mrope_position_ids_for_generation`, so no override has to
  reason about the MRO.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The vision helpers in `vision_utils.py` take their settings as plain
arguments, but several of those arguments were only available on the
*module*: hardcoded in `__init__` (`self.interpolation_mode = "bilinear"`),
derived there (`self.num_grid_per_side = int(config.num_position_embeddings
** 0.5)`), or passed as a literal at the call site (`merge_temporal=True`).
That makes the helpers uncallable from a saved config alone — you need a
built module to know how to call them.

Declare all of them on the vision configs instead, and have the modules
read them back, so config is the single source of truth:

- fields for the hardcoded values: `interpolation_mode`,
  `interpolation_align_corners`, `interpolation_padding`,
  `merge_temporal_attention` (was a literal `True` in
  `get_vision_cu_seqlens`), `include_temporal_position_ids` (likewise in
  `get_vision_position_ids`), and `resample_before_merge`, which records
  that a family resamples its position grid at patch resolution — the
  reason kimi_k25 / muse_glimmer / paddleocr_vl pass a merge size of 1
  while the qwen3_vl family passes its real one.
- properties for the derived ones: `num_grid_per_side`, muse_glimmer's
  `window_size`, and `spatial_merge_size` on kimi_k25 / muse_glimmer,
  which spell it `merge_kernel_size` / `merge_size` — so every vision
  config now exposes that name.

Five config classes cover all nine vision modules, since six of them
inherit `Qwen3VLVisionConfig`. minicpmv4_6 is left alone: its
`get_vision_window_index(..., spatial_merge_size=1, patch_size=1)` call
passes an already-merged grid, which is a property of that call site
rather than a config setting.

Verified by resolving every helper argument off a bare
`CONFIG_MAPPING[...]()` and running the helpers with no model or module
in play, for all ten families.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`configuration_utils` imports `modeling_rope_utils` for
`RotaryEmbeddingConfigMixin`, so that module has to stay importable in a
torch-free install — which is why every torch annotation in it is
deferred. Hosting 800+ lines of tensor code there was a standing trap:
CI caught it once already when a `dtype: torch.dtype = torch.long`
default was evaluated at import time.

The layouts belong with the mixins that call them anyway:
`modeling_multimodal_utils` imports torch unconditionally and is only
reached from modeling code. Move `get_mrope_index`, the four layouts and
their helpers, `get_mrope_vision_positions` / `get_mrope_image_positions`
/ `get_mrope_text_positions` and `uses_mrope` there, leaving
`modeling_rope_utils` as the frequency machinery it was.

Splits the tests the same way: `tests/utils/test_modeling_multimodal_utils.py`
holds the layout tests, `test_modeling_rope_utils.py` the frequency ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Now that every M-RoPE layout and vision knob is declared on the config,
the exporter no longer needs a model (or module introspection) to know
how to call the shared helpers:

- `precompute_export_inputs(config, inputs) -> dict` is config-only and
  no longer mutates its argument. It took a model solely for
  `model.get_rope_index`, which is now a thin wrapper over
  `get_mrope_index(config, ...)`, so it calls the layout directly. A
  component without a config (`embed_tokens`) simply has nothing to
  precompute, which the caller decides rather than the function
  silently accepting a module.
- The grid_thw preparer reads `merge_temporal_attention` off the config
  instead of inferring clip-level attention from
  `hasattr(module, "get_vision_frame_index")`, and derives the resample
  merge size the same way the modules do
  (`1 if resample_before_merge else spatial_merge_size`).
- Drop the `lm_head` component: nothing consumed it (the runtime never
  built a runner for it), so it was an exported graph holding a single
  matmul. Rename `language_model` -> `text_decoder`, which is what
  `get_decoder()` returns and no longer reads like the `decode` phase.

Adds `test_precomputed_inputs_match_eager`: the other export tests feed
the precomputed tensors to both the eager reference and the exported
program, so a wrong-but-consistent value cancels out and passes. This
compares the model run with precomputed inputs against the same model
run without them, at exact tolerance — verified to fail when the
preparer derives them at the wrong merge size.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`materialize_cache_layers` derived the KV cache geometry as
`num_key_value_heads x head_dim`, which is wrong for multi-head latent
attention: deepseek_v3 and kimi_k25 cache a single compressed latent head
of `kv_lora_rank`. The runtime therefore handed the decode graph a
`[batch, num_key_value_heads, 0, head_dim]` cache while the graph had been
traced on `[batch, 1, seq, kv_lora_rank]`, failing its own guard
(`past_key_values[0].size()[1] == 1`) — the multi-token decode path that
multimodal models use, since they export no standalone prefill graph.

Branch on `kv_lora_rank`, the config-level marker for MLA, so both the
capture and the runtime derive the same shape from the same place. Also
casts the generate inputs to the model dtype in
`_assert_generate_matches_eager`, which the half-precision models need
(kimi's tower is bf16 while the regenerated inputs were fp32).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Cuts the torch.export sweep from ~171 to ~55 real failures, and removes a
CUDA cascade that was reporting 504 phantom failures on top of that.

Runtime / decomposition:
- granite_speech names its audio tower `encoder`, which the canonical
  `get_encoder(modality="audio")` lookup does not search, so the model
  exported as a single graph with a baked audio-feature length. Its
  masked_scatter then overran on CUDA, poisoning two xdist workers. Answer
  the accessor and it decomposes properly (inherited by granite_speech_plus).
- Anyres vision towers (llava_next, llava_next_video, llava_onevision,
  granite4_vision) export as tower + projector with the packing done in the
  runtime, so `image_sizes` is no longer baked into the graph. granite4_vision
  gets per-layer deepstack support; the grid side comes off the projected
  tensor, not the vision config, since a deepstack projector downsamples.
- Route each modality's own generate kwarg onto its graph's feature input,
  filling only what nothing else named. A video getter that declares the
  generic `pixel_values` was being handed the images.
- Keep the 2D padding mask when the decode graph was traced on one, and pad it
  to the static cache width. `prepare_inputs_for_generation` upgrades it to 4D
  for any compileable cache, which is wrong for a graph traced on the 2D form.
- Rebuild an `EncoderDecoderCache`'s cross half with full attention: generate
  builds both halves from the same decoder config, sliding layer_types and all.
- Move every tensor a growing cache layer already holds onto the target device;
  a sliding layer builds `_sliding_window_tensor` in `__init__` and relies on
  its own lazy init to place it.
- Choose growing vs fixed-size by the layer's kind, not `get_max_length()`. A
  `DynamicSlidingWindowLayer` grows while reporting its window as a max length,
  so it took the static path and got rank-1 empties the graph cannot index.
- Read a growing sliding layer's length off its keys tensor while tracing, and
  normalise the step counter in the pytree context, so a graph traced at one
  step accepts a cache at another.
- Concatenate `token_type_ids` when merging decode steps; left at one step it
  specialised the merged graph's query axis back to 1.
- Drop kwargs that are `None` before tracing when the parameter defaults to
  `None`. torch.export records them as placeholders carrying no value, so the
  graph declares an input the runtime then has to fill for no benefit.

Masking:
- `prepare_padding_mask` pads by `sym_max(...)` instead of branching on a width
  that can be unbacked (opt sizes its own mask from the cache length).
- `_ignore_causal_mask_sdpa` returns before reading that width while tracing;
  it already returned False there, so the comparisons were only a guard source.

Models:
- internvl / qianfan_ocr: `pixel_shuffle` compared a dimension against a float
  `scale_factor`, which is both vacuous (any integer % 0.5 == 0) and asks the
  exporter for a value range over a float. Check divisibility by the reciprocal
  through `torch_compilable_check`.
- ernie4_5_vl_moe: hoist the temporal even/odd gather behind
  `get_vision_temporal_slice_index`, following the existing
  `get_vision_*_index` convention, and forward kwargs to the resampler so the
  precompute reaches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ExportedGenerator` carried a text-generation loop and a vision/audio one in the
same class: three methods plus two module-level helpers (~200 lines) that a
text-only runtime never reaches, and a fork in `forward` between feeding token
ids and running them through an embedding graph.

Introduce `_text_feed` as the seam — the base hands the ids straight to the
graph — and move the rest to `ExportedMultimodalGenerator`: the modality merge,
each modality's feed, the M-RoPE positions, and the modality half of
`_validate_model_kwargs`. `from_runners` already branched on whether the decode
graph takes embeddings, so it picks the class there and stays the single entry
point. Cache construction, masks and the decode loop are untouched and shared.

Also give `OnnxModelRunner` and `ExecutorchModelRunner` the `mask_rank` the
dynamo runner had, so the "feed the mask kind this graph was traced with" rule
is not dynamo-only: ORT reports input shapes on the session, ExecuTorch through
`input_tensor_meta`. Declared on `ModelRunner` so the contract is documented
once and the call sites can read the attribute directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- `ModelRunner.cache_input_shapes` has no reader anywhere: it predates
  `kv_geometry`, which is what actually sizes the cache the runtime feeds back.
  Remove the attribute, the helper behind it, and all three backends'
  computations of it.
- `_fresh_recurrent_cache` is referenced only by its own definition.
- `_traced_mask_dict_ranks` and `_traced_mask_rank` ran the same spec check and
  the same placeholder scan, the single rank being the first element of the list
  the other built. One `_traced_mask_ranks` returns both, and exactly one of the
  two is set — a graph either took a plain mask or a per-type dict.

1264 -> 1193 lines, no behaviour change: dynamo 54 passed (including nemotron_h
for the per-type mask path), and ONNX/ExecuTorch unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Runtime:
- `forward` fed the graph-declared kwargs twice. The second pass selects the
  same set as the slicing loop above it (`name in input_names and name in
  kwargs and name not in feed`), so it can never add anything — instrumented
  across llava_next / bloom / csm / paligemma and it fired zero times.
- Declare `mask_dict_ranks` on `ModelRunner` beside `mask_rank`, so the
  ExecuTorch runner has it and `_mask_feed` reads the attribute rather than
  probing with `getattr`. Both are the same question — which mask kind the graph
  took — and exactly one of them is ever set.

Export:
- `decompose_multimodal` had the body of `_modality_owner` copied out inline;
  call the helper, which leaves its `base` local unused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_ORT_TO_TORCH_DTYPE` was seven entries maintained by hand, and `.get()` on it
turned anything absent into `None`. ORT spells its element types in ONNX's own
vocabulary, so go through ONNX's enum and torch's own mapping
(`JitScalarType.from_onnx_type`): every dtype ONNX has is covered, including the
fp8 types the table omitted, and a non-tensor type (`seq(tensor(float))`) still
resolves to `None` because it names no tensor.

Lives next to `_session_kv_geometry`, above the runner that uses it, the way the
dynamo and ExecuTorch helpers sit with theirs. `onnx` is an optional dependency,
so the import is local — only an ONNX runner ever asks.

bloom/opt ONNX unchanged at 10 failed / 6 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Runtime: beam-expanded cache batch, graph-derived cacheless generation
(openai-gpt), pytree kwargs matched by leaf names on every backend
(encoder_outputs, mask dicts, the ONNX input. prefix for mutated inputs),
prompt ids for modality graphs that declare them, dict-keyed cache
write-back, encoder-side feature scatter for multimodal encoder-decoders
(florence2).

Exporters: value_head_dim on Cache.early_initialization (MLA caches build
both buffers key-wide otherwise, silent OOB on CUDA), scalar counters
normalized before the pytree walk (orphaned SymInt leaf broke every ET
sliding-window model), output names baked into the .pte from the
program's out_spec, best-effort assert-feeder erase (fx SystemError on
audio towers), aten.mul.Scalar / aten.rsub.Scalar ONNX translations.

Models: qwen2_audio get_audio_features, omni rope-delta branch on the
embeds path, openai-gpt consuming generation-loop kwargs, MiniMaxCache
init sync. Skips recorded with evidence: csm, xlnet, xlm, vibevoice_asr,
kyutai, deepseek_v4, minimax, xlstm, blt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…time

Exporters / runtime:
- Keep the captured prefill as its own component when the decode cache carries
  recurrent state (conv / linear-attention / lightning layers) or is the only
  writer of cross-attention K/V: a graph traced mid-generation bakes its
  step-regime branches (conv mask width, lightning chunk boundary, warm-cache
  cross reads), which a prompt on a fresh cache cannot satisfy. Unblocks
  multi-token decode for lfm2_vl, inkling, minicpmv4_6, qwen3_5(-moe),
  minimax_m3_vl and idefics.
- Materialize the sparse-indexer cache slot before the `is_initialized` skip:
  `early_initialization` marks the layer done while the indexer buffer and its
  cpu counter are untouched (deepseek_v32 / axk2 / glm_moe_dsa fed the graph a
  cpu scalar for a cuda input).
- Normalize `cumulative_length_int` out of the pytree context, as the growing
  counter already was — a static sliding layer otherwise pins the graph to the
  step it was traced at (exaone4_5, muse_glimmer, t5gemma, t5gemma2).
- Restore `forward` by deleting the capture wrapper instead of assigning the
  bound method back: a later `copy.copy` of the module carried a `forward`
  bound to the original, so patches and getter captures on the copy silently
  never fired (empty modality capture on the omni thinkers).
- Modality specs carry their aux kwargs (`*_position_mask`,
  `*_attention_mask`, `*_sizes`) so they are neither trimmed as per-step
  tensors nor mistaken for a grid by the precompute; scatter features by an
  explicit position mask for a model with no placeholder token id.
- Audio: bridge the getter's `feature_attention_mask` seam to the packed
  `feature_lens` preparer, and derive `audio_seqlens` for the audio-merged
  M-RoPE layout at both the export precompute and the runtime rebuild.
- ONNX: translate `higher_order.associative_scan` to a dynamic-trip-count
  `Loop`, keyed on the higher-order op itself.
- ExecuTorch: turn the mixers' `use_associative_scan` off for the trace — no
  lowering for the op, and no loop primitive to run it as.
- Build the skip-scope facets once and prefix each with the backend, so
  `<backend>.<facet>` keys work without mirroring the list by hand.

Models:
- mamba / falcon_mamba / jamba: pass the cached SSM state into the selective
  scan (`initial_states`, the contract mamba2 and the deltanet models already
  implement). A multi-token continuation over a warm cache restarted from zero
  state — prefill 6 then continue 6 diverged 2e-2 from the full forward, now
  bit-exact. Covered by `test_multi_token_continuation`; the generic
  continuation test skips every `cache_params` model, hence unnoticed.
- idefics: cache the gated cross-attention K/V (a `cross_attention` cache
  layer, immune to text rollbacks) instead of raising NotImplementedError and
  re-feeding image features every step. `layer_types` names the gated layers'
  slots; decode steps carry text and mask only, and assisted generation now
  runs.
- kosmos2_5: add `get_image_features` (vision tower + image-to-text
  projection) so the decomposition has a modality seam to split out.
- minicpmv4_6: forward `**kwargs` through `get_image_features` /
  `get_video_features` so the precomputed vision tensors reach the tower.
- qwen2_5_omni: `torch_compilable_check` in place of `sum(...tolist())`.

Skips recorded with evidence: rwkv (state outside the cache API), prophetnet
(its own single-token assertion), idefics cross-attention routing,
perception_lm (videos through a second image getter), sam2 dynamic (sympy),
zamba (per-head mixer), and the ExecuTorch-only SSM multi-token limitation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four ONNX-side bugs, found by the full sweep (91 failures, ~67 of them these):

- `patch_model_outputs` captured the graph's input names *after* running the
  traced forward, and appended on every trace. A forward mutates its pytree
  kwargs in place, so a cache whose states the model creates on that very call
  contributed leaf names the graph never took as inputs — shifting every later
  name onto the wrong tensor: a recurrent model's `attention_mask` was declared
  as `cache_params.layers.0.conv_states.0` (rank-2 int64), and feeding it
  displaced the real input. Names are now taken before the forward and replace
  rather than accumulate.

- `_ort_to_torch_dtype` went through `JitScalarType`, which aliases ONNX's
  plain integers onto torch's *quantized* dtypes (`int32` -> `qint32`,
  `int8`/`uint8` -> `qint8`/`quint8`). Casting a real tensor to one raises
  `empty_strided not supported on quantized tensors`, which every grid VLM hit
  on its int32 `cu_seqlens` (31 tests). Resolved by name against `torch`
  instead, covering fp8 and complex as well.

- The ONNX runner zipped the graph's declared cache inputs against the cache's
  tensor leaves *positionally*. A recurrent layer's states are `None` until a
  step creates them, so the leaves are a subset of what the graph declares and
  in a different order: the feed silently lost inputs (`Required inputs
  (['input.cache_params.layers.0.conv_states.0']) are missing`) and could pair
  a name with the wrong tensor. Now keyed by each input's dotted leaf path,
  with a `None` slot fed the zeros the layer's own lazy initialization would
  have made, sized from the declared shape.

- Guard `_prepare_masked_omni_audio_inputs` on the model actually having the
  chunked-audio helpers: unlike `feature_lens`, the `feature_attention_mask`
  marker is not omni-specific, so it fired on qwen2_audio and asked its module
  for `chunk_and_pad_features`.

Also: keep a *cpu* ONNX export on the sequential SSM scan — the associative
scan's `generic` combine mode (what a cpu tensor selects) lowers through
`vmap`, which `run_decompositions` cannot take apart.

Recorded with evidence: llava_next_video hits the same onnxscript
`SplitToSequence` constant-folding crash as prophetnet/zoedepth (optimizer
off under dynamic shapes); higgs_audio_v2 cannot be driven by the generic
generate loop — its `prepare_inputs_for_generation` masks the audio ids the
cache already holds and in decode drops `input_ids` to pass only the last
codebook row (export and per-component parity still run).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two torch-export fixes found by the full sweep:

- `_materialize_layers` moved a layer's own `__init__` tensors to the target
  device only on the growing branch, and below the `is_initialized` skip that
  `early_initialization` has already tripped. A sparse-index layer's counter
  (minimax_m3_vl's `idx_cumulative_length`; `StaticLayer.lazy_initialization`
  moves `cumulative_length` and nothing else) was therefore left on cpu among
  cuda leaves, and the graph's input spec mismatched on device. The move now
  runs first, for every layer kind.

- Don't drive the runtime for a static-shape multi-modal export: those models
  embed their text in a graph of its own, captured on the prompt, so under
  static shapes it is specialized to the prompt's length and cannot serve the
  loop's 1-token steps (`Guard failed: input_ids.size()[1] == 39`). The
  static-shape exports themselves are still asserted; driving them needs a
  length-generic embedder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_patch_reshape` clones a non-contiguous input before `reshape`/`view`, and
asked `input.is_contiguous()` to decide. Deciding contiguity compares the
strides against the products of the sizes, so on a tensor sized by *unbacked*
symbols — a NaViT packer's data-dependent patch count — the question has no
answer and plain `is_contiguous()` raises `GuardOnDataDependentSymNode`,
failing the whole export on a check meant only to pick a fast path. Nothing in
the model was at fault: the same components export fine through
`torch.export` (and through `run_decompositions` with the ONNX table) once
this patch is out of the way, which is why only the ONNX backend saw it.

`torch._prims_common.is_contiguous_or_false` answers "not provably
contiguous", which lands on the copy — always correct, and the copy is what
makes the view legal in the first place.

Fixes the data-dependent guard failures for lfm2_vl, minicpmv4_6,
paddleocr_vl and modernvbert (20 tests); wavlm / wav2vec2_bert, the models the
patch was written for, keep passing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…er-less VLM

The branch that keeps a captured prefill as its own component was keyed on "a
multi-modal model with no modality getter", which was meant for idefics (whose
gated layers cache the image K/V the prefill writes). It also caught mllama:
that model's cross-attention reads a per-step `cross_attention_mask` no generic
loop grows, so driving it indexed past the mask — an out-of-bounds device-side
assert that poisons the whole xdist worker, turning one real failure into ~148
cascaded ones. Its multi-token variant is skipped for exactly this reason.

Keyed on the layer kind now (`DynamicCrossAttentionLayer`), which is what the
branch was always about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r-tie ids

- `_modality_feed` matched a graph's declared inputs by exact name, so a
  *list-valued* input never reached it: ernie4_5_vl_moe's `temporal_slice_index`
  (the even/odd gather pair) is declared as the flattened leaves
  `temporal_slice_index.0` / `.1`, and ORT rejected the feed for missing them.
  It now asks `_declares`, which already knows a graph naming only the leaves
  still takes the pytree — the same rule the decode feed uses.

- Compare generated ids step by step, and only while eager's own top-2 gap says
  the choice is not a coin flip. A tiny random model routinely has gaps of a few
  1e-3, and how far a backend's kernels drift is model-specific (an ONNX
  chameleon drifts past 1e-3 where kosmos2_5 stays at 6e-5, dynamo lands at
  ~1e-9), so the argmax flips on that drift while saying nothing about
  correctness — which is what made chameleon and kosmos2_5 fail only in some
  orderings. Numeric fidelity is still asserted per component above; half
  precision, whose rounding scale is uniform, still compares scores too.

- Skip git's multi-token capture with the measured evidence: its forward
  corrects for the image tokens only on a single-token step (`position_ids`
  offset by the past length, `attention_mask` widened by the cached image
  tokens, both gated on `seq_len == 1`), so a capture traced at 2 tokens bakes
  those branches off and every later 1-token step runs without them — the
  second step's scores drift ~1e-2 from eager. The single-token capture keeps
  the branches and is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`MllamaModel.forward` indexes `cross_attention_mask[:, :, arange(seq_len) +
past_seen_tokens]`, which only stays in bounds because the model's own
`_update_model_kwargs_for_generation` appends a row per decode step. Any other
caller — an exported runtime, a hand-written loop — indexes past the end, and an
out-of-bounds gather is a device-side assert, not an exception: it kills the
CUDA context, so in a parallel test run one such call turned into ~148 cascaded
failures on the poisoned workers.

The appended row is a copy of the last one, so clamping to the mask's own last
row is exactly what that growth would have produced, and cannot run off the
end. mllama's eager suite is unchanged (the 4 pre-existing static-cache /
multi-gpu failures, one test more passing).

Its export skip now records what this unblocks and what still holds it: the
out-of-bounds is gone and the remaining blocker is routing `pixel_values`
through the runtime, for which the idefics recipe (cross slots declared in
`layer_types`) is verified to work — held back only by mllama's static-cache
compile path, since that layer type is necessarily dynamic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ExecutorchModelRunner` never set `dtype`, so it inherited the base class's
fp32. The generation layer sizes the cache it feeds from that, so every
half-precision export — a grouped-mm MoE, a varlen-attention VLM — was handed
fp32 cache leaves for a program declaring bf16 ones, and the method refused
them while binding its inputs. That is the whole `set_inputs ... error 0x12`
cluster: 109 failures across ~40 MoE models in the ExecuTorch sweep, all from
one missing attribute. The precision is in the `.pte` metadata already; read it
off the first floating-point input.

`set_inputs` failures no longer count as a backend limitation, anywhere.
Binding inputs is our side of the contract — it fails when the runtime hands
the method something it never declared — and treating it as ExecuTorch's
ceiling is what hid this for every MoE model at once. Only `execute()` keeps
that treatment (a portable kernel refusing the step's shapes mid-run), and the
generate drive now absorbs it there too, the way the component checks already
did: deepseek_v3's merged decode reaches exactly that ceiling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
IlyasMoutawwakil and others added 28 commits October 4, 2026 14:32
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rwritten

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	src/transformers/cache_utils.py
#	src/transformers/generation/utils.py
#	src/transformers/masking_utils.py
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…NO runner

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ences

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… export tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cross-attention)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…fig fields

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fresh process had none registered, so torch.export.load could not rebuild the input spec.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The trace specializes an int logits_to_keep, so as an input it was a port nothing read. It now joins the output
flags prepare_for_export pops, and the traced forward receives the flags the config doesn't declare as fixed
values.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…th ONNX

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… declares

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ength, shorter metadata accessors

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
executorch 1.4.1 prunes empty cat inputs with guard_or_true and 1.5.0 writes non-finite constants for flatc
itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eed; honour fft norm on OpenVINO

Each cited model exports and matches eager without them. The OpenVINO fft patch ignored norm, giving wrong values
for ortho/forward.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…thod, one ONNX/OpenVINO compare

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

CI recap

Dashboard: View test results in Grafana
Latest run: 37316589019:1
Result: failure | Jobs: 16 | Tests: 190,729 | Failures: 6 | Duration: 13h 10m

Saving goes through each backend's `save_artifact` (and `save_pretrained`), so `OnnxConfig.output_path`,
`external_data` and `export_params` and `OpenVINOConfig.output_path` were a second, partial way to save; the ONNX
ones were also passed to `torch.onnx.export`, where they do nothing without `f`. `compress_to_fp16` stays: it
compresses the converted model, not the saved file.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants