Skip to content

Document MiMo v2.5 video/audio input handling and video sampling controls #540

Description

@em2275

Could you clarify the supported audiovisual input path for xiaomi-mimo-v2-5 on POST /api/v1/chat/completions?

The published chat schema documents an MP4 data URI in video_url.url and raw base64 WAV in input_audio.data with format: "wav". Xiaomi's direct MiMo documentation additionally describes fps and media_resolution on video content parts. These controls are not documented in Venice's chat video schema.

  1. Does Venice pass MP4 video and a separate WAV together to MiMo's native visual and audio inputs? Is the MP4's own soundtrack also processed, or should clients supply the separate WAV?
  2. How are source duration, frame timestamps and audio start offsets represented to the model? Are frames resampled or arranged into composite images?
  3. Are MiMo's fps and media_resolution accepted and forwarded by Venice? If so, please document their exact placement, defaults, limits and effect on token accounting.
  4. What WAV sample rate/channel layout is recommended, and is there response metadata that confirms which modalities were processed?

A minimal documented request example would help clients diagnose missed speech or incorrect event order without guessing unsupported parameters. This is a documentation question; no account credentials, user media or private conversation content are included.

References:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions