MiMo-V2.5 is an open-weight Mixture-of-Experts model from Xiaomi for agentic work and multimodal understanding. It runs 15B active parameters out of 310B, takes a 1M-token context, and is natively omnimodal: text, image, video, and audio in, text out.
It's the only open-weight model in the is*ai lineup that reads audio, and its focus is sharp perception (charts, documents, video) alongside frontier-level agentic ability.
With is*smart, MiMo-V2.5 is already deployed and optimized inside is*hosting, so the hardware a 310B multimodal model needs is handled. Here is where it earns its place.
The audio encoder is what sets it apart. It can work over meeting recordings, calls, and voice notes alongside text, images, and video, which fits tasks that mix formats: summarizing a recorded call against a slide deck, or answering questions across a document and its audio together.
MiMo-V2.5 reads charts, tables, scanned documents, and screenshots and reasons over what it sees. That makes it a fit for visual question answering, chart and figure analysis, and pulling structured meaning out of documents that aren't plain text.
Many agent tasks stall because the model can't interpret what's on the screen. MiMo-V2.5 pairs perception with tool use, so an agent can read a dashboard, interpret a chart, and act on it in the same loop instead of handing off to a separate vision step.
Because it's efficient for its size, perception-heavy pipelines that would otherwise need a frontier-priced model can run here at a lower cost per task, without dropping to a model that can't see or hear.
MiMo-V2.5 is built around perception. When you don't need it, a leaner model is a better match:
MiMo-V2.5 reads image, video, and audio but outputs text only, so it understands media rather than generating it. It's a large model, so perception-heavy calls cost more than a Flash-tier text model. For plain text work, a lighter model is cheaper. Match it to tasks that actually use its perception or audio and it earns its keep; otherwise you're paying for capability you won't touch.
MiMo-V2.5 is MIT-licensed, so the weights are free to run and build on. Through is*smart, it's already hosted and optimized inside the is*hosting infrastructure, so your files, prompts, and media stay in the environment: no third-party APIs, no data leaving, and no vendor lock-in. You get an omnimodal, agentic model without sourcing the GPUs a 310B model needs.
Subscribe to is*smart to get instant access to MiMo-V2.5 and put audio, vision, and agentic reasoning behind a single model.