This flagship release is MIT-licensed, allowing developers to download, customize, and run the system on their own hardware without paying licensing fees.
The launch is significant because it provides frontier-level performance at a fraction of the cost of most proprietary rivals.
Xiaomi’s pricing for its API (a bridge for software to talk to the model) is roughly $0.87 per million output tokens for the Pro version, while a smaller "Flash" version offers similar multimodal capabilities for $0.28 per million.
By achieving high scores on tasks involving long-horizon reasoning and tool use, Xiaomi is positioning these models as high-value infrastructure for "agentic AI"—systems capable of executing complex, multi-step workflows like software engineering and scientific research.
To achieve these results, Xiaomi utilized a massive reinforcement learning (RL) program, a training method where the model learns through trial and error across hundreds of thousands of simulated trajectories.
The company has open-sourced not only the model weights but also the technical frameworks used to train them, including tools designed to prevent "reward hacking," where AI finds shortcuts to solve tasks without actually performing the work correctly.
This release affects the broader AI industry by offering a reproducible blueprint for scaling agent performance using specialized training environments rather than just increasing raw computing power.