According to Microsoft, the model supports 60 languages and is specifically designed to handle "messy" real-world audio, such as recordings with heavy background noise or overlapping speech.
The release is a key part of Microsoft’s strategy to build its own "frontier-class" models—highly advanced AI systems—to reduce its dependence on partner OpenAI.
By developing in-house technology for specific tasks like transcription, Microsoft can lower its own operational costs for services like Teams and Word while undercutting the pricing of specialized competitors.
Technical benchmarks from Artificial Analysis indicate the model is five to ten times faster than rival products, achieving a high level of efficiency that allows Microsoft to maintain profitability even at the lower price point.
Beyond basic transcription, the model bundles several advanced features into its base price, including speaker diarization—the ability to identify and separate different speakers in a recording—and code switching, which handles conversations that alternate between languages like "Hinglish" or "Spanglish." While the model currently leads several industry benchmarks for accuracy and speed, technical buyers should note that the $0.10 rate is a launch offer.
The model is currently available for testing and deployment via the MAI Playground and Microsoft Foundry environments.