Microsoft Unveils MAI-Transcribe-2 Streaming Voice AI, Beats ElevenLabs
Serge Bulaev
Microsoft has launched MAI-Transcribe-2-Streaming, a voice AI model that may offer higher accuracy and lower cost than other options like ElevenLabs. Early data suggests it has a low word error rate of 2.5 percent and very short delay, supporting 60 languages. The model appears to be the top performer in recent vendor benchmarks, but independent checks are still limited. Microsoft says this system could make live transcription faster and more accurate across its products.

Microsoft has launched MAI-Transcribe-2-Streaming, a powerful voice AI model designed for simultaneous, real-time transcription and speech generation. The company asserts that this new model provides higher accuracy, lower latency, and reduced costs compared to rivals like ElevenLabs and is integrating it across its product ecosystem.
What Makes MAI-Transcribe-2-Streaming Unique?
Microsoft's MAI-Transcribe-2-Streaming features a unified architecture that performs transcription and speech generation simultaneously, not sequentially. This integrated approach reportedly improves speed and quality, reducing the delay between a user speaking and the AI responding, which creates a more natural conversational experience for users.
The key differentiator for MAI-Transcribe-2-Streaming is its unified architecture that handles both transcription and speech generation in a single, simultaneous process. According to Microsoft AI CEO Mustafa Suleyman, this tandem operation improves both quality and speed, contrasting with typical pipelines where these tasks are sequential. This design aims to eliminate delays between understanding and responding for more fluid conversations.
How Does Its Accuracy Compare to ElevenLabs and Google?
Microsoft supports its performance claims with benchmark data, stating its model is faster and more accurate than key competitors. According to the company, MAI-Transcribe-2-Streaming delivers significantly faster performance than competing models while maintaining higher accuracy.
Microsoft reports improved Word Error Rate (WER) performance with low latency compared to existing solutions. However, these benchmarks still await broader independent verification.
Where Will MAI-Transcribe-2-Streaming Be Used?
Microsoft is deploying this technology across its product portfolio to replace external speech models with its own vertically integrated solution. Key integration targets include:
| Product | Application |
|---|---|
| Microsoft Teams | Real-time meeting transcription and speech generation |
| Dragon Copilot | Clinical documentation and other healthcare voice workflows |
| Copilot assistant | Enhanced voice mode for Microsoft's primary AI assistant |
How Can Developers Access the Model?
Microsoft is offering several access points for developers during its public preview period.
Current access methods:
- Microsoft Foundry: The primary channel for API access.
- MAI Playground: An in-browser environment for testing and demos.
- APIs and SDKs: An OpenAI Realtime-compatible WebSocket API and the native Azure Speech SDK.
- Partner Platforms: Integrations with various development platforms.
Important Limitation: The model is currently in public preview, which means it comes with no SLA and is not recommended for production workloads at this time.
Is It Really the "Cheapest" Voice AI?
Microsoft has marketed MAI-Transcribe-2 as offering competitive pricing for speech recognition. The company has announced promotional pricing for early adopters, with support for multiple languages.
However, this is a promotional rate with a set expiration, and standard pricing has not been released. Organizations should perform direct cost comparisons based on their specific workloads, as total cost depends on volume, concurrency, and other factors.
Key Specifications at a Glance
| Attribute | Claimed Performance |
|---|---|
| Streaming WER | Improved performance |
| Latency | Low latency |
| Speed vs. Competitors | Significantly faster |
| Languages Supported | Multiple languages |
| Pricing | Competitive rates |
| Availability | Public Preview (No SLA) |
Microsoft's release represents a significant bet on the vertical integration of voice AI. While the claimed performance advantages hold promise, the model's true competitive position will be determined by independent, production-scale evaluations against established players.