Microsoft Unveils Real-Time Streaming Transcription Model
On October 2, IT Home reported a significant technological advancement. Yesterday, October 1, Microsoft released an official announcement declaring the launch of their inaugural real-time streaming voice transcription model, MAI-Transcribe-2-Streaming. This formidable tool can continuously output text precisely as the speaker talks. Furthermore, it comprehensively covers sixty distinct languages and inherently supports automatic language detection.
Pricing and Accessibility
Regarding financial implications, the MAI-Transcribe-2-Streaming model currently resides in a promotional phase. The price stands at a competitive 0.54 dollars per hour. This translates to roughly 9 dollars per 1,000 minutes. Developers can eagerly experience or integrate this technology through various accessible channels. These include the Microsoft Foundry, the MAI Playground, and OpenRouter.
IT Home previously noted that Microsoft had already launched the non-streaming MAI-Transcribe-2. During the Azure Speech public preview phase, its price was an incredibly modest 0.10 dollars per hour. The streaming iteration returns highly relevant results in real time, making it utterly ideal for live conversations and immediate subtitling. Conversely, the non-streaming version processes the entire file before returning an exhaustive result. Therefore, it excels at complete transcription tasks involving recorded audio files, comprehensive meeting minutes, and detailed clinical documentation.
Exceptional Low-Latency Performance
Regarding latency, MAI-Transcribe-2-Streaming will swiftly generate the initial batch of “partial text” within roughly 100 milliseconds after receiving the audio input. Subsequently, it continuously refines the transcription based on expanding contextual information. Finally, it rapidly submits a steadfastly stable result immediately after the sentence concludes.
Microsoft proudly claims that during their rigorous real-time evaluations, text can manifest a mere 320 milliseconds after the speech actually occurs. In independent testing conducted by Artificial Analysis, the model achieved a final transcription latency of precisely 0.13 seconds. Furthermore, it recorded an astonishingly low final word error rate of 2.50 percent. Consequently, it emphatically claimed the undisputed top position among twenty-eight competing streaming speech-to-text models.
Transforming Real-World User Experiences
In actual practical application, voice-activated software no longer needs to wait for a user to finish their sentence. Instead, it can proactively comprehend the underlying intent, initiate complex reasoning, or summon necessary tools prematurely. For example, a customer service intelligent agent can begin identifying a specific problem while the caller is still speaking. Similarly, live subtitling systems can approach a genuinely seamless “see what is said” experience.
Regarding comprehensive functionality, MAI-Transcribe-2-Streaming flawlessly supports sixty languages. It also possesses a robust, continuous automatic language detection capability. Consequently, users absolutely do not need to manually specify a language prior to commencing a session. The intelligent system effortlessly detects subtle language shifts based entirely on the spoken content.
Leading the Industry Benchmarks
Concerning benchmark performance, Artificial Analysis published their definitive streaming speech-to-text leaderboard on September 28. MAI-Transcribe-2-Streaming achieved an exemplary final word error rate of 2.50 percent. This exceptional score comfortably undercuts Grok Voice Transcribe 2.0 Streaming’s 2.73 percent and ElevenLabs Scribe v2 Realtime’s 3.59 percent. Finally, Microsoft emphatically stressed that their model successfully ranked first in both crucial metrics: final transcription accuracy and initial partial transcription speed.











