Science, Discovery, Tech and Environment · 5 April 2026
Microsoft unveils MAI-Transcribe-1 multilingual speech-to-text model
Exam-focused facts from the 5 April 2026 current affairs briefing.
Key facts
- MAI-Transcribe-1 is Microsoft’s new speech-to-text model supporting 25 languages and available on Microsoft Foundry.
- The model achieves the lowest Word Error Rate on the FLEURS 25-language benchmark, outperforming Scribe v2, Whisper-large-V3, GPT-Transcribe, and Gemini 3.1 Flash-Lite.
- It delivers batch transcription speeds 2.5× faster than the current Microsoft Azure Fast offering.
- MAI-Transcribe-1 is priced at $0.36 per hour of audio.
- The model handles noisy environments, background noise, low-quality recordings, and overlapping speech.
- Offline applications include subtitle generation, podcast transcription, video accessibility, meeting archives, compliance recording, legal discovery, call centre QA, customer insight extraction, searchable audio libraries, and large-scale data pipelines for ML training, search indexing, and summarisation.
- Online applications comprise meeting transcription, video closed captioning, and dictation.
- MAI-Transcribe-1 integrates with MAI-Voice-1 text-to-speech and an LLM to build voice agents.