Science, Discovery, Tech and Environment · 5 April 2026

Microsoft unveils MAI-Transcribe-1 multilingual speech-to-text model

Exam-focused facts from the 5 April 2026 current affairs briefing.

Key facts

  • MAI-Transcribe-1 is Microsoft’s new speech-to-text model supporting 25 languages and available on Microsoft Foundry.
  • The model achieves the lowest Word Error Rate on the FLEURS 25-language benchmark, outperforming Scribe v2, Whisper-large-V3, GPT-Transcribe, and Gemini 3.1 Flash-Lite.
  • It delivers batch transcription speeds 2.5× faster than the current Microsoft Azure Fast offering.
  • MAI-Transcribe-1 is priced at $0.36 per hour of audio.
  • The model handles noisy environments, background noise, low-quality recordings, and overlapping speech.
  • Offline applications include subtitle generation, podcast transcription, video accessibility, meeting archives, compliance recording, legal discovery, call centre QA, customer insight extraction, searchable audio libraries, and large-scale data pipelines for ML training, search indexing, and summarisation.
  • Online applications comprise meeting transcription, video closed captioning, and dictation.
  • MAI-Transcribe-1 integrates with MAI-Voice-1 text-to-speech and an LLM to build voice agents.