Paris: Mistral AI, high-performance AI company introduced Voxtral Transcribe 2, the company’s second-generation speech-to-text models aimed at delivering faster, more accurate, and lower-cost transcription. The release includes two variants, Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live audio applications with latency configurable down to sub-200 milliseconds.
In a post on X, the company described it as “next-gen speech-to-text” offering state-of-the-art transcription, speaker diarization, and sub-200ms real-time latency.
Read more: India’s Sarvam AI Beats Google Gemini, ChatGPT on OCR Benchmarks
Also read: Microsoft Takes on AI Rivals With 3 New Foundational Models
The Paris-based AI firm said the new models are designed to compete directly with leading transcription services while significantly reducing costs. Voxtral Mini Transcribe V2 is priced at $0.003 per minute for batch jobs, which the company says is roughly one-fifth the cost of competing offerings such as ElevenLabs’ Scribe v2.
Read more: Meta Unveils Muse Spark As Its First Proprietary AI Model
According to Mistral’s internal benchmarks, the models deliver about a 4%-word error rate on the FLEURS dataset, outperforming several well-known transcription systems while also processing audio up to three times faster than some rivals. The company added that the real-time model can match batch-level accuracy at higher latency settings suitable for live subtitling, while lower latency modes introduce only a small increase in error rates.
Voxtral Mini Transcribe V2 includes features such as speaker diarization, word-level timestamps, and context biasing that allow users to add up to 100 domain-specific terms for improved accuracy. Voxtral Realtime, meanwhile, is built for voice agents, live captioning, and call-center automation.
Notably, Voxtral Realtime is released under the Apache 2.0 license, allowing organizations to deploy it on-premises without relying on external APIs. With a 4-billion-parameter footprint capable of running on edge devices, the models are positioned for industries with strict data-privacy requirements, including healthcare and finance.
California: Meta introduced its first pair of camera-free AI glasses on Wednesday at its Connect…
Bengaluru: Anthropic has combined Claude's chat and Cowork into one interface, so users no longer…
California: OpenAI has reportedly acquired Glass Imaging, a startup specializing in AI-enhanced smartphone camera technology,…
New Delhi: OpenAI has launched GPT-6 Astra, its latest and most advanced AI model, with…
New Delhi: Adobe has acquired Rilo, an India based marketing intelligence startup, in a deal…
New York: Anthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, its latest AI…