Artificial Intelligence

Mistral AI Releases Voxtral Transcribe 2, Targets Speed and Cost Improvements

Paris: Mistral AI, high-performance AI company introduced Voxtral Transcribe 2, the company’s second-generation speech-to-text models aimed at delivering faster, more accurate, and lower-cost transcription. The release includes two variants, Voxtral Mini Transcribe V2 for batch processing and Voxtral Realtime for live audio applications with latency configurable down to sub-200 milliseconds.

credits: mistral ai

In a post on X, the company described it as “next-gen speech-to-text” offering state-of-the-art transcription, speaker diarization, and sub-200ms real-time latency.

Read more: India’s Sarvam AI Beats Google Gemini, ChatGPT on OCR Benchmarks

Also read: Microsoft Takes on AI Rivals With 3 New Foundational Models

The Paris-based AI firm said the new models are designed to compete directly with leading transcription services while significantly reducing costs. Voxtral Mini Transcribe V2 is priced at $0.003 per minute for batch jobs, which the company says is roughly one-fifth the cost of competing offerings such as ElevenLabs’ Scribe v2.

Read more: Meta Unveils Muse Spark As Its First Proprietary AI Model

According to Mistral’s internal benchmarks, the models deliver about a 4%-word error rate on the FLEURS dataset, outperforming several well-known transcription systems while also processing audio up to three times faster than some rivals. The company added that the real-time model can match batch-level accuracy at higher latency settings suitable for live subtitling, while lower latency modes introduce only a small increase in error rates.

Voxtral Mini Transcribe V2 includes features such as speaker diarization, word-level timestamps, and context biasing that allow users to add up to 100 domain-specific terms for improved accuracy. Voxtral Realtime, meanwhile, is built for voice agents, live captioning, and call-center automation.

Notably, Voxtral Realtime is released under the Apache 2.0 license, allowing organizations to deploy it on-premises without relying on external APIs. With a 4-billion-parameter footprint capable of running on edge devices, the models are positioned for industries with strict data-privacy requirements, including healthcare and finance.

Tea4Tech Team

The Tea4Tech News Desk is our collaborative editorial team dedicated to bringing you the latest breaking news, industry updates, and trending stories from the world of technology. Comprised of experienced journalists, tech enthusiasts, and digital researchers, the News Desk works around the clock to curate, verify, and deliver timely content that keeps our readers informed and ahead of the curve.

Recent Posts

Microsoft Launches its First Cybersecurity AI Model, MAI-Cyber-1-Flash

San Francisco: Microsoft unveiled its first cybersecurity-focused AI model, MAI-Cyber-1-Flash at an event in San…

2 weeks ago

Facebook Blocks PM Modi’s NEET Video, Meta Calls It an Error

New Delhi: The Indian government was caught off guard when social media giant Meta blocked…

2 weeks ago

Moonshot’s Kimi K3 Becomes Largest Open-Weight AI Model Ever

BEIJING: Moonshot AI releases Kimi K3, a 2.8 trillion-parameter model that instantly becomes the largest…

2 weeks ago

Fireworks AI Hits $17.5B as Enterprises Flee Frontier API Prices

SAN MATEO, Calif.: Fireworks AI closes a $1.505 billion Series D at a $17.5 billion…

2 weeks ago

AegisAI Lands $36M Series A to Secure the Agentic Enterprise

SAN FRANCISCO: AegisAI raises $36 million in Series A funding to secure enterprises against threats…

2 weeks ago

Paper Raises $34M as AI Coding Agents Redraw Design’s Borders

SAN FRANCISCO: Paper raises $34 million in Series A funding led by Accel and ICONIQ…

2 weeks ago