# voxtral-mistral-ai-speech-understanding-2025-07-15

## Veille

Voxtral — open source speech understanding models from Mistral AI: multilingual transcription, audio Q&A, Apache 2.0 license (mistral.ai)

## Titre Article

Voxtral | Mistral AI

## Date

2025-07-15

## URL

https://mistral.ai/news/voxtral

## Keywords

Voxtral, Mistral AI, speech understanding, open source, speech recognition, language models, transcription, multilingual, Q&A, automatic summarization, function-calling, AI, voice interface, Apache 2.0, 24B model, Mini 3B

## Authors

Mistral AI

## Ton

**Profile:** Technical product launch | Institutional | Informative-Promotional | Expert

Mistral AI adopts an announcement voice positioning Voxtral as a major open source release in voice models. The emphasis on the Apache 2.0 license and multilingual capabilities speaks to the open source community. Precise technical language (24B model, Mini 3B, function-calling, Q&A) demonstrates the breadth of capabilities. The confident, professional tone is typical of a European AI startup positioning itself against OpenAI/Google. The capability-centered structure, with availability details, facilitates developer adoption. Typical of AI lab product launches (Anthropic, Cohere style) targeting a developer community committed to open source with commercial-grade capabilities.

## Pense-betes

- **2 sizes**: 24B variant (production), 3B Mini (local/edge), under **Apache 2.0 license**
- **API access** + open source download
- **Bridges the gap** between high-error-rate open source ASR and costly proprietary APIs
- **State-of-the-art accuracy** + native semantic understanding at **less than half the price** of proprietary solutions
- **Long context**: up to **30-40 minutes of audio** (30 min transcription, 40 min understanding)
- **Built-in Q&A and summarization**: direct audio querying without chaining ASR + LLM
- **Natively multilingual**: automatic language detection, **8+ languages** (English, Spanish, French, Portuguese, Hindi, German, Dutch, Italian)
- **Function-calling from voice**: spoken intentions → actionable system commands
- **Voxtral Mini Transcribe**: **less than half the price of OpenAI Whisper**, and outperforms it
- **Voxtral Small matches ElevenLabs Scribe** for **less than half the price**
- **Benchmarks**: outperforms Whisper large-v3, competitive with GPT-4o mini Transcribe and Gemini 2.5 Flash
- **Roadmap**: speaker diarization, audio annotations (age, emotion), word-level timestamps, non-speech audio recognition

## RésuméDe400mots

Mistral AI unveils **Voxtral**, an innovative suite of open source speech understanding models designed to transform human-machine interaction through voice. Considering voice as humanity's original interface, Voxtral aims to overcome the limitations of current systems, whether unreliable or proprietary, by offering robust, multilingual, and deeply intelligent voice tools. The models are available in **two versions**: a 24B variant built for production-scale applications, and a more compact 3B variant, **Voxtral Mini**, suited to local and edge deployments. Both are released under the **permissive Apache 2.0 license** and accessible via the Mistral AI API, with **Voxtral Mini Transcribe**, optimized for transcription, offering unmatched cost/latency efficiency.

**Two-tier architecture and cost advantage**

Voxtral stands out by bridging the gap between open source ASR systems with high error rates and costly proprietary APIs. It delivers **state-of-the-art accuracy and native semantic understanding** within an open framework, at **less than half the price** of comparable proprietary solutions. This cost efficiency makes high-quality voice intelligence accessible and controllable at scale for a wide range of applications.

**Advanced capabilities beyond transcription**

The models offer several advanced capabilities beyond simple transcription. They support **long audio contexts**, up to **30 minutes for transcription and 40 minutes for understanding**, enabling full processing of extended conversations or recordings. A standout feature: **built-in Q&A and summarization**, which allow direct querying of audio content or generation of structured summaries **without chaining separate ASR and language models**. Voxtral is natively multilingual, with automatic language detection and state-of-the-art performance across many widely used languages: **English, Spanish, French, Portuguese, Hindi, German, Dutch, Italian**. It also enables **function-calling directly from voice**, translating users' spoken intentions into actionable system commands. The models retain the strong text understanding capabilities of their foundation, the Mistral Small 3.1 language model.

**Competitive benchmarks**

Benchmark results underscore Voxtral's superior performance. It **outperforms Whisper large-v3 overall**, the reference open source transcription model, and surpasses GPT-4o mini Transcribe and Gemini 2.5 Flash on various tasks. **Voxtral Small achieves state-of-the-art results** on short-form English and Mozilla Common Voice, demonstrating strong multilingual capabilities, and **matches ElevenLabs Scribe** for premium use cases at a significantly reduced cost. **Voxtral Mini Transcribe outperforms OpenAI Whisper at less than half the price**.

**Roadmap and vision**

Mistral AI plans to introduce **speaker diarization, audio annotations (age, emotion), word-level timestamps, and non-speech audio recognition**. The company is expanding its audio team and encourages developers to integrate Voxtral via local download on Hugging Face, via the API, or by trying it in Le Chat's voice mode, with advanced enterprise features: private deployment, domain-specific fine-tuning, and dedicated integration support.

## GrapheDeConnaissance

- Mistral AI —publie→ Voxtral (TECHNOLOGIE, 0.99)
- Voxtral —est_basé_sur→ Mistral Small 3.1 (TECHNOLOGIE, 0.97)
- Voxtral —utilise→ licence Apache 2.0 (CONCEPT, 0.99)
- Voxtral Small —surpasse→ Whisper large-v3 (TECHNOLOGIE, 0.96)
- Voxtral Small —surpasse→ GPT-4o mini Transcribe (TECHNOLOGIE, 0.95)
- Voxtral Small —surpasse→ Gemini 2.5 Flash (TECHNOLOGIE, 0.95)
- Voxtral Small —concurrence→ ElevenLabs Scribe (TECHNOLOGIE, 0.93)
- Voxtral Mini Transcribe —surpasse→ OpenAI Whisper (TECHNOLOGIE, 0.94)
- Voxtral —permet→ function-calling depuis la voix (CONCEPT, 0.97)
- Voxtral —permet→ Q&A et résumé audio intégrés (CONCEPT, 0.97)
- Voxtral —observé_dans→ Hugging Face (TECHNOLOGIE, 0.96)
- Le Chat —utilise→ Voxtral (TECHNOLOGIE, 0.95)
- Mistral AI —collabore_avec→ Inworld (ORGANISATION, 0.88)

---
Canonical: https://www.thekb.eu/en/fiches/voxtral-mistral-ai-speech-understanding-2025-07-15/
