Voxtral Mini 4B Realtime: Mistral's real-time speech recognition model
Mistral published Voxtral-Mini-4B-Realtime-2602 on Hugging Face, a 4B parameter real-time automatic speech recognition model built on the Ministral-3-3B base. It uses vLLM as its inference library and supports 16 languages. An associated arXiv paper is referenced.
PUBLISHED2026-01-21
OBSERVED2026-08-11
AGE7mo
SOURCES1
- Pipeline: automatic-speech-recognition
- Library: vLLM (voxtral_realtime tag)
- Base model: mistralai/Ministral-3-3B-Base-2512
- Languages: en, fr, es, de, ru, zh, ja, it, pt, nl, ar, hi, ko
- Referenced paper: arxiv:2602.11298
- Repository created 2026-01-21, last modified 2026-03-11
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.