Voxtral-Mini-3B-2507 — ONNX (speech-llm)
ONNX export of mistralai/Voxtral-Mini-3B-2507
(© Mistral AI, Apache-2.0) for onnx-asr with the
speech-llm model type.
Load it with the TigreGotico/onnx-asr fork,
branch feat/speech-llm-voxtral — it adds the language_suffix_ids config key that
Voxtral needs. This is the offline Voxtral; the realtime/streaming variant is not part
of this export.
Three-graph layout: encoder.onnx (Whisper-style audio encoder + multi-modal projector,
output already in the language model embedding space), embed_tokens.onnx, and a merged
prefill+decode decoder.onnx with a 30-layer KV cache. The transcription-request token ids
come from mistral-common at export time and are baked into config.json, so no tokenizer is
needed at runtime. fp32 output matches native transformers character-for-character on four
FLEURS clips (2 en, 2 pt). fp32 and int8 graphs are included.
Usage
import onnx_asr
model = onnx_asr.load_model("speech-llm", "path/to/this/repo") # or quantization="int8"
print(model.recognize("audio_16khz.wav"))
print(model.recognize("audio_16khz.wav", language="pt"))
The language argument takes an ISO 639-1 code (en, fr, de, es, it, pt, nl,
hi) or the English name of the language. Without it, Voxtral detects the language itself.
Audio is padded or cut to one 30 s window, the same as every Whisper-preprocessor model in onnx-asr. Cut longer recordings into segments before you send them.
Sizes
| Graph | fp32 | int8 |
|---|---|---|
encoder |
2.65 GB | 0.67 GB |
embed_tokens |
1.61 GB | 0.40 GB |
decoder |
14.45 GB | 3.61 GB |
| total | 18.7 GB | 4.7 GB |
See REPORT.md for the parity numbers, the speed measurements and the graph contract.
License: Apache-2.0, inherited from the source model. Voxtral is a model of Mistral AI.
- Downloads last month
- 314
Model tree for OpenVoiceOS/voxtral-mini-3b-onnx
Base model
mistralai/Voxtral-Mini-3B-2507