Voxtral-Mini-3B-2507 — ONNX (speech-llm)

ONNX export of mistralai/Voxtral-Mini-3B-2507 (© Mistral AI, Apache-2.0) for onnx-asr with the speech-llm model type.

Load it with the TigreGotico/onnx-asr fork, branch feat/speech-llm-voxtral — it adds the language_suffix_ids config key that Voxtral needs. This is the offline Voxtral; the realtime/streaming variant is not part of this export.

Three-graph layout: encoder.onnx (Whisper-style audio encoder + multi-modal projector, output already in the language model embedding space), embed_tokens.onnx, and a merged prefill+decode decoder.onnx with a 30-layer KV cache. The transcription-request token ids come from mistral-common at export time and are baked into config.json, so no tokenizer is needed at runtime. fp32 output matches native transformers character-for-character on four FLEURS clips (2 en, 2 pt). fp32 and int8 graphs are included.

Usage

import onnx_asr
model = onnx_asr.load_model("speech-llm", "path/to/this/repo")  # or quantization="int8"
print(model.recognize("audio_16khz.wav"))
print(model.recognize("audio_16khz.wav", language="pt"))

The language argument takes an ISO 639-1 code (en, fr, de, es, it, pt, nl, hi) or the English name of the language. Without it, Voxtral detects the language itself.

Audio is padded or cut to one 30 s window, the same as every Whisper-preprocessor model in onnx-asr. Cut longer recordings into segments before you send them.

Sizes

Graph fp32 int8
encoder 2.65 GB 0.67 GB
embed_tokens 1.61 GB 0.40 GB
decoder 14.45 GB 3.61 GB
total 18.7 GB 4.7 GB

See REPORT.md for the parity numbers, the speed measurements and the graph contract.

License: Apache-2.0, inherited from the source model. Voxtral is a model of Mistral AI.

Downloads last month
314
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenVoiceOS/voxtral-mini-3b-onnx

Quantized
(31)
this model

Collections including OpenVoiceOS/voxtral-mini-3b-onnx