Voxtral Mini 3B 2507
यह Soniqo पेज स्थानीय speech-swift / speech-core implementation में Voxtral Mini 3B 2507 को दस्तावेज़ करता है। Hugging Face bundle links integration notes के बाद दिए गए हैं।
पहले आंतरिक पेज
Landing cards और docs menus पहले इसी पेज पर आते हैं; source model और bundle links यहीं उपलब्ध रहते हैं।
सारांश
| मॉडल | Voxtral Mini 3B 2507 |
|---|---|
| भूमिका | High-accuracy multilingual offline speech-to-text |
| Backend | Native MLX on Apple Silicon |
| Output | Plain-text transcription |
| भाषाएँ | English, French, German, Spanish, Italian, Portuguese, Dutch, and Hindi |
| लाइसेंस | Apache-2.0 |
| स्थिति | Published FP16, INT5, and INT8 bundles; INT5 is the default |
| Source | Mistral Voxtral Mini 3B 2507 |
| Swift product | VoxtralASR |
| CLI / runtime | speech transcribe --engine voxtral |
उपयोग
नीचे का snippet मौजूदा speech-swift API या command से मेल खाता है।
# INT5 is the default.
speech transcribe recording.wav --engine voxtral
# Select another published precision and pass a language hint.
speech transcribe recording.wav --engine voxtral --model int8 --language fr
मॉडल लिंक
implementation notes
- The audio frontend resamples mono Float32 PCM to 16 kHz and packs up to 30 seconds of audio per request.
- INT5 is the default: on the validated English FLEURS run it used a 3.77 GiB bundle, 6,012 MiB physical footprint, and 0.0739 mean RTF.
- The decoder projects only the final prompt state through the 131,072-token language-model head, reducing quantized-model RTF without changing transcripts.
- MLX has no INT7 affine kernel; use INT8 as the supported higher-quality option.
- This is a non-streaming engine. Benchmark results cover English read speech, not conversational or speaker-heavy audio.