Voxtral Mini 3B 2507

توثق هذه الصفحة من Soniqo نموذج Voxtral Mini 3B 2507 كما هو منفذ في speech-swift / speech-core. روابط Hugging Face موجودة أدناه بعد ملاحظات الدمج.

الصفحة الداخلية أولا

بطاقات الصفحة الرئيسية وقوائم الوثائق تشير إلى هذه الصفحة أولا؛ وتبقى روابط النموذج والحزم داخلها.

لمحة سريعة

النموذجVoxtral Mini 3B 2507
الدورHigh-accuracy multilingual offline speech-to-text
BackendNative MLX on Apple Silicon
الإخراجPlain-text transcription
اللغاتEnglish, French, German, Spanish, Italian, Portuguese, Dutch, and Hindi
الرخصةApache-2.0
الحالةPublished FP16, INT5, and INT8 bundles; INT5 is the default
المصدرMistral Voxtral Mini 3B 2507
منتج SwiftVoxtralASR
CLI / runtimespeech transcribe --engine voxtral

الاستخدام

المقتطف أدناه يطابق API أو الأمر الحالي في speech-swift.

# INT5 is the default.
speech transcribe recording.wav --engine voxtral

# Select another published precision and pass a language hint.
speech transcribe recording.wav --engine voxtral --model int8 --language fr

# --model also accepts a Hugging Face model ID or a local directory.
speech transcribe recording.wav --engine voxtral --model /models/voxtral/int5

قياس الأداء

قيس في 2026-07-22 على 80 نطقًا مقروءًا بالإنجليزية من FLEURS (759.56 ثانية إجمالًا)، على Apple M5 Pro بذاكرة 48 غيغابايت ونظام macOS 26.5.2. نُفذ كل متغير في عملية مستقلة من بناء release.

VariantBundleWERΔ WERMean RTFOverall ×RTFootprint
FP168.71 GiB4.633%0.13058.05×10,568 MiB
INT53.77 GiB4.744%+0.110 pp0.073914.47×6,012 MiB
INT85.18 GiB4.578%-0.055 pp0.090611.79×7,233 MiB

البصمة الفعلية هي مقياس الذاكرة الموحدة المهم عند النشر، لأن MLX قد يربط ملفات الأوزان بالذاكرة. وFLEURS كلام مقروء بالإنجليزية: لا تدّعي هذه الأرقام تكافؤًا على الصوت الحواري أو الهاتفي أو الصاخب أو متعدد المتحدثين.

روابط النموذج

ملاحظات التنفيذ