Cohere Transcribe 2B
This first-party Soniqo page documents Cohere Transcribe 2B from the local speech-swift / speech-core implementation. Hugging Face bundles are linked below after the integration notes.
Internal Page First
Landing cards and docs menus now point here first; source model and bundle links remain available from this page.
At a Glance
| Model | Cohere Transcribe 2B |
|---|---|
| Role | High-accuracy multilingual offline speech-to-text |
| Backend | Native MLX on Apple Silicon |
| Output | Plain-text transcription |
| Languages | 14 languages |
| License | Apache-2.0 |
| Status | Published FP16, INT5, and INT8 bundles; INT5 is the default |
| Source | Cohere Transcribe |
| Swift product | CohereTranscribeASR |
| CLI / runtime | speech transcribe --engine cohere |
Use
The snippet below mirrors the current speech-swift API or command exposed by the repo.
# INT5 is the default.
speech transcribe recording.wav --engine cohere
# Select another published precision and pass a language hint.
speech transcribe recording.wav --engine cohere --model int8 --language de
Model Links
Implementation Notes
- The runtime resamples mono Float32 PCM to 16 kHz and handles long recordings in overlapping chunks.
- INT5 is the practical default: on the validated English FLEURS run it used a 1.62 GiB bundle, 2,582 MiB physical footprint, and 0.0150 mean RTF.
- MLX affine quantization supports 2, 3, 4, 5, 6, and 8 bits, not INT7; choose INT8 when prioritizing quality.
- This is a non-streaming engine. The CLI rejects --stream instead of silently changing behavior.
- The published benchmark covers English read speech and does not claim parity with conversational, telephone, noisy, or speaker-heavy benchmarks.