
在 MacBook 上全双工运行,并调用真实工具
本地运行 NVIDIA Nemotron VoiceChat 11B 的真实会话:一次对话轮次、句子中途的打断,以及一次写入 Apple 提醒事项的完整 MCP 工具调用。在 M5 Pro 上 RTF 0.92,数据全部不离开本机。

本地运行 NVIDIA Nemotron VoiceChat 11B 的真实会话:一次对话轮次、句子中途的打断,以及一次写入 Apple 提醒事项的完整 MCP 工具调用。在 M5 Pro 上 RTF 0.92,数据全部不离开本机。

说出指令,Android 立即执行:从语音到真实设备操作再到语音回复的完整闭环,在一部手机上以 1.2 GB 内存运行 — Silero VAD、Parakeet STT、FunctionGemma 工具调用与 Pocket TTS。

4 分钟开源库导览:Nemotron Streaming 实时转写、PersonaPlex 本地语音对话、VoxCPM2 48 kHz 声音克隆 — 全部在笔记本上运行。
手机与 Mac 上的 VAD → STT → LLM → TTS · iPhone 约 1.2 GB · S23 约 1.5 GB · 桌面 <4 GB
同一条 VAD → STT → LLM → TTS 语音代理管线跑在 iPhone、Galaxy S23 与 Mac 上,并给出各内存预算的实测值。
每个分组都涵盖多个由 Soniqo 组件拼接而成的子场景。投入音频,即可在本地、实时获得对话、转录或合成语音。
上面的用例流水线全部由这些模型拼接而成。点击组件查看其架构、CLI、Swift API 与基准测试。全部支持 Apple Silicon,多数也支持 Android 与 Linux。
52 langs, RTF 0.06, 4-/8-bit
Native CoreML: 1.40% WER, 6 s model load
MLX INT5: 128K offline context, RTF 0.028
Small, medium, large-v3, and turbo exports
32× real-time on Neural Engine
120M, 25 langs, streaming partials + EOU, ~232 MB on-device
1,672 languages, 300M–7B
14 langs, INT5 default, RTF 0.015
8 langs, INT5 default, RTF 0.074
Word-level timestamps, 80 ms
Streaming with punctuation
9 langs, zero-shot cloning, bf16 / 8-bit
12 Hz codec LM, faster than real-time
48 kHz, 30 langs, voice design + cloning
Zero-shot cloning, emotion + tempo controls, RTF 1.0
99M, 31 langs, 44.1 kHz
CoreML voice cloning, TTFT 0.27s, RTF 0.59
600+ langs, NAR diffusion cloning
Hindi emotion TTS + raw reference cloning
Style markers + zero-shot cloning
Zero-shot cloning in 16 flow steps, RTF 0.57
4B conversational cloning, emotion tags, RTF 0.78
54 voices, iOS-ready, 10 langs
English, ~130 ms streaming TTFA
90-min podcasts / audiobooks
9 langs, 5 baked voices, streaming
IndexTTS2, CosyVoice, Chatterbox Flash, Qwen3-TTS ICL
SpeechBrain ECAPA, 107 languages, 21M
Masked WeSpeaker + native PLDA/VBx, ~32 MB
Community-1 + Pyannote + Sortformer
WeSpeaker / CAM++ for ID
Silero v6.2.1, Pyannote, FireRedVAD
KWS Zipformer, 26× real-time
44.1 kHz stereo, variable-length music
Text → 30 s music, RTF 0.36
Open-Unmix + HTDemucs v4, 4 stems
DeepFilterNet3, 48 kHz real-time
Two-stream echo cancellation, 16 ms latency
Sidon denoise + dereverb → 48 kHz
48 kHz audio super-resolution
Streaming on-device LLM
Structured tool / function-call grammar
Many-to-many translation, 400+ languages
Streaming speech translation, FR/ES/PT/DE → EN
Moshi-family full-duplex speech-to-speech
Asymmetric duplex speech-to-speech, INT5 / INT8