On-screen text
🚨 NVIDIA JUST DROPPED
A 0.6B CPU-ONLY
SPEECH MODEL
parakeet.cpp (q8_0) vs NVIDIA
CPU, no GPU
identical output
Well, I don't wish to see it
24 ms
NVIDIA NeMo (PyTorch)
CPU, no GPU
Well, I don't wish to
82 ms
WHAT'S UP, GUYS?
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
SO NVIDIA DROPPED
Well, I don't wish to see it any more, observed Phoebe, turning
away her eyes. It is certainly very like the old portrait.
138 ms
57%
NVIDIA NeMo (PyTorch)
CPU, no GPU
Well, I don't wish to
192 ms
88%
A SPEECH RECOGNITION
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
MODEL WITH ONLY
241 ms
31x realtime
fastest
100%
NVIDIA NeMo (PyTorch)
CPU, no GPU
Well, I don't wish to see it any more, observed Pho
away her eyes. It is certainly
303 ms
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
0.6 BILLION PARAMETERS.
OKAY? IT'S CALLED
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
NEMOTRON 3.5 ASR
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
IT SUPPORTS OVER
SAME MODEL - SAME CPU - IDENTICAL OUTPUT
parakeet.cpp
(q8_0)
NVIDIA NeMo (PyTorch)
60B
3x faster on the same CPU, no GPU
github.com/mudler/parakeet.cpp
241 ms
12x
https://huggingface.co/
nvidia/nemotron-3.5-asr-
streaming-0.6b
40 LANGUAGES,
LOCALAI
from the LocalAI team
github.com/mudler/parakeet.cpp
huggingface.co/mudler/parakeet.cpp
AND IT HAS
REAL TIME STREAMING
OUT.