Hook
More breakout videos from this creator.
China just most advanced open source while this new tool let's run a parameter model ON A LAPTOP 8 GIGS OF 8 GIGS OF RAM AND NO G all that happened in AI this week in seconds 6 /kyutai Pocket TTS Supports Python 3.10, 3.11, 3.12, 3.13 and 3.14. Requires PyTorch 2.5+. Does not require the gpu version of PyTorch. Demo | GitHub Repository | Hugging Face Model Card | Tech report | Paper | Documentation Main takeaways Runs on CPU Small model size, 100M parameters Audio streaming Low latency, ~200ms to get the first audio chunk Faster than real-time, ~6x real-time on a CPU of MacBook Air M4 Uses only 2 CPU cores Python API and CLI Voice cloning Multi-language support: english, french, german, portuguese, italian, spanish Can handle infinitely long text inputs Can run on client-side in the browser Additional languages can be added in the future. Using it as a Python library You can try out the Python library on Colab here. Install the package with pip install pocket-tts # or uv add pocket-tts is literally one command 5 reference NOU HERMES SKILLS HUB over 90,000 for your to work across 200 200 CATEGORIES OpenAI ANTHROPIC OpenAI NVIDIA installable into your agent 4 Grok Imagine IMAGE 2.0 A PRECISE TOOL FOR THE IMAGINATION and change plus background 5 reference smart ratio smart resize It's now NO. 2 IN THE WORLD OF IMAGE EDITING 3 Muse Code. Muse Code it's 1st coding agent terminal background across session in parallel your repo Qwen3.8 Qwen3.8-Max Qwen3.8-Max - now available via QwenCloud: 2.4T parameters (95B active), with open weights releasing next week comprehensive improvements across coding, work, research, and long-horizon tasks end-to-end and dependable delivery of complex tasks Call via API on QwenCloud. 2.4 trillion parameters only 95 billion 1 Explore the Epic Story of Life Start Exploring 1 MILLION TOKEN CONTEXT WINDOW native video understanding the highest Chinese model 1 Kimi K3 K3 pulsar research deck with P-Pdot chart and report. kimik3 kimik3-in-c A 2.78-trillion-parameter model. One CPU. 8 GB of RAM. Kimi K3 inference in portable C99. No BLAS. No framework. No GPU. CI passing license Apache 2.0 C99 portable platform Linux x86-64 version 1.0.0 2.78T 1.56 TB checkpoint on disk 8.24 GB peak RSS, measured 176 KB the whole engine KIMI K3 ENGINE MIXTURE-OF-EXPERTS SPARSE ROUTING 16 of 896 fire CPU ONLY FROM SCRATCH IN C FIXED-SIZE STATE context costs nothing extra STREAM THE TRUNK a dial, not a floor RESIDENT IN RAM (TOTAL: 8.24 GB) LM HEAD 163,840 logits ARGMAX (Next Token) DETOKENIZER (BPE -> Text) OUTPUT "Paris." MLA KV CACHE 2.37 MB per position (Latent) EXPERT LRU CACHE 17.56 MB slots (Hot Experts) EXPERT PAYLOAD 25.58 GB per token, 1,472 reads TOTAL ~135 GB / TOKEN BYPASSED PAGE CACHE Direct I/O O_DIRECT No Page Cache EXPERT POOL 896 per layer, 82,432 total, 1.45 TB OFFLINE PREPARATION (ONE-TIME) BUILD PINNED TRUNK PACK & INDEX EXPERTS WRITE SHARDS TO NVMe Crossing Gates Storage Offline Reductions Disabled kimik3 of your disk loading them Generate Assets into memory Paris. The first token out of a 2.78 trillion parameter model running on a CPU in a few gigabytes of RAM is the right answer. The trailing quote and the "The Eiffel" continuation are not a defect: with no chat template a no instruction tuning in the loop, this is a base model continuing text. It has decided it is inside a JSON list of sentences about France, and it is continuing that list. with no 100% 100% OPEN SOURCED them out? AI