Hook
More breakout videos from this creator.
Giant MoE models. Consumer GPUs. You can now run a 700 billion parameter AI model on two 16 gigabyte graphics cards and I'm gonna show you hit that save button. Let's get in. Let's get in again. We're not shrinking the model. We're doing something very similar to what we did in the last video, but we're streaming from the SSD straight into the GPU this time. This is Pulsar, a new Rust + CUDA engine. Its job is to keep the brain of the model, the current brain, the most used experts and the currently being used layer, only those in your VRAM. So now these giant MoE models that used to require an entire data center, you can now run at home. It's not set up for Kimi K3, but it will run Kimi K 2.7 at 1.2 tokens/sec. GLM 5.2 has 743 billion parameters and it'll run at 2.7 tokens/sec. If you use the 295 billion parameter model, you'll get 7 tokens/sec, which if you know from the last video is way faster than the previous method. This is really cool that it treats your SSD basically as an extension of your GPU. So now we can basically do anything we want. I told you, the tech is gonna continue to improve.