Hook
More breakout videos from this creator.
This is it. This is Jalapeno. This is Jalapeno, you very graciously allowed me to get my hands on it. Ed Ludlow Bloomberg News allowed me to get my hands on it. And again, the whole emphasis is on throughput and ultra-low latency. And actually, in the ultra-low latency piece, you'd associate that with SRAM based systems. This is both HBM and SRAM. Talk a little bit about its design and architecture, but why that ultra-low latency matters? So, it is a HBM-based design, it's HBM4. Richard Ho OpenAI Head of Hardware It's one of the first, you know, processing to go to volume with HBM4, and we're right behind Nvidia's products, and it is very low-latency. The reason how we got it was really taking a blank slate approach to this design. So my team arrived at OpenAI, and we basically were able to work with the research team to understand where the bottlenecks were in large language models. And with a blank slate design, we figured out that you had to reduce your data movement a lot, and we figured out a way to do it. And so this architecture is different from GPUs, is different from TPUs and other chips in that nature. It really reduces the data movement, it optimizes the algorithm, and then it basically is able to get this very-low latency, which matters for the agentic workloads.