Loading video…
Hook

Their other posts in the index, biggest breakout first.
Nvidia just open sourced a computer vision model that's 10x faster than top models and it's kind of insane. Most vision models today predict bounding boxes step by step, corner by corner, token by token. But this one changes that. It uses something called parallel box decoding. Instead of predicting pieces, it predicts the entire box at once. That's why it's up to 10x faster than models like Qwen 3 VL. And they trained it on massive data, over 100 million queries and hundreds of millions of boxes. It's fully open sourced on Hugging Face and GitHub.