Hook
More breakout videos from this creator.
So much wasn’t possible last week… and I’ll say that next week too late last night I was scrolling on X, I think it's like one of the best places for, like, up to the most up-to-date AI news, whenever something comes out, it's just on Twitter or X, whatever. I don't know what to call it. But the Spark community there is amazing. People are finding tons of different models that work perfectly on single Spark setups, dual Spark setups, quad Spark setups. Pretty incredible. The one that I've started running now, deep, deep seek with a million context tokens, also working with D Spark, which, if you don't know what D Spark is, it's kind of the latest evolution of MTP, multi-token prediction, DFlash, speculative decoding, making the model way more efficient, having tokens cued up for it. But DFlash came out two days ago, and I'm already using it in my setup, and it's the best local setup that I've ever tried. I'm getting like 55 to 62 tokens per second. Instead of just bigger and bigger models, we're just gonna get more and more efficient locally. DeepSeek passes at 1M, the decisive result. 851,210 token needle test: secret=PASS, since=PASS at 851K tokens of context, DSpark correctly retrieved both the mid-context secret cores and the notice. This is the deep-context correctness measurement that has never been published for DSpark at 1M, and it works on your cluster. Full scorecard: Decode, Prefill, Context, 12.7K, 1.446 tok/s, 34.1 tok/s, PASS, secret. 149.5K, 1.764 tok/s, 61.9 tok/s, PASS, secret. 851.2K, 1,070 tok/s (TTFT ~13 sec), 55.6 tok/s, PASS (secret + nonce). Decode at depth (55-62 tok/s) actually beats MTP (~38-50). Spec-decode accepted context at deep context (per-position 2099/4/2/1) - faster, deep, as research predicted but lossless - correctness: solid. So your condition is met: DSpark works at 1M and runs faster than MTP. bigger models, we're just gonna get more and more efficient locally.