Original caption
China’s latest AI model is getting surprisingly close to the frontier—while charging dramatically less. Moonshot AI’s Kimi K3 is a 2.8-trillion-parameter model, but it doesn’t activate the entire network for every token. Its Mixture-of-Experts architecture routes each token through only 16 of 896 experts. Kimi Delta Attention compresses long-context memory, Attention Residuals improve how information moves between layers, and low-precision quantization helps reduce the cost of running the model. The result? Kimi charges $3 per million input tokens and $15 per million output tokens. GPT‑5.6 charges $5 and $30, while Claude Fable 5 charges $10 and $50. K3 probably doesn’t match those models in real-world use yet—and Moonshot admits the overall experience still trails both. But it shows that Chinese labs are learning to convert limited compute into more capability. Export controls may raise their costs, but they haven’t locked China out of the frontier. And if near-frontier intelligence keeps getting cheaper, the biggest American labs may have a much smaller moat than their valuations suggest. #KimiK3 #opensource #AInews #airesearch