Back to feed
Dev.to
Dev.to
7/17/2026
Moonshot AI's Kimi K3 Is Here: A 2.8 Trillion Parameter Open MoE Model That Pushes Long-Context AI Forward

Moonshot AI's Kimi K3 Is Here: A 2.8 Trillion Parameter Open MoE Model That Pushes Long-Context AI Forward

Short summary

Moonshot AI released Kimi K3, a 2.8 trillion parameter open Mixture-of-Experts model with a 1 million token context window. Key innovations include Kimi Delta Attention for 6.3x faster decoding on long contexts, Attention Residuals for 25% higher training efficiency, and Stable LatentMoE with 2.5x better scaling. Benchmarks show it outperforming competitors on 6 of 35 categories, though it trails in software engineering and reasoning tasks. The model includes upstream vLLM contributions for prefix caching support.

  • 2.8T parameter open MoE model with 1M token context window from Moonshot AI
  • Kimi Delta Attention delivers 6.3x faster decoding; Attention Residuals improve training efficiency 25%
  • Outperforms GPT-5.6 Sol on 6 of 35 benchmark categories but trails in software engineering tasks

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more