Dev.to
7/29/2026

Let me rewrite this headline following the rules. The original is "Kimi K3: A New Benchmark for Open-Weights Mixture-of-Experts"
Original: Kimi K3: A New Benchmark for Open-Weights Mixture-of-Experts
Short summary
Kimi K3 is a 2.8T parameter open-weights Mixture-of-Experts model with 104B active parameters, claiming 2.5x scaling efficiency over Kimi K2. Key innovations include Kimi Delta Attention for deep transformer coherence and a 1-million-token context window optimized for agentic workloads. The full weights are publicly available, positioning K3 between current open-weights models and frontier proprietary systems like Claude Fable 5 and GPT-5.6 Sol.
- •2.8T parameter open-weights MoE model with 104B active parameters and 1M token context
- •2.5x scaling efficiency over predecessor via architectural refinements including Kimi Delta Attention
- •Full weights released, enabling fine-tuning and inspection for RAG, coding, and agentic use cases
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



