Dev.to
7/1/2026

The original title is: "I Benchmarked DeepSeek, Qwen, Kimi & GLM for 30 Days — The Numbers"
Original: I Benchmarked DeepSeek, Qwen, Kimi & GLM for 30 Days — The Numbers
Short summary
Benchmarked 4 Chinese LLM families (DeepSeek, Qwen, Kimi, GLM) across 1,247 prompts over 30 days. DeepSeek V4 Flash offers best price-to-quality ratio at $0.25/M tokens with 60 tokens/sec latency. Kimi excels at reasoning; Qwen spans broadest pricing ($0.01–$3.20/M) but weak correlation between cost and quality.
- •DeepSeek V4 Flash: $0.25/M tokens, 60 tokens/sec, best overall price-to-quality
- •Kimi K2.5: strongest reasoning (4.6/5) but slowest (410ms TTFT)
- •Qwen: broadest range and competent, but median across all benchmarks
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



