Back to feed
Dev.to
Dev.to
7/1/2026
The original title is: "I Benchmarked DeepSeek, Qwen, Kimi & GLM for 30 Days — The Numbers"

The original title is: "I Benchmarked DeepSeek, Qwen, Kimi & GLM for 30 Days — The Numbers"

Original: I Benchmarked DeepSeek, Qwen, Kimi & GLM for 30 Days — The Numbers

Short summary

Benchmarked 4 Chinese LLM families (DeepSeek, Qwen, Kimi, GLM) across 1,247 prompts over 30 days. DeepSeek V4 Flash offers best price-to-quality ratio at $0.25/M tokens with 60 tokens/sec latency. Kimi excels at reasoning; Qwen spans broadest pricing ($0.01–$3.20/M) but weak correlation between cost and quality.

  • DeepSeek V4 Flash: $0.25/M tokens, 60 tokens/sec, best overall price-to-quality
  • Kimi K2.5: strongest reasoning (4.6/5) but slowest (410ms TTFT)
  • Qwen: broadest range and competent, but median across all benchmarks

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more