Dev.to
7/7/2026

Benchmarking China's Top 4 LLMs: DeepSeek, Qwen, Kimi, and GLM on Cost, Latency, and Quality
Original: I Benchmarked China's Top 4 LLMs — The Numbers Don't Lie
Short summary
A consultant benchmarked DeepSeek, Qwen, Kimi, and GLM across 200 production prompts measuring TTFT, throughput, cost, and blind-rated quality. DeepSeek V4 Flash delivers the best value at $0.25/M output tokens with 58 tokens/sec throughput and quality within measurement noise of GPT-4o, while Qwen3-32B at $0.28/M is the recommended generalist and GLM-4-9B at $0.01/M is cheapest for low-stakes tasks. Kimi's K2.5 at $3.00/M positions as a premium reasoning model that doesn't compete on price.
- •DeepSeek V4 Flash best overall value: $0.25/M output, 58 tok/s, quality near GPT-4o
- •Qwen3-32B recommended generalist at $0.28/M; Qwen3-Omni-30B uniquely handles audio+video+image
- •Kimi K2.5 at $3.00/M is 12x pricier than GLM-4-9B; positioning is reasoning over cost
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



