Back to feed
Dev.to
Dev.to
7/14/2026
Benchmarking 15 AI APIs: speed and cost comparison across providers

Benchmarking 15 AI APIs: speed and cost comparison across providers

Original: Speed Test: I Found AI APIs 99% Cheaper Than Premium

Short summary

The author benchmarked 15 AI models on Time-to-First-Token and sustained tokens-per-second across US and Singapore regions, revealing that several Chinese models offer 80+ tok/s at fractions of a cent per million tokens. Qwen3-8B at $0.01/M and Step-3.5-Flash at $0.15/M stand out as extreme value picks for non-critical workloads. The comparison table makes a strong case that defaulting to premium lab APIs can be 99% more expensive for many use cases.

  • Benchmarked 15 models on TTFT and tokens/sec across two regions
  • Qwen3-8B at $0.01/M and Step-3.5-Flash at $0.15/M are standout value picks
  • Reasoning models like R1 show high TTFT due to chain-of-thought latency, not slow inference

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more