Dev.to
6/30/2026

The original headline is 9 words: "How Coinbase Halved AI Costs Without Limiting Engineer Access"
Original: Coinbase Cut Its AI Spend in Half Without Throttling Engineers - Here's the Playbook
Short summary
Coinbase slashed AI spend 50% by defaulting engineers to cheaper open-weight models (GLM 5.2 at $1.40/M tokens vs. Opus 4.8 at $5/M), boosting cache hit rates from 5% to 60%, and routing by task complexity—without access caps. The five-tactic playbook (gateway defaults, task routing, aggressive caching, lean context, spend visibility) required no hard restrictions. This mirrors shifts at Snowflake and Lindy, signaling direct revenue pressure on Anthropic and OpenAI.
- •Coinbase reduced AI infrastructure costs 50% via five independent levers: gateway defaults to cheaper models, task-based routing, cache optimization (5→60% hit rate), lean context, and per-engineer spend visibility
- •Open-weight Chinese models (GLM 5.2, Kimi, DeepSeek) cost 3–6x less than Opus/GPT-4, driving enterprise-wide adoption signals beyond one-off experiments
- •Three immediately actionable tactics: audit cache hit rates (target 20%+), classify and route by task complexity, default to cheaper models with engineer opt-up
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



