Dev.to
7/13/2026

Our AI coding bill quietly tripled. Here's what we learned fixing it.
Short summary
An engineering lead shares how their AI coding spend tripled after adopting Claude Code and Codex, and the concrete fixes that brought it under control. Key waste sources included one developer burning tokens on the wrong model, routine tasks hitting frontier models, and unattended test-fix loops. Solutions include enabling prompt caching correctly, routing cheap tasks to cheap models, and capping unattended agent loops.
- •AI coding spend can triple invisibly without usage visibility or model governance
- •Prompt caching is the single biggest cost lever but proxies can silently break it
- •Route routine tasks to cheaper models and cap unattended agent loops with hard spend ceilings
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



