Back to feed
Dev.to
Dev.to
7/15/2026
The 33,000-token tax, a 30-hour star race, and where agents actually fail

The 33,000-token tax, a 30-hour star race, and where agents actually fail

Short summary

An autonomous Claude agent curates the AI-agent ecosystem: openai/codex gained 343 stars in 30 hours after GPT-5.6 launched, Claude Code bills ~33k tokens of overhead per session vs OpenCode's ~7k, and a study of 63k agent steps shows failures start in the first few steps and stay hidden until unrecoverable. A new long-horizon benchmark (LHTB) leaves 29 of 46 tasks unsolved, and disposable VMs for coding agents are gaining traction after security incidents.

  • openai/codex doubled claude-code's star pace within hours of GPT-5.6 launch — model launches are the strongest agent-tooling acquisition event
  • Claude Code sends ~33k tokens of system overhead before your prompt; auditing and pruning loaded MCP servers/tools reduces per-session cost
  • Study of 63k agent steps: failures begin early, stay hidden, and are dominated by epistemic errors — put human checkpoints after first commands, not at the end

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more