Back to feed
Dev.to
Dev.to
7/19/2026
The Harness Effect: orchestration overhead drives 80% of agent costs, not inference

The Harness Effect: orchestration overhead drives 80% of agent costs, not inference

Original: Your Agent Bills While It Waits. Here's the Fix.

Short summary

A peer-reviewed study called 'The Harness Effect' found that model inference is only ~20% of total agent cost; infrastructure, orchestration, tooling, and governance consume the other 80%. Switching orchestration layers cut cost per task by 41% — more than switching models. Re-sent context alone accounts for 62% of agent inference bills, and Goldman Sachs projects a 24x increase in token consumption by 2030. The article breaks down four categories of agent waiting costs: tool latency, human approval gates, retry backoff, and polling loops.

  • Inference is only ~20% of agent cost; orchestration and infrastructure eat 80%
  • Switching orchestration layers reduced cost per task by 41%, more than switching models
  • Re-sent context (system prompts, tool definitions, state history) accounts for 62% of inference bills

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more