The Cost Control Plane

Measure AI cost per accepted outcome, normalize usage telemetry, evaluate caching and routing, and enforce budgets without assuming a universal optimization order.

The Cost Control Plane — AI

In a hypothetical comparison, an autonomous coding agent can produce a $20 patch, spend $200 discovering its first idea was wrong, or burn $2,000 re-reading the same repository while six subagents argue about it. All three runs can end with identical code. That is the whole problem in one sentence: cost has decoupled from output. The engineering task is not to buy a cheaper model. It is to build the control plane that connects spend to accepted work.

This is written for the company that already went all-in. Flagship models — Anthropic's Fable 5, OpenAI's GPT-5.6 "Sol" — pinned at maximum reasoning. Auto-approve on. Loops that run unattended for hours. Velocity is excellent and quality is high. Then the bill arrives, or the usage wall hits mid-batch, and someone asks the question nobody scoped: what is this actually costing, and per what?

TL;DR

A cost problem is hard to diagnose without visibility and enforceable limits. Build that foundation, then prioritize improvements using measured cost, accepted quality and latency. The following are candidates to evaluate, not a universal ROI ranking:

  1. See it. Put every agent call behind one instrumented gateway. Tag each request with task, PR, feature, phase, model, and outcome. If you can't attribute a dollar to a merged PR, you don't have a strategy — you have an invoice.
  2. Measure reusable prefixes. Put stable context before volatile content where the API supports prefix caching. Measure reuse and the applicable model’s cache-read/write prices; savings depend on traffic and expiry.
  3. Tier the models. Flagship for architecture and hard judgment; a mid-tier workhorse for implementation; a cheap or local model for search, classification, and bulk edits.
  4. Test advisor–executor. Let an executor handle bounded work and consult an advisor at consequential decisions. Compare complete cost and accepted quality against a simpler single-model baseline.
  5. Govern the loops. Hard per-task budgets, aggregate concurrency caps, a circuit breaker on token velocity, and a kill switch that fires *before* the call leaves your system.
  6. Choose the billing surface. Compare subscription terms and metered automation. Use discounted Batch only for supported requests that tolerate asynchronous completion.

Optimize cost per accepted outcome, not cost per token. Establish a workload baseline and the quality threshold before estimating savings; vendor maxima and isolated case studies are not a transferable target.

Where the money actually goes

Before you cut anything, get the accounting identity straight. For a single agent job:

job_cost = sum(request_cost(r) for r in every_billable_attempt) + tool_and_infrastructure_charges
request_cost(r) = sum(tokens(r, category) * rate(r, category) for category in disjoint_billable_categories)
# Normalize provider counters first: cached tokens or reasoning may already be included.
# Retries are additional requests in the sum, not an extra charge added again.

Architecture changes the number and composition of billable requests. Repeated context, branches, retries and reasoning settings are explanatory drivers, not independent accounting multipliers.

billable_attempts = initial_requests + retries + subagent_requests + verification_requests
# Attribute each attempt once; apply its actual model, token categories and rates.

A lower token rate changes one part of the sum. Request volume, context reuse, retries and generated output determine whether that rate produces a lower cost for the complete accepted job.

DriverMechanismTypical failure mode
Output & reasoningGenerated tokens plus hidden reasoning, billed at the output rate on the major providersMax effort on formatting; regenerating whole files
Context growthHistory and tool results re-sent on every turnA long loop re-pays for old information each turn
Cache missesA changed prefix can't reuse the cached oneA timestamp near the top voids the whole cache
Subagent fan-outEach child gets its own context and outputSix agents each ingest the same snapshot
Retries & flailingFailed calls and rejected diffs still cost tokensNo spec, no stopping rule, so it wanders
Tool loopsTool output re-enters context and is reconsidered40k lines of build log persist for 20 turns

Two numbers anchor the intuition. On flagship pricing (mid-2026), output costs 5× input ($10 vs $50 per million tokens for Fable 5), and a generated token is roughly 50× the price of a cached input token. Reasoning effort can affect both result quality and resource use. Evaluate settings against the workload’s acceptance requirements and complete cost; neither maximum effort nor a lower setting guarantees the best result.

Context growth is the invisible one. Suppose a loop starts at 25,000 input tokens and appends 5,000 tokens of history per turn. Over 20 turns it sends ~1.45 million input tokens — not the 120,000-token final context a trace would show you. Hold each turn near 35,000 and you cut that to ~700,000. Nobody notices this in a demo; everybody pays for it in production.

Fan-out can duplicate context: six agents receiving the same 100,000-token snapshot initially process 600,000 input tokens before accounting for cache reuse. Anthropic’s June 2025 research-system account observed roughly 4× chat tokens for agents and 15× for multi-agent research in its data. Those are workload observations, not constants for current coding systems; measure whether parallel coverage earns its cost.

Tokenizer changes can alter token counts for identical text, but the effect depends on the model, language and workload. Measure representative prompts rather than assuming a universal30% increase. Repeated context can dominate volume even when output has a higher unit rate.

The two billing worlds

Every token you buy comes from one of two economies, and confusing them is the most expensive mistake on this list.

DimensionFlat subscription seatMetered API
Marginal costNear zero inside the allowanceEvery token is charged
Best forInteractive, human-paced workParallel automation, CI, batch
AttributionCoarse unless separately instrumentedNative, per request
Budget controlIndirect — via rate limitsExplicit caps and spend controls
Main riskTreating "included" as unlimitedUnbounded concurrency, runaway loops

A subscription’s allocated effective cost is its fixed fee divided by the chosen usage or accepted-outcome measure. It may be attractive within an included allowance, but that cannot be assumed universally. Optional overage and limits must be tracked separately.

As checked in September 2026, Anthropic documents included Fable access on eligible premium plans, with a model-specific allowance and faster consumption of regular limits. Other plans and optional overage use credits. July cutoff proposals are historical; API prices do not establish a subscription’s exact depletion formula.

And the arbitrage shifts under your feet. In mid-2026 Anthropic announced that headless and Agent-SDK usage would leave the subscription pools onto metered credits — then paused the change before it took effect. Flat-rate automation still works, but it has been openly flagged as temporary. Do not architect a business-critical pipeline on the assumption that subscription headless usage stays free. The clean split survives any pricing change: subscriptions for human-paced interactive work, metered API for elastic automation you can budget and attribute.


Part I — See it

🔎
You cannot optimize what you cannot attribute. Provider dashboards tell you how much; they rarely tell you per what. The first deliverable is not a cheaper model — it's a number next to every PR.

Put all agent traffic behind one instrumented layer — a gateway or wrapper — and record, per request: provider, model, tier, billing channel, reasoning setting; input / output / cache-read / cache-write / reasoning token counts; estimated and invoiced cost; and the metadata that makes it *mean* something — task ID, run ID, PR, commit, feature, repo, team, developer, phase (planning / execution / validation / review), and final outcome (merged, rolled back, abandoned, still running).

The tooling splits into layers that do different jobs. Pick one trace plane and one invoice source of truth — five dashboards produce five different totals.

LayerToolsJob
Invoice truthAnthropic Console + Cost/Usage APIs; OpenAI usage & costs dashboard + Admin APIReconciliation, limits, org-level totals
Gateway telemetryHelicone; LiteLLM; an internal LLM proxy (the Shopify / Uber pattern)Central logging, cost estimation, policy, per-team analytics
Agent traces + evalLangfuse, LangSmith, Arize PhoenixEnd-to-end runs, evals, datasets, debugging
Open standardOpenLLMetry / OpenTelemetry GenAI semantic conventionsPortable, vendor-neutral traces and metrics
Enterprise APMDatadog LLM Observability, New RelicFold agent cost into existing ops

Collect each client’s supported usage telemetry and record the client version, enabled exporters and coverage gaps. Normalize token categories before estimating API cost: input counters may already include cached tokens, and reasoning may already be included in output. Use a versioned rate table and flag unknown models instead of silently assigning zero cost. Reconcile estimates with provider billing records. API-equivalent token cost is not a measurement of subscription charges or allowance depletion.

The unit of management is cost per outcome, not cost per token:

MetricWhat it reveals
Cost per merged PREngineering-delivery economics (pair with change size + rollback rate so it can't be gamed with tiny PRs)
Cost per accepted taskWhether retries and routing produce useful completion
Retry-waste %Spend on rejected attempts and repeated failures
Cache-read ratioReuse of stable context — the caching scoreboard
Flagship-token shareWhether tiering is actually working
Escalation precisionWhether the router is too timid or too trigger-happy

The harness can change cost per accepted result through context, retries, tool execution and human review. Compare those complete costs for the same workload; a lower per-token rate can lose its advantage if it produces more rejected attempts.

Two disciplines make this real. Attribution: compute cost at the leaf (tokens × your price table) and roll it up by whatever key you stamped on the request — task, PR, team. Keep low-cardinality tags (team, feature, env) for charts and high-cardinality ones (user, repo, run ID) for forensics, or your dashboards explode. Subscription spend is fixed and needs an allocation convention — call it showback, not precise accounting, and never weaponize per-developer cost into a productivity ranking. Anomaly detection on behavior, not just dollars: a cache-hit-rate collapse, an output-token spike, or a jump in flagship-token share shows up *before* the invoice moves — the best leading indicators are tokens-per-task and retries-per-request. The teams that run this well route every AI request through an internal gateway, precisely so guardrails, analytics, and model swaps live in one place.


Part II — Cut it

Choose the next lever from measured cost and quality bottlenecks. The order below is a menu of candidates, not a universal return-on-investment ranking.

Cache like it's free money

When an agent repeatedly sends a substantial stable prefix, caching is a useful optimization candidate. Anthropic’s documented Fable 5 rates charge cache reads at 0.1× base input and five-minute cache writes at 1.25×; later models can use different multipliers. Under those rates, one write plus one read costs 1.35× instead of 2× for two uncached uses. Check minimum prefix size, expiry and actual reuse before estimating the task-level saving.

The catch is that caching is a prefix game: any change early in the prompt invalidates everything after it. Order context most-stable-first — system instructions, tool definitions, durable repo conventions, reference material — and push volatile content (timestamps, current state, logs) to the very end. Put a request ID near the top and you've turned your cache off without noticing.

Moving volatile content behind a reusable prefix can improve cache reuse. Measure cache-read ratio and total accepted-task cost before and after the change; an unnamed team’s invoice result is not evidence for a general savings estimate.

Build a portfolio, not a monoculture

Model routing can concentrate expensive reasoning at consequential decisions. Whether it preserves quality must be tested against the same acceptance criteria, including retries and review.

Task classDefault tierEscalate when
Architecture, migrations, hard debuggingFlagshipKeep a human on irreversible, high-blast-radius calls
Ambiguous planning, novel failuresFlagship / strong midSystems interact or the evidence is thin
Normal implementation, bounded fixesMid (Sonnet-class)Tests keep failing or scope creeps
Bulk edits, boilerplate, refactor executionCheap / mid — flagship only *plans*Behavior or a repo-wide invariant changes
Search, summarize, triage, extractCheap (Haiku / nano / local)Conclusions drive a consequential decision
Classify, format, routeSmall / local modelValidation fails or confidence is low
Security, auth, destructive opsFlagship + deterministic checksAlways add review appropriate to the risk

Compare current rates for the exact model and billing surface. GitHub moved Copilot to token-based AI Credits on June 1, 2026; request multipliers remain a legacy concept for certain existing annual plans. Do not apply those multipliers to current credit billing. Evaluate routing against the same accepted-quality threshold, including retries and review.

Local or open-weight models can be useful for suitable workloads, but no parameter-count cutoff decides self-hosting economics. Hardware utilization, latency, capacity, operational effort and accepted-result quality must be compared with the chosen API.

Advisor–executor: the backbone

One pattern worth evaluating — the advisor–executor split — puts a cheap executor on the main loop (reading files, editing, running tools, iterating) and calls a flagship advisor only at decisions that change the shape of the work. Most tokens bill at the executor rate; judgment stays near the frontier. This reframes orchestration as an economic control surface, not just a quality trick.

The whole system lives or dies on the escalation rule. Escalate on *observable* triggers, not the model's self-reported "confidence" (which isn't calibrated): the same error twice, a change that crosses an API or data boundary, a diff past approved scope, a validation loop failed more than once, or 70–80% of the task budget consumed without acceptance. Ask the advisor to rubber-stamp every edit and you've built an expensive single-model system with extra latency.

The economics, illustratively — hold a workload at 20M input / 2M output tokens, executor at ~10% of flagship cost:

AllocationIllustrative costvs all-flagship
100% flagship$300—
80% executor / 20% flagship$84−72%
90% executor / 10% flagship$57−81%
100% executor$30−90%

The 90/10 split is the interesting operating point: most tokens cheap, hard calls still frontier-grade. The exact ratio isn't the point; the *shape* is.

Route by exception, not by habit

In a simplified two-stage cascade, cheap-first costs Ccheap + (1−p)×Cflagship. It beats always-flagship when Ccheap < p×Cflagship only if success is detected reliably, accepted quality is equivalent, fallback cost is unchanged and verification overhead is absent or included. Real routing must measure those assumptions, false acceptance and additional latency.

RouterArena provides a framework for comparing routers across quality and cost. Its results depend on the model pool, tasks and evaluation metric; there is no universal savings percentage. Start with an inspectable baseline and use held-out outcome data before adding learned routing.

Context is a budget, not a bucket

Compare retrieval and full-context approaches on the same tasks, with both quality and complete cost measured. Larger context can preserve useful evidence but also add repeated input, distraction and long-context charges. The best choice is workload-dependent.

Tool-result deletion, summarization and subagent isolation are distinct techniques. Anthropic reported an84% token reduction from context editing in a specific100-turn web-search evaluation; that is not a universal compaction saving. Deleting stale results need not run a summary model, while model-based compaction incurs its own request cost. Validate preserved information and total usage for the chosen method.

Batch the rest

Use Batch for supported independent requests that can wait for asynchronous results. Check model and endpoint eligibility, tool dependencies, expiry, partial results, retention and orchestration costs. Anthropic documents a 50% Batch discount, a 24-hour expiration window and best-effort cache hits; expiration can leave requests unfinished, so it is not a guarantee that every request succeeds. Where a 0.5× Batch multiplier stacks with a 0.1× cache-read rate, eligible cache-read tokens cost 0.05× base input—a 95% reduction for that category only. Cache writes, output, tools and other charges remain separate.


Part III — Govern it

🛑
A budget is executable policy, not a spreadsheet you read after month-end. Every autonomous loop needs caps it cannot talk its way past.

The failure mode that produces horror-story invoices is not one greedy agent — it's five agents, each within budget, drawing from one pool, blowing 3× past the aggregate before any dashboard refreshes. So the controls have to be structural:

GuardrailRequired behavior
Preflight reservationEstimate worst-case cost, reserve it from the budget, call only if the reservation clears
Soft threshold (50%)Reassess projected cost; use only validated context or effort changes that preserve acceptance requirements, otherwise checkpoint and escalate
Escalation threshold (80%)Freeze new fan-out; require advisor sign-off before spending the rest
Hard threshold (100%)Checkpoint and pause — never continue optimistically
Progress detectorKill repeated errors, unchanged diffs, cyclic tool sequences
Aggregate concurrency capCap the shared pool, not just per-agent
Kill switchStop one run, one queue, one model, or the whole fleet — before the call leaves
Rate-limit handlingPause, jitter, resume once — never retry-thrash the wall

The non-negotiable is the pre-call kill switch: estimate worst-case cost, reserve it atomically, make the call only if the reservation succeeds, reconcile after. This is the exact line between observability and control — your traces *measure and alert*, but only a gateway in the request path can *block* a charge. A budget that lives in a dashboard is a cost report, not a cost control (and note the sharp edges: some providers' "spend limits" are soft — they alert while the requests keep flowing). Cap velocity, not just totals — a per-hour or per-session ceiling trips a fast retry loop long before a daily budget notices, and loop detection (the same tool call two or three times over) kills the classic runaway. And pause-on-limit must not become retry-on-limit: hit a wall, persist state, compute the next window, add jitter, and resume *once* — thrashing a rate limit just wastes capacity and duplicates work.

Budget *transitions* matter more than the numbers: reserve a per-task budget up front (a routine task $5, a hard debug $25, a migration $100); at 50% consumed, reassess and make only validated changes that preserve acceptance requirements; at 80%, forbid new fan-out and require escalation; at 100%, checkpoint and pause. A loop without these isn't autonomous engineering. It's an unbounded recursive purchase order.


The recommendation: measure, then prioritize

Instrumentation and enforceable limits establish the baseline. After that, choose caching, retrieval, routing or batching according to measured workload behavior and the required quality. No fixed ordering guarantees the largest gain.

  1. Instrument first. Unmeasured optimization just relocates cost. Get cost-per-task and cost-per-PR before you touch a model contract.
  2. Turn on caching. Validate compatibility and measure the actual gain. Fix prefix ordering and measure cache-read ratio.
  3. Evaluate the architecture. Compare advisor–executor and model routing with a simpler baseline. Keep them when accepted quality and complete cost justify the coordination overhead.
  4. Check caching and context hygiene. Remove avoidable repeated work while preserving required evidence. Compare retrieval and compaction on complete task outcomes; neither guarantees an improvement.
  5. Choose billing deliberately. Use current subscription terms for interactive work and an appropriate metered route for automation; choose Batch only when its operational constraints fit.
  6. Put hard budgets and kill switches on every loop. Re-evaluate when models, rates, workload or acceptance requirements change.

Monday morning, concretely: put every call behind one wrapper; require task_id, PR, phase, and tier on each request; reconcile yesterday's estimate against the provider invoice; publish cost-per-merged-PR, retry-waste, and cache-read ratio; move stable instructions into a cached prefix; stand up one advisor–executor workflow with objective escalation triggers; set per-loop dollar, token, concurrency, and wall-clock caps; and review your ten most expensive *failed* runs before you argue about model choice. Use that evidence to choose the next improvement instead of assuming which lever will win.

Four hot takes

  • Maximum effort throughout should be a measured workload choice. Evaluate any proposed change against the required quality and complete cost. A budget threshold alone is not evidence that lower effort is acceptable.
  • A reusable prefix can be a major lever. If repeated input dominates cost, evaluate caching alongside model choice. Prioritize the measured saving while holding accepted quality and latency requirements fixed.
  • If you can't attribute cost to a PR, you don't have a cost strategy — you have a bill. Provider totals explain procurement. They say nothing about engineering economics.
  • Rate limits are a scheduler signal, not a budget. Hitting a wall proves only that capacity was consumed. It says nothing about whether the work was worth buying.

You can't afford infinity

The point of the AI-native shop was to remove the human bottleneck from engineering. It worked — and it moved the bottleneck to the meter. The answer is not to throttle ambition back down to a spreadsheet. It's to build the plane that lets you spend aggressively *where it compounds* and refuse to spend *where it doesn't*: judgment at the frontier, volume at the floor, every dollar attributed, every loop bounded.

A cost control plane makes continuous operation measurable and bounded. It lets the team decide which workloads justify their resources, which need a different design, and which should stop. The result still has to earn its cost through accepted work.

Current implementation details: Claude pricing and cache multipliers; Batch eligibility, expiry and partial results. Checked September 20, 2026.