Your Agents Eat Your Subscription

Subscription-authenticated agents can share interactive usage allowances. Check authentication, current rate cards and spend-control semantics before budgeting automation.

Your Agents Eat Your Subscription

Subscription-authenticated unattended runs can consume the allowance needed for interactive work. A headless process does not automatically have a separate budget.

The billing route depends on authentication and the product’s current plan terms. Distinguish included subscription usage, prepaid credits and API-key billing before planning a batch.

AI tooling budgets span subscriptions, usage credits and company plans. Confusing those billing surfaces can produce expensive assumptions. The comparisons below describe the plan structure discussed in August 2026.

An Example Subscription Stack

An engineer can use individual subscriptions or centrally managed seats. That choice changes both capacity and account administration.

LineWhat it isCost
Claude Max 20xInteractive Claude Code plus eligible subscription-authenticated runs$200 / month
Codex Pro 20xSecond opinions, reviews, an independent pair of eyes$200 / month
Claude usage creditsOptional prepaid overagea separately chosen illustrative allocation

This is a hypothetical subscription stack. An unattended loop may run a headless client, but its credentials and service configuration determine the allowance or wallet it consumes. No purchase history is implied.

There Is No Automation Budget

The assumption almost everyone makes: interactive sessions come out of your subscription, and headless runs — CI jobs, unattended batches, agent loops — come out of something else. An API bucket. A separate allowance. Something.

When authenticated through eligible subscription flows, interactive and unattended work can draw from the same allowance. An API key can select a separate billing route.

Anthropic’s June 15 update paused a proposed change to Agent SDK billing. Its current note says eligible subscription-authenticated Agent SDK and claude -p use still draw from plan limits; API-key billing remains separate. The earlier proposal on that page is historical.

For Codex with ChatGPT authentication, local messages and cloud chats share the plan’s allowance, with limits that depend on the plan and model. API-key usage is billed separately. Check the current pricing documentation for the exact route.

⚠️
The consequence is not theoretical. A long unattended batch can rate-limit you out of your own terminal. Your agents and your hands compete for the same quota, and the agents are faster.

What “Included” Actually Means

Claude runs a five-hour session window plus a weekly cap that applies across all models. Anthropic does not publish the numbers. Max 5x is described as five times Pro per session, Max 20x as twenty times Pro. What that means in messages or tokens, you discover in Settings → Usage, or by hitting the wall.

OpenAI’s pricing guide, checked 20 September 2026, publishes the following estimated local messages per five-hour period. They are workload-dependent estimates, not fixed message entitlements.

ModelPro 5x ($100/mo)Pro 20x ($200/mo)
GPT-5.6 Sol50–500 messages200–2,000 messages
GPT-5.6 Terra125–1,000 messages500–4,000 messages
GPT-5.6 Luna1,250–10,000 messages5,000–40,000 messages

The ranges vary with model, context, reasoning, tools and caching. They are planning estimates, not a guarantee that a particular batch fits before a reset. Consult the current usage panel and include retries in the budget.

What Overage Actually Costs

Extra usage may be metered at nominal API-equivalent rates while the effective cash cost depends on purchased-credit terms and discounts. Keep the meter’s units, credit purchase price and actual cash charges separate. Headless mode itself determines none of them.

Current reference rates are available in the Claude API rate card and Sol model card. The table below was reconciled to those documents and the Codex credit table on 20 September 2026; quoted API rates do not establish a subscription allowance.

ModelInput / MOutput / M
Claude Opus 5$5.00$25.00
Claude Opus 5 (fast mode)$10.00$50.00
GPT-5.6 Sol$4.00$20.00
GPT-5.6 Terra$2.00$12.00
GPT-5.6 Luna$0.20$1.20

At the current standard rates shown here, Opus 5 costs $5/$25 per million input/output tokens and Sol costs $4/$20. Different token usage, cache behavior and service tiers can still change the cost per accepted task. The earlier $5/$30 Sol figures are not used as a current comparison.

Do not over-read this. Two models do not consume the same tokens for the same task, Terra and Luna are dramatically cheaper if your work fits them, and Claude's prompt caching cuts input cost by up to 90% while batch processing halves it. The real lesson is narrower and more useful: read the rate card, do not reason from vibes about which vendor is cheaper at the margin.

A balance and a spending meter are different measurements

Here is a trap in monthly spend controls.

A positive credit balance can appear beside a monthly usage bar reading $0.00 spent, 0% used.

Credit purchases and usage consumption may have separate meters. Confirm what the cap measures, its reset period and how it interacts with prepaid balances and auto-reload before treating it as a cash-spending ceiling. A positive balance beside a zero usage meter does not establish that the cap authorizes an additional purchase.

Do not infer spending-control semantics from a dashboard screenshot. Use the current product documentation or a documented controlled test; if the behavior is unclear, keep additional charging disabled until it is understood.

Choose auto-reload deliberately

Both vendors offer automatic top-up: the balance drops below a floor, the card gets charged and work continues. A manual monthly purchase makes three trade-offs explicit.

Predictability under a fixed budget. An auto-reload that fires repeatedly during a runaway loop can exceed an intended monthly ceiling.

Credit-pack discounts, eligibility and expiry can change. Compare the checkout terms for manual purchases and auto-reload; do not assume either receives a particular discount.

Disabling auto-reload limits automatic credit purchases. Verify the provider’s treatment of in-flight and delayed usage, and configure the controller to detect exhaustion and save a resumable stop. Manual funding alone does not establish that the batch will park gracefully.

Manual credit purchases can interrupt unattended work, so test both stopping and recovery. Treat manual funding as a spending control whose enforcement must be verified, rather than proof of an exact cash ceiling.

The Seat Math Nobody Runs

Anthropic’s Team plan guide, checked 20 September 2026, specifies per-session allowances of 1.25× Pro for Standard and 6.25× for Premium, with separate weekly limits. These ratios are not guarantees of identical task throughput. The listed US seat prices exclude taxes.

When a company says “we'll give you a seat,” that sounds like the larger option. Run the arithmetic. Claude Team has two seat types, and Claude Code is included on both:

What you getUsage per sessionPrice
Team Standard seat1.25× Pro$20/mo annual · $25 monthly
Team Premium seat6.25× Pro$100/mo annual · $125 monthly
Personal Max 20x20× Pro$200/mo

A Standard seat is roughly one sixteenth of the session capacity of a personal Max 20x. Even a Premium seat lands at about a third. For someone running agents all day, the company seat is not automatically the bigger option — depending on the tier it can be a fraction of a personal plan, at a comparable price point.

Which means the useful question, when someone offers you a seat, is not “do we have a plan?” It is two questions: which seat type, and what credit allocation. “We have a Team plan” is not an answer — the tier is the whole number.

On the corporate side, Team and seat-based Enterprise credits are configured org-wide by an owner, with a monthly spend limit and per-user limits. Codex Business and Enterprise are credit-based with their own spend controls. Both are real accounting machinery. The question is whether you need it.

How Two Lanes Became One

Three possible architectures show where account-separation complexity comes from.

The first is a wrapper: a PATH shim with a directory-scoped launcher. Inside one project tree it verifies the company identity and refuses to launch if verification fails; outside that tree it verifies an individual identity. Session-start hooks can warn when the wrong account is in use. This can fail closed, but it creates another authentication mechanism to maintain.

That wrapper requires separate token stores, verifiers, guards and login paths. The launch ordering becomes another dependency for unattended workers.

The second is a machine boundary: one identity on a laptop, another on a cloud development machine. The separation moves from custom routing code to distinct environments.

The third is one identity per vendor across both machines. It removes the account-routing wrapper, with the provisioning and offboarding trade-offs discussed below.

In that single-account arrangement, the laptop and cloud development machine use the same identity. No directory-sensitive shim or account-selection branch is needed.

💡
When you find yourself writing code to decide which credentials a directory deserves, the problem is usually the number of credentials, not the absence of a wrapper.

There is a real argument on the other side, and I want to state it fairly. Centralized seats give you provisioning, offboarding and audit — when someone leaves, one admin action revokes their access. Reimbursed personal subscriptions give you none of that; you get a receipt and a trust relationship. If you have compliance obligations or meaningful headcount churn, that machinery earns its keep. For a small team of senior people, it is overhead with a login flow attached.

One argument for a tool budget is that it leaves the choice of harness with the engineer. A centrally selected seat chooses that tool for them.

What I'd Do On Monday

Concretely, if you run agents for a living and someone else is paying:

Find your real ceiling. Open the usage panel on both vendors before you plan anything. On Codex you get published numbers; on Claude you get your own consumption history, which is the same information the hard way.

Budget subscription-authenticated unattended runs against the shared allowance described by that product. Budget API-key runs against their separate wallet and rate card. Verify which route the worker actually uses.

Confirm what the spending cap measures before configuring it. Credit purchases, consumed usage and cash charges are separate events; a cap is not a verified cash ceiling until its behavior is established.

Choose automatic reload only within an explicit budget and authority. Check purchase terms and stopping behavior; do not rely on a presumed volume discount or an undocumented meter.

Ask which seat, not whether there's a plan. A Standard seat and a Max 20x differ by a factor of sixteen and nobody volunteers that.

A supported API-key route can isolate unattended usage from a subscription allowance. It introduces separate billing and controls; compare effective rates and validate its own spending limits.

None of this is exotic. It is just that the pricing pages describe subscriptions and the docs describe APIs, and almost nobody writes down the part where an unattended agent quietly drinks from the cup you are holding.

The useful operational question is which meter each worker consumes, how it stops and where its verified handover is stored. Make those properties explicit before an unattended batch starts.