I Fired My Busiest AI Reviewer
Same $700, one provider fewer. What 620 review receipts say about the cut, what it costs, and how two Claude accounts now route by directory.
On September 29 I wrote that I would not drop Grok just to make a new pricing card fit, and that I would trade a third provider for more Codex "only when the work shows that this is the exchange I need." Five days later I cancelled Grok.
The uncomfortable part comes first: in my own receipts, Grok completed more independent reviews than Claude or Codex. I cut it anyway. The same session brought a second change: directory-based routing for two Claude accounts.
| ⚡ | TL;DR Money. Same $700 a month at list prices: Claude Max at $200 and ChatGPT Pro 500 at $500, which includes Codex. The $300 that paid for SuperGrok Heavy moved to the OpenAI seat. Models. Codex runs GPT-6 Astra at Receipts. 620 review runs since September 6. Grok completed 269 reviews, Codex 138, Claude 114. When Codex was tried first, it failed 60% of the time, almost always on its usage limit. What the receipts can't say. Anything about review quality. This is a purchase made on availability evidence, not a proof. Accounts. Work for the company I work for runs on the company's Claude Team seat; everything else runs on my personal plan. The directory decides, and a hook blocks a prompt before any model call when it catches a session on the wrong account. |
The same $700, rearranged
These are list prices before taxes and extras, the same basis as in I Love Trios. The company's Claude seat is not part of this budget.
Pro 500 needs precise words. It is a ChatGPT subscription with 25 times the Plus allowance and access to Astra Ultrafast. It is not $500 of credits; the September 29 article has the breakdown. Pro plans currently have no five-hour limit, and weekly limits can still apply.
Both models stay at their ceilings. I did not go back to Fable 5.1. Anthropic's own numbers put Opus 5.5 at Fable's level on most work for $4/$20 per million tokens, against Fable's $10/$50. These are per-token prices, not a measure of subscription consumption; I simply judge Fable too costly for the seat.
What 620 receipts say
Every independent review in my harness goes through one helper, and the helper writes a receipt for every run: who asked, which providers it tried in which order, which one completed, and the classified reason for every failure. Between September 6 and October 4 it wrote 620 of them. The order of the two other providers was random. On an availability failure (credits, rate limit, login, outage) the run moved on, and a fresh child of the caller's own provider was the last resort.
The failures tell three different stories. Grok failed 116 times, 109 of them on an exhausted balance. Codex failed 175 times, 171 of them on its usage limit. Claude failed 158 times: 95 authentication errors, 30 invalid results, 29 rate limits and 4 other. The 16 attempts missing from those sums were interrupted or never wrote a final status. Of the 620 runs, 521 ended with a completed review and 48 found every reviewer unavailable. The remaining runs failed for other reasons, were interrupted, or never finalized.
Fallback order distorts the totals, so here is each provider when it was the first one tried:
Grok was not a backup singer. 152 of its 269 completed reviews came from runs where it was the first choice. Another 83 came after the first provider had already failed, and 34 were Grok reviewing Grok as the last resort. Any story in which Grok was dead weight is fiction.
Codex was the bottleneck. It logged 171 usage-limit failures, carrying three different reset dates: September 14, September 19 and September 26. That is not 171 outages. It is the same wall, hit again and again until each reset. When Codex was tried first and failed, the limit was the reason 124 times out of 128. The $200 plan ran dry repeatedly, and the third provider kept reviews moving while it was dry.
Claude's failures were a different animal. 92 of its 95 authentication failures were HTTP 403s, and 72 of those landed in two bursts, on September 16 and September 21. They were classified as authentication failures, not capacity, and more Codex will not fix Claude's access failures.
What the receipts cannot say. Nothing in them measures review quality, accepted outcomes, or whether a Grok finding ever changed a decision. Five days ago I listed the evidence I wanted before trading the third seat: which tasks a Codex limit delayed, and whether the third provider found something material. The receipts answer neither. They measure availability.
So this is a purchase, not a proof. My bet, labelled as mine: more Codex capacity will do more for my work than a third provider did. The receipts show a Codex plan that kept running dry; they cannot show what Grok's reviews were worth. Pro 500 buys capacity from the provider I want doing more of the work, at the same total.
What I give up
The case for a third provider was never that three models agreeing makes something true. The case was another way to be wrong in a different direction, and somewhere for a review to go when one vendor is out.
Removing Grok loses that third perspective and one fallback option. A bigger Codex allowance buys capacity. It does not buy a third set of model weaknesses.
The review contract is now four lines:
- A Claude caller asks Codex. A Codex caller asks Claude.
- If that other provider is unavailable, a fresh, isolated child of the caller's own provider reviews. Its receipt says
same_provider_fallback, and it is never reported as cross-provider agreement. - An explicitly named reviewer gets one attempt and no fallback.
- No recursion, and no reviewer's answer counts as permission to act.
A fresh child does not inherit the parent session's context, which helps. It does inherit the provider, and probably its blind spots. I am trading diversity for fewer walls.
Pulling a provider out is a refactor
Cancelling the subscription took one click. Removing the assumption that there are three providers took a working session.
Grok was wired into the review runner and its contract tests, the hooks, the statusline, the usage parser, the measurement ledger, limit watch and takeover, the generated skill adapters in every project, the autonomous lane's failover, a headless loop, instruction files and memory notes. Removing it touched 167 files in six repositories: 3,517 lines added, 3,803 removed. Afterwards 985 harness tests pass; 26 tests in one older module fail exactly as they did before.
Three decisions mattered more than the deletions:
- History stays readable. The ledger keeps 1,497 Grok turn aggregates, and the review store keeps every receipt that named Grok. Reports show them as a retired provider, and nothing new is recorded. Deleting evidence would only weaken the next comparison.
- Stopped work stays stopped. A few autonomous batches were paused with Grok saved as their lane. Their resume path now maps a retired lane to the default, Codex, instead of failing or quietly bringing Grok back. Nothing was restarted.
- The removal has a regression check. The Grok CLI and its own settings are still installed; the tool itself is fine. But the global check now fails if any harness hook, rule, adapter or statusline points at it again, and the project check fails on any repository that declares a Grok adapter.
One person, two Claude accounts
The second change is about who pays. Work for the company I work for must run on the company's Claude Team seat. Everything else runs on my personal Max plan. That has to hold for interactive sessions, headless runs, loops, the autonomous lane and second-opinion children. A rule that only covers the terminal I happen to be watching is decoration.
Claude Code's authentication docs name the mechanism: "To stay signed in to multiple accounts at once, such as work and personal accounts, give each account its own configuration directory." CLAUDE_CONFIG_DIR points at that directory, and the directory carries its own login, settings, history and .claude.json.
The registry
One small, non-secret file decides. The longest matching root wins, and everything outside every root belongs to the default account. In this illustration the personal directory is only a label: the router leaves CLAUDE_CONFIG_DIR unset for personal work and sets one fixed absolute path for work.
{
"default": "personal",
"accounts": {
"personal": {"email": "me@personal.example", "config_dir": "~/.claude"},
"work": {"email": "me@work.example", "config_dir": "~/.claude-work",
"roots": ["~/Projects/work"]}
}
}The trap I walked into
My first instinct was to make every launch explicit, including CLAUDE_CONFIG_DIR=~/.claude for the personal account. Wrong. With the variable pointing at the default directory, claude auth status reported loggedIn: false, and the run created a fresh state file inside ~/.claude. Unset and set-to-the-default are two different login states.
The Keychain shows why. On my Mac the default login lives in the entry Claude Code-credentials. The work login lives in Claude Code-credentials- plus eight hex characters, and those eight characters are exactly the start of the SHA-256 of the CLAUDE_CONFIG_DIR string I set. Pointing the variable at the default directory therefore selects an entry that was never logged in. None of this is documented and it can change. A different spelling of the same folder is a different string, so the router allows exactly one spelling and checks it.
The rules that follow: the personal account runs with the variable unset; the work account runs with exactly one registered absolute path; anything else is a mismatch.
Routing every entry point
Choosing the directory is not enough if a credential variable outranks the login. ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN and CLAUDE_CODE_OAUTH_TOKEN all sit above the stored subscription login in Claude Code's authentication order. The router strips all three.
- Terminal. A
claudeshell function asks the registry about$PWD, then sets or unsets the variable and strips the overrides in a subshell. Mycandcxaliases go through it. - Headless runs and loops.
ai-workspace claude-exec --root DIR -- claude ...checks the config directory, the login recorded in it, an optional pinned organization ID and the absence of overrides. If anything is off, it exits with code 3 and launches nothing. - Second opinions. The review runner takes
--rootand puts the Claude child on the account that owns the reviewed work. If that account is not logged in, the attempt counts as an authentication failure and routing continues; it never borrows the other login. A child started from a scratch directory inside a work session keeps the work account.
claude() {
local routing
routing="$(ai-workspace claude-account --env --root "$PWD")" || return 3
( eval "$routing" && command claude "$@" )
}The guard
Routing does the work, and a hook catches what routing misses. For a UserPromptSubmit hook, Claude Code's hooks reference is explicit about what a block does: it "Blocks the prompt, so it never reaches Claude." Mine compares the account the project needs with the account the session is actually on. It takes the session's config directory from the hook's transcript_path, which survives an environment scrub, and reads the login recorded there. On a mismatch it blocks, with the exact command to restart correctly. SessionStart warns earlier, and the statusline shows acct personal, acct work, or WRONG ACCOUNT with both emails.
I tested it from the work tree with a plain, unrouted claude -p. The hook blocked it: zero turns, zero tokens, no model in the result. The same prompt through the router answered on the company seat. Both organization IDs are pinned too, so a login to the right email in the wrong organization is refused as well.
The limits are documented, and I keep them in view:
- A
UserPromptSubmithook that hits its timeout is cancelled, and the prompt "still reaches Claude". The default timeout is 30 seconds; my check is a local file read. --bare,disableAllHooksand managed settings can switch hooks off. That is why launch routing is primary.- IDE and desktop-app launches are not routed by my shell. When the guard is active and completes its check, a wrong-account prompt there gets blocked, not fixed.
Why not just /login
I have tried account switching before. An older wrapper of mine swapped tokens in place, and a /login in the middle of a session could re-point a session that was still running. I deleted it in August. Now each account has its own config directory, and I restart on the right account instead of switching inside a running session.
Staying logged in
Both accounts use ordinary /login OAuth logins, and those still expire: Claude Code warns at startup when the login you created with /login is "within three days of expiring", and once it expires and can't be refreshed, requests fail until you log in again. I skipped claude setup-token. Its one-year token "can only make model requests", so claude.ai connectors are gone, and it lives in an environment variable, which outranks whichever stored login the active directory holds. The router strips exactly that variable. I also moved to Claude Code 2.1.289, whose changelog lists a VS Code fix reverting a 2.1.288 change to claude auth status "that may have made sign-outs more frequent"; my account checks call that command too.
One visibility gap remains. The statusline docs list the plan-window rate_limits field for claude.ai Pro and Max subscribers, plus a gateway case, so I can't count on it for the Team seat. My limit watch may learn about that seat's limits only from errors. Limit state is now stored per account, so a personal window never raises a warning in a work session, and the other way around.
What I measure next
The receipts showed capacity walls and a busy third provider. They cannot tell me whether this reallocation improves the work. The next report needs four things, each tied to the actual reviewer and account:
- Codex interruptions. Which tasks a usage limit actually stops, and whether Pro 500 makes that go away.
- Review contribution. Material findings, decisions changed and corrections accepted, per reviewer.
- Fallback rate. How often cross-provider review is unavailable and a same-provider child steps in.
- Account routing. Blocked prompts, refused launches and login failures, kept apart from capacity failures.
Those numbers decide the next round of harness work. Further out, I want my own front end for all of it: an app built around this harness, a custom interface for working with it.
Grok completed the most reviews, and I cut it anyway. The next month of receipts will show whether more Codex capacity removes the interruptions. What Grok's own reviews were worth, no future receipt can tell me.