Handling AI Usage Limits and Clean Handovers
Detect provider usage limits, save verified task state, preserve stop markers and hand over cleanly without confusing account limits with model quality.
Design for interrupted work
A coding assistant can become unavailable in the middle of a task. Without a durable handover, a successor must reconstruct changed files, verified results and unfinished work from incomplete conversation history.
A headless loop also needs explicit failure classification. Retrying an exhausted allowance as though it were a transient tool failure can produce repeated empty iterations.
Record a provider’s reported error category, scope and reset information when available. Keep authentication, billing, rate-limit and service failures distinct; do not infer a percentage from a refusal.
The design below combines native observations with advisory warnings and saved state. It applies to multi-provider workflows where each client exposes different signals.
A limit is not an outage
A usage limit and a service outage need different handling. Some providers report allowance percentages or reset times; others only refuse a request. Preserve those distinctions and do not invent a reset time or assume another provider has capacity.
A signal is useful only when it leads to a defined response: stop starting work, leave a consistent state and write a handover that the successor can verify.
Limits can apply to an account, product, model or usage window. Preserve the scope and reset information reported by the provider. A capacity limit does not itself mean that model quality declined. Repeatedly retrying an exhausted allowance is not a recovery strategy.
What each program will tell you
That is the whole basis of the design: read what is native, store it, act on it. No model call is involved anywhere in the signal path.
One snapshot per provider
A native hook can keep a local JSON snapshot per provider: observed windows, reported reset times, refusal state and an advisory level such as ok, warning, critical or rejected. Store snapshots locally with owner-only permissions. Summaries injected into hosted sessions are sent to the selected provider.
Three rules keep the snapshot honest. A window drops out when its reset time passes, so yesterday’s number cannot scare today’s session. A refusal clears when the exhausted window resets; when no reset is known, which is Grok’s case, the refusal is held for six hours and then shown as stale, not as ok. And a later, healthier-looking observation never wipes a live refusal, because a statusline at 30 percent does not mean the account stopped refusing.
Thresholds: short windows (Claude Code’s five-hour window and Codex’s primary, also five hours) warn at 85 percent and turn critical at 95. Weekly windows warn at 97 and turn critical at 99.5, because a week at 86 percent with three days left is information, not a reason to wrap up. All four are configurable, and all four are guesses until I have seen how often they fire.
LIMIT WATCH
When a level is reached, the hook injects one LIMIT WATCH block per session and level, through whichever native event each program supports: Claude Code on UserPromptSubmit, SessionStart and PostToolUse; Codex on SessionStart and PostToolUse; Grok Build on its next PreToolUse. Once per session and level means a session hears “warning”, “critical” and “rejected” once each, not on every tool call, and hears it again after a compaction or a resume, when the earlier text is gone from its context. The Claude Code statusline shows it too, as “!! LIMIT WARNING 90%”.
The harness is advisory by design. It never blocks a prompt or a tool call. It never asks the model to continue after a stop. It never retries a provider. And it stays silent inside subagents, which have no business wrapping up the parent’s work. Those four rules are the part I care about most: a limit watcher that blocks calls or auto-retries is a second source of surprises, and the point of this was to remove one.
Wrap up, do not push through
A warning is useless without a fixed answer to “so what now”. That answer is one contract file, the same for all three programs, generated into a native “takeover” skill in each. The order:
- Stop starting new work. Warning: reach a safe stopping point after the current step. Critical or rejected: finish only the step in flight, abort what cannot finish safely, and say so.
- Leave every repository consistent. Commit and push under the normal rules: no skipped hooks, no force push, no STOP marker removed to make a tree look clean.
- Update the ledgers the work already keeps, so the lasting state lives in the repository, not in the chat.
- Write the takeover record (the command below).
- Tell me in one line: which limit, when it resets, and “open another agent and say: takeover from claude”. Then end the turn. Nobody waits for the window inside the same session.
ai-workspace takeover write --provider claude --root "$PWD" --reason "five_hour 96%" --body-file - <<'END'
## Task
## Done (verified)
## Next, in order
## Loops and automation
## Open risks
ENDThe toolkit records what it can check on its own: branch and head, dirty and unpushed counts, and the provider’s limit snapshot. It also records the resume pointers under the project: loop directories with their STOP markers, the saved runtime lane, and the last limit or failover lines of their logs. The session writes the body: the task, what is done with evidence, next steps in order, which loops exist and how to relaunch them, open risks.
The machine records facts it can check; the model records the task judgment and the evidence behind it. A useful handover needs both.
Takeover from claude
The other side of the contract is just as fixed. A session told “takeover from claude”, in any of the three programs, starts with two commands:
ai-workspace takeover status
ai-workspace takeover show --root "$PWD"The first command reports which providers are limited as of the last observation, including the one you are running in. The second prints the latest record for the project plus a fresh scan of its loop pointers, because loop state can move after a record is written. The successor reads the record, every named ledger and the applicable repository rules in full; inspects the current branch, HEAD and uncommitted or unpushed changes; and reconciles differences with the saved record before any pull or other checkout update. Update the checkout only under the repository’s normal policy, preserving existing work.
The record is a pointer and a to-do list. It is not authority. It cannot grant a permission the previous session did not have, and “the last session said so” is not a reason to do anything the repository’s rules forbid. The successor re-derives the work from those rules.
Loops get extra rules. A loop I stopped stays stopped. A loop paused by a limit is relaunched only on the successor’s own provider lane, with the loop’s documented command, and only when the loop’s policy allows it. When the work is picked up, the record is marked consumed with ai-workspace takeover done. If the successor’s own provider is also limited, it does not start. It says which providers remain usable and stops there. Three limited assistants is a real state, and the right answer to it is a sentence, not an attempt.
When all three are full
If every permitted provider is unavailable, do not launch another attempt. Leave the handover pending until a route becomes available, or continue authorized work that does not require an assistant.
Additional usage is a separate purchasing decision, not an automatic fallback. A capacity policy should state which routes are allowed and preserve that boundary. Unaided work can still include reading sources, drafting specifications and checking engineering assumptions.
The loops and the console
A headless controller and its interactive operator session need separate recovery paths. The controller can classify provider failures and preserve batch state; the interactive session still needs its own handover if it cannot continue.
Do not change a running controller’s template to repair its recovery behavior. Prepare and verify the change separately, then apply it at the documented safe boundary.
Parity and the second machine
Each program has its own hook events and its own native skill, so drift is the default state. The toolkit’s check verifies the hook registrations in all three programs, the generated skills and the contract together, in one run. Eighteen tests cover the new module, seven of them added after an independent review of the first version found real gaps: a refusal that could age out instantly, a status line that could wipe it, a resumed transcript that could resurrect it. The suite is at 306, all passing.
A second machine gets the same files through an allowlisted packet exported from a reviewed dotfiles revision. The rule I wrote down today for that path: replicate, never reinvent. A harness rebuilt from memory on a second machine is a different harness, and a different harness has different gaps.
What is not known
Some of this is scaffolding, and I would rather list the holes myself.
- Grok Build exposes no allowance percentage to hooks or the statusline, so a Grok session gets no warning before a refusal and will get none until the program exposes one. After a refusal, the record mostly serves the other two programs and the next takeover check.
- The thresholds are a first guess. I have no data on how often the warning fires, how often it fires too early, or how much time a takeover saves against manual reconstruction. I am not claiming savings; I am claiming a procedure that exists and is tested.
- A session that ignores the advisory simply runs into the native limit, the same as before. The harness never blocks, so it cannot force a wrap-up. In that case there is no record, and the state has to be rebuilt from the ledgers and the chat.
- Records are machine-local. Nothing is synced between machines. A limit hit on one machine is invisible on the other.
- The Codex transcript field and Grok’s silence are dated observations from 10 September 2026. Either may change with the next release, and the parity check was not built to notice a renamed transcript field; a silent format change would most likely show up as a warning that never fires.
Usage limits will recur. The intended improvement is a clean handover from a saved record, with current state checked before the successor continues.