The Harness: Two AIs, Zero Trust
A nine-minute documentary about everything around the models: one rulebook for Claude and Codex, shared memory, cross-provider review, account guards, clean handovers and loops that know how to stop.
A background run once waited an hour and thirty-nine minutes on a single confirmation prompt. A finished loop once came back to life and kept working. Failures like these shaped the rules in the film above.
The models are the easy part: Claude Opus 5.5 at max effort and Codex on GPT-6 Astra at ultra. Choosing them leaves the hard questions open. Who checks the work? Which account pays for it? What survives a limit, a restart or a stale terminal tab? Who is allowed to call a task finished? The harness is my answer, and the documentary is its tour, chapter by chapter.
What the film shows
- Two minds. Both models start at medium effort unless told otherwise, so every pin here says maximum. Claude's settings file silently drops
max, so it lives in an environment variable. A new Codex model must beat Astra at maximum and cost less, which is why GPT-6 Sol was rejected on launch day. - One rulebook. One shared policy generates both global instruction files. In the knowledge base
CLAUDE.mdis the source and everyAGENTS.mdis a thin forwarder. Slash commands become generated Codex skills stamped with the hash of their source, so drift fails a check. - Total recall. A local search engine (QMD) indexes 3,646 markdown files with three small local models and serves both agents from one daemon. A hook blocks grep over the knowledge folders. A decision graph was tested against plain search and did not earn an expansion.
- The tribunal. Claude writes, Codex reviews, and the reverse. The reviewer is a fresh text-only process with the full evidence inline. A limit allows a disclosed same-provider fallback. A harsh verdict never does.
- The guards. Two Claude accounts on one laptop, routed by folder, verified by email and organization, enforced by a hook that blocks a wrong-account prompt before the model sees a token.
- Out of fuel. When a provider refuses, the agent gets a
LIMIT WATCHtelling it to wrap up, write a takeover record and point to the other agent. - The loop. Headless iterations start with an empty context, a separate oracle decides "done", and a changed verifier file stops everything.
- The mirror. A private ledger measures every observed response and leaves three fields honestly blank: actual charges, accepted outcomes and quality.
Steal this
- Keep one canonical rulebook, generate each agent's native files from it and fail a check on drift.
- Give the reviewer complete evidence in a fresh, tool-free process. Disclose fallbacks, and never switch reviewers to escape a harsh verdict.
- Route accounts before launch, verify the identity literally and block the prompt, not the bill.
- Write handovers to disk: task, verified work, next steps, open risks. Make the next agent re-derive the rest from the repository.
- Separate the worker from the judge, hash the judge, and make STOP survive every restart.
A successful process exit cannot certify the work.