Make Love Not War

Two rival frontier models, both at maximum reasoning, in one autonomous engineering loop. Why independent review, observability, and verified output matter more than picking a side.

Make Love Not War — AI

The frontier-model war is the best show in tech. Anthropic versus OpenAI, benchmark leaderboards, new releases, weekly arguments about who's number one. The labs have to compete. On your own machine, you don't have to pick a side. You run both.

Right now I run Anthropic's Fable 5 and OpenAI's Codex GPT-5.6 Sol at the same time — both at maximum reasoning, both full-auto, no per-action approvals — inside one autonomous loop. They don't compete. They collaborate. The war stays upstream, at the labs. Downstream, on my machine, it's cooperation.

The scoreboard isn't your architecture

Model comparisons matter when you're choosing one model. They matter less once your system can use more than one.

A leaderboard asks which model gives the better answer on a fixed test. An autonomous engineering loop has a different problem: producing reliable work across planning, implementation, verification, infrastructure, testing, and delivery. Nothing says all of those jobs belong to the same model. Model choice still matters; test whether combining models improves the complete workflow.

Two models, one loop

Two frontier models from two rival labs have different training, different failure modes, different blind spots. That's the entire point. Run them together and you get what no single model can hand you: a second opinion from a different model family, with its own training and its own blind spots.

In practice it's a division of labor. Fable 5 advises while Codex executes — reserve the expensive reasoning for the decisions with the most leverage, then let the executor carry out the work. One model proposes an architecture; the other challenges it. One implements; the other checks the result. When both independently reach the same design, that's signal. When they disagree, that's a flag worth reading before a line gets written.

I've written about running a single model off the leash. This is the next step: two of them, off the leash, at once, consolidating before anything merges.

Autonomy needs receipts

Full-auto changes the control model. The loop doesn't stop before each action to ask permission, and it can run for hours or days with nobody watching. That's exactly why observability is part of the system, not a dashboard bolted on later.

Ticket comments can summarize progress, decisions and open questions. Because the agent generates that narration, it can omit or misdescribe actions. Treat comments as a navigation aid to the evidence, not a complete audit record.

Audit evidence should come from independently recorded actions, versioned inputs, tool results, test outcomes and exact commits. Preserve the link from each claim to the underlying artifact, including failures and interrupted runs.

Humans retain accountability for authorization and acceptance. Automation can collect evidence and make review easier; it cannot move that responsibility into a generated ticket comment.

Fast and good at the same time

The aim is to improve speed and accepted quality together. That requires comparison with a single-model baseline, not an assumption that adding a model improves both.

The ambition is complete engineering work: architecture, implementation, tests, infrastructure, applied research, and verification. Generated volume alone does not establish that the result is useful or that a smaller team could not have produced it more efficiently.

Speed alone was never the hard part; anyone can generate a mountain of bad code fast. A large pile of plausible code has limited value. A complete solution — architecture, implementation, tests, infrastructure, verification, reporting, retained knowledge — is a different unit of output. The loop isn't typing faster. It's compressing an engineering project.

Compare complete workflow costs

Running multiple frontier models continuously can consume substantial resources. The relevant comparison is the complete cost of a verified deliverable, including model use, infrastructure, human review, and rework.

A subscription price and an engineering budget measure different things. Neither establishes productivity on its own.

MeasureWhat to compare
CostModels, infrastructure, human review, and rework
ThroughputAccepted deliverables against the same specification
ReliabilityDefects, failed runs, and recovery work
Human attentionTime spent supervising and correcting the workflow

The useful comparison is between workflows that meet the same acceptance criteria. Measure completed work, elapsed time, human attention, reliability, and total cost together. Continuous operation is a capability; it is not proof that the output replaces a team.

The All-In episode with Jensen Huang published March 19, 2026 includes a chapter beginning at 20:48 on revenue capacity and employee token allocation. It frames compute as an engineering input; it does not establish a universal spending target.

💸
Evaluate the complete workflow: more compute is useful when the additional accepted value justifies its full cost.

Compute can be a productive engineering input rather than a cost to minimize in isolation. Whether more of it creates value depends on the work accepted, the quality retained, and the alternatives available.

Where this goes

The same question sits underneath predictions about a one-person billion-dollar company: how much useful work can one accountable operator coordinate?

A “unicorn” describes a startup valuation of at least $1 billion, not a headcount. In 2024, Sam Altman discussed the possibility of a single-person billion-dollar company with Alexis Ohanian; Ohanian’s own follow-up confirms that discussion. Treat it as a prediction, not evidence that a particular workflow produces that outcome.

The mechanism worth testing is one person coordinating multiple models across a complete workflow. Scaling that into a company still requires customers, distribution, reliable operations, and accountable decisions. Model output does not establish a valuation.

Make them work

Combining models is an option to evaluate, not a reason to spend more by default. Compare the combination with each model alone on accepted results, human attention and total cost. Keep the workflow that meets the requirements.

Make love, not war — at least downstream.