The Human Is Not the Runtime

Automating PR-review handoffs through repository-state reconciliation while retaining human acceptance.

The Human Is Not the Runtime — AI

Putting an AI product into production changes the risk calculation immediately.

Production changes need review policies matched to their consequences. A high-impact module can justify deeper review, especially when mistakes are difficult to reverse.

Consider an illustrative workflow in which pull requests touching that module require additional review and explicit human acceptance.

Agents may write code, open pull requests, respond to comments, run tests and update tickets. If every handoff waits on a person noticing a notification and rebuilding context, an otherwise useful gate can accumulate unattended work.

The design question is how to preserve that review requirement while automating the mechanical handoffs.

⚡
Design objective: automate discovery and feedback transport while retaining substantive review and human acceptance. The workflow below is illustrative; its effect on throughput and safety must be measured.

The bottleneck is review transport, not review

Across the tooling that matters in 2026, agents can now carry most pre-merge stages under supervision — planning and decomposition, multi-file writing, test generation, bug reproduction and fix, opening the PR, and responding to review comments. Capability varies enormously by repo and task, but the direction is not in dispute. The place the loop conventionally stops is merge, which most teams keep behind an explicit human approval.

When changes arrive faster than reviewers can process them, a queue can grow. Waiting can increase sharply as arrivals approach review capacity; the relationship is not a fixed proportion of code-generation speed.

Where the loop standsWhat it means
Pre-merge stages can be assistedPlanning and decomposition, multi-file writing, test generation, bug reproduction and fix, opening the PR, and responding to review comments — the degree of supervision depends on the task, repository and policy.
Merge is notThis proposed workflow keeps an explicit human acceptance gate at merge; other tools and deployments can use different policies.
Measurement scopeDelivery research and this review design answer different questions. Evaluate this workflow using its own baseline and outcome measurements.

Delivery outcomes and review queues need separate measurements. A correlation between AI adoption and delivery metrics would not establish this particular queueing mechanism. Measure arrival rate, review capacity, unattended time and escaped defects for the workflow being evaluated.

One design to evaluate is to keep substantive review and human acceptance while automating discovery, transport and re-entry.


An illustrative reconciliation loop

The following is a proposed workflow, not a report of a private deployment.

What happensWho does it
1An implementation agent opens a PR containing the change rationale, risk and review context.implementation agent
2A scheduler periodically queries the authorized repository set for eligible reviews or revisions.review scheduler
3Two independently configured review systems inspect the same immutable commit and produce one consolidated review.review agents
4The implementation agent addresses accepted feedback, pushes a revision and updates the work record.implementation agents
5A later sweep sees a new head SHA needing re-review. Back to step 3.review scheduler

Push, review, revise and re-review can cross a repository boundary while the responsible humans inspect the evidence and decide whether to accept the result.

✋
Who merges? A human does in this proposed workflow. Agents discover eligible work, prepare reviews and track revisions; an authorized person reads the consolidated evidence and makes the acceptance decision.

A maintainer message can supply useful context: intent, risk and what deserves attention. It need not authorize execution. In this design, the trigger comes from repository state, while the PR body and documented invariants supply review context.


Repository state supplies the trigger

The obvious architecture is event-driven. A PR requests review, a webhook fires, a service validates it, an agent starts, a result gets posted. Immediate and efficient, on the whiteboard.

In production it is an integration surface: an endpoint, authentication, signature validation, secret rotation, durable receipt, retry handling, ordering, and observability for all of it. If delivery is at-most-once, one lost event strands a PR indefinitely. If it is at-least-once — which is what serious event infrastructure actually gives you — duplicates are normal and every downstream side effect has to be idempotent anyway.

Exactly-once processing is real inside a single transactional boundary. What you cannot get is one transaction spanning a webhook, an agent run, repository state and a posted review — four systems that share no commit. Retries, deduplication keys and idempotent effects stay on your plate regardless. Effectively-once is the honest name for what you can actually build.

🔑
For review eligibility specifically, re-query state before you act on a notification.

"Review requested" may already be false by the time the handler runs. The PR may have closed, gained three commits, changed reviewers, or already been reviewed. A correct handler queries current state anyway — which means the event is a hint about when to look, not an instruction to act. That is a real service (it cuts full scans and improves responsiveness) but it is not authority.

The sweep just starts from the authority. Its question is not "what event fired?" It is "which PRs require action from me right now?"

That reframing turns the workflow into a reconciliation loop, and this is the part worth stealing whether or not you care about code review. Desired state: no eligible PR is sitting unreviewed. Observed state: the current set of open PRs, head commits, review requests and prior reviews. Each sweep computes the difference and acts on it.

The useful Kubernetes analogy is state-based reconciliation. Watch events and retries trigger work; periodic authoritative queries can repair missed triggers where needed. An informer resync may reprocess cached objects rather than fetch fresh state, so it is not interchangeable with that repair query. Choose the combination according to latency, API load and recovery requirements. The engineering objective is convergence on the desired state.

Why this is quietly more robust

PropertyHow the proposed sweep supports it
A stable work identityThe review unit is repo + PR number + head commit SHA + reviewed base/diff identity + reviewer + policy version. Reuse completion only while that full work identity and its acceptance conditions remain current. A changed head or base creates new validation work. Being precise: this identifies the work, it does not by itself make the side effect idempotent — two overlapping sweeps, or a crash after posting but before recording completion, will still double-post. Use an atomic claim and a durable completion record. Reconcile an uncertain posting outcome using a durable remote work identifier before retrying. The identity is what makes those possible.
Recovery of durable outstanding workMiss three scheduled runs and there is no backlog to replay and no question about which messages vanished. The next sweep looks at reality and finds everything still waiting — provided the obligation is still visible. A closed PR or a withdrawn review request is gone, and no sweep will resurrect it.
Independence from the notification channelSlack can be down, a notification can be buried, someone can tag the wrong thread — none of it stops the loop, because Slack is not the queue; the repository is. That is channel independence, not the absence of single points of failure: the scheduler, the credentials, the forge API and the model providers are all still exactly that.
Natural batchingOne run discovers the entire review workload across every repo, instead of spawning a process per event burst.
Stale-review protectionA force-push moves the branch ref; it does not alter the commit you already reviewed. So pinning the reviewed SHA is necessary but not sufficient — the head can move between discovery and posting. The job captures the head SHA, reviews that immutable commit, then checks for a changed head before posting. This reread is not atomic with posting: attach the review to the immutable commit and validate the current integration candidate at acceptance. Record the reviewed base/diff identity and rerun affected checks when the base changes, even if the PR head is unchanged.

A slower trigger can make a faster system

🔥
For slow, stateful agent work, build the reconciliation path before depending on notifications.

Event-triggered reconciliation can reduce latency; a periodic authoritative repair query can recover durable work missed by the trigger path. Combining them adds API load and operational complexity, so choose the interval and trigger strategy from measured latency and recovery requirements. An hourly interval is an illustration, not a universal default.

Events tell you something happened. Reconciliation checks what still needs doing.

This is emphatically not a claim that polling wins everywhere. High-volume streaming, fraud signals, anything latency-critical — different problem, different answer, and an hourly sweep would be malpractice. But code review is durable, inspectable and slow, and the work stays visible until it is resolved. That is an unusually good fit for periodic reconciliation.

And the metric people optimise here is the wrong one. It is not milliseconds from review request to process start. It is how long an eligible PR stays unattended, including after partial failure. On that number, self-healing beats instant.


Two models is not one model twice

The review itself runs two models on one branch — Claude Opus 5 and OpenAI Codex over the same diff, independently, consolidated into a single comment. Strictly these are two independently configured review systems rather than two bare models: each brings its own harness, prompting and tooling, and Codex in particular wraps a model that can change under me. That distinction matters for what the pairing actually buys.

Run one model twice and you get two different answers. But randomness is not independence. The same model carries the same training distribution, the same architectural tendencies, the same learned conventions and the same blind spots into both runs. Sampling variation changes the path. It does not replace the map.

Different models may make different mistakes, but overlapping training data and similar harnesses can preserve correlated failures. Treat lower error correlation as a hypothesis to evaluate on representative defects, not a guarantee purchased by adding a second model.

What the two models doWhat it means
Both find itAgreement worth investigating; shared failure modes can still produce the same incorrect finding.
Only one finds itThe minority report is valuable precisely because the other model missed it. This is the case a single reviewer structurally cannot produce.
They disagreeThe disagreement is review output. It marks a place where the diff or its surrounding contract is ambiguous — which merits focused investigation.

Which means the consolidation step has real work to do. It must not flatten two reviews into a polite average, and it must not treat a majority vote as proof. Shared findings, unique findings and conflicts carry different weight and have to survive into the final comment with their provenance attached.

A second review system can produce a useful minority report, but that benefit must be tested. Compare findings against a labelled defect set and preserve disagreements; model diversity alone does not establish better malicious-code detection.


The transcript is part of the control surface

Design the evidence trail explicitly: repository diffs, exact reviewed commits, versioned inputs, tool results, review outputs and acceptance decisions. Ticket or chat comments can summarize progress, but summaries are not complete action logs.

Where audit requirements apply, the workflow must satisfy the applicable controls rather than merely produce more prose.

🧾
An instrumented workflow produces a more complete action trail than an uninstrumented one. That is the real comparison — not human versus agent.

An uninstrumented review, human or automated, records a conclusion. Files were scanned, architecture was recalled, something was run locally, a judgement formed — and what survives is "LGTM" plus four comments.

An automated pipeline is simply much easier to force into leaving the path: which commit was reviewed, which inputs were read, which tool calls ran, what each reviewer returned, how the feedback changed the code. One caveat worth stating plainly: model prose is not a record of model reasoning. An explanation can be post-hoc or incomplete. What you can trust is the action trail — versioned inputs, tool calls, outputs, decisions — not the narration wrapped around it.

A human workflow can be instrumented to the same standard. In practice, under time pressure, it usually is not.

And more logging is not automatically more auditability. Records need stable identities, timestamps, model and policy versions, links to exact commit SHAs, access control, retention rules, and a hard line between a generated summary and the underlying action record. On sensitive data the accurate framing is minimisation and authorisation rather than an absolute: regulated data can be processed by systems explicitly approved to handle it, under the contractual, access, retention and security controls that come with that. A workflow can deliberately exclude sensitive data from prompts and logs; that is an implementation boundary, not a general statement about what the law permits.


Where this breaks

An autonomous review loop can fail while looking busy. These are five failure modes to test for.

Failure modeWhat it looks like
Correlated agreementBoth models approve the same defect because both learned the same conventional pattern and neither knows the domain exception. The evidential value of agreement depends on reviewer accuracy and error correlation. Shared blind spots can substantially limit its added value.
Review theatreLong comments, broad checklists, many low-value suggestions, an impression of rigor, and no test of the dangerous invariant. Comment volume is not review quality.
Rubber-stampingModels are trained to be agreeable. A reviewer accepts a plausible explanation; an implementer satisfies the wording of a comment without addressing its cause. Two agreeable agents can run an immaculate ceremony around a bad change.
Context lossLarge diffs exhaust attention well before they exhaust a token limit, and the behaviour that matters often lives in callers, schemas, migrations or deploy config that never appears in the patch. Oversized changes to the critical module get decomposed or escalated.
Convergence without resolutionAgents alternate between locally reasonable fixes, introduce regressions, or rewrite code to satisfy each other. Iteration count and repeated findings need escalation thresholds. An autonomous loop has to know when to stop being autonomous.

And underneath all five, the oldest failure in automated controls: treating the existence of the gate as evidence that what passed through it is safe. A green check proves a configured process ran. It does not prove the process asked the right questions.

Humans still own the threat model, the critical-module boundary, the review policy, exception handling, model selection, access permissions and final acceptance of residual risk. They sample the output, challenge repeated agreement, study every defect that escapes, and rewrite the rules when production teaches something new. Moving the human out of the runtime does not move the responsibility out with them.


Evaluate the gate and the handoffs separately

Automating handoffs may reduce unattended queue time, but deeper review still costs wall-clock time and can produce false positives or repeated revisions. Compare those costs with accepted-result quality and escaped defects rather than assuming a free throughput improvement.

The design changes which layer humans operate at; its operational effect remains something to measure.

We decideThe agents execute
Which module gets exceptional scrutinyDiscovery — sweeping every repo, every hour
What the review pattern isTwo independent passes over the same diff
That repository state is authoritativeConsolidation into one comment
Which traces have to existFeedback transport, revision, re-entry
Judgment, exceptions, residual riskThe waiting

In the illustrative sequence, a PR appears, a sweep discovers it, independent reviews become one consolidated result, an implementer revises the code, and a later sweep checks the new head. Notifications improve visibility, but repository state remains the source of review eligibility.

A credible evaluation needs a baseline, comparable workloads, a stated observation window and defect outcomes. This article does not publish private rollout measurements or claim that throughput was unchanged.

The intended deliverable is an accepted change with a versioned record of the review process, subject to human judgment.