The Human Is Not the Runtime
Automating PR-review handoffs through repository-state reconciliation while retaining human acceptance.
Putting an AI product into production changes the risk calculation immediately.
Production changes need review policies matched to their consequences. A high-impact module can justify deeper review, especially when mistakes are difficult to reverse.
Consider an illustrative workflow in which pull requests touching that module require additional review and explicit human acceptance.
Agents may write code, open pull requests, respond to comments, run tests and update tickets. If every handoff waits on a person noticing a notification and rebuilding context, an otherwise useful gate can accumulate unattended work.
The design question is how to preserve that review requirement while automating the mechanical handoffs.
The bottleneck is review transport, not review
Across the tooling that matters in 2026, agents can now carry most pre-merge stages under supervision — planning and decomposition, multi-file writing, test generation, bug reproduction and fix, opening the PR, and responding to review comments. Capability varies enormously by repo and task, but the direction is not in dispute. The place the loop conventionally stops is merge, which most teams keep behind an explicit human approval.
When changes arrive faster than reviewers can process them, a queue can grow. Waiting can increase sharply as arrivals approach review capacity; the relationship is not a fixed proportion of code-generation speed.
Delivery outcomes and review queues need separate measurements. A correlation between AI adoption and delivery metrics would not establish this particular queueing mechanism. Measure arrival rate, review capacity, unattended time and escaped defects for the workflow being evaluated.
One design to evaluate is to keep substantive review and human acceptance while automating discovery, transport and re-entry.
An illustrative reconciliation loop
The following is a proposed workflow, not a report of a private deployment.
Push, review, revise and re-review can cross a repository boundary while the responsible humans inspect the evidence and decide whether to accept the result.
A maintainer message can supply useful context: intent, risk and what deserves attention. It need not authorize execution. In this design, the trigger comes from repository state, while the PR body and documented invariants supply review context.
Repository state supplies the trigger
The obvious architecture is event-driven. A PR requests review, a webhook fires, a service validates it, an agent starts, a result gets posted. Immediate and efficient, on the whiteboard.
In production it is an integration surface: an endpoint, authentication, signature validation, secret rotation, durable receipt, retry handling, ordering, and observability for all of it. If delivery is at-most-once, one lost event strands a PR indefinitely. If it is at-least-once — which is what serious event infrastructure actually gives you — duplicates are normal and every downstream side effect has to be idempotent anyway.
Exactly-once processing is real inside a single transactional boundary. What you cannot get is one transaction spanning a webhook, an agent run, repository state and a posted review — four systems that share no commit. Retries, deduplication keys and idempotent effects stay on your plate regardless. Effectively-once is the honest name for what you can actually build.
"Review requested" may already be false by the time the handler runs. The PR may have closed, gained three commits, changed reviewers, or already been reviewed. A correct handler queries current state anyway — which means the event is a hint about when to look, not an instruction to act. That is a real service (it cuts full scans and improves responsiveness) but it is not authority.
The sweep just starts from the authority. Its question is not "what event fired?" It is "which PRs require action from me right now?"
That reframing turns the workflow into a reconciliation loop, and this is the part worth stealing whether or not you care about code review. Desired state: no eligible PR is sitting unreviewed. Observed state: the current set of open PRs, head commits, review requests and prior reviews. Each sweep computes the difference and acts on it.
The useful Kubernetes analogy is state-based reconciliation. Watch events and retries trigger work; periodic authoritative queries can repair missed triggers where needed. An informer resync may reprocess cached objects rather than fetch fresh state, so it is not interchangeable with that repair query. Choose the combination according to latency, API load and recovery requirements. The engineering objective is convergence on the desired state.
Why this is quietly more robust
A slower trigger can make a faster system
Event-triggered reconciliation can reduce latency; a periodic authoritative repair query can recover durable work missed by the trigger path. Combining them adds API load and operational complexity, so choose the interval and trigger strategy from measured latency and recovery requirements. An hourly interval is an illustration, not a universal default.
Events tell you something happened. Reconciliation checks what still needs doing.
This is emphatically not a claim that polling wins everywhere. High-volume streaming, fraud signals, anything latency-critical — different problem, different answer, and an hourly sweep would be malpractice. But code review is durable, inspectable and slow, and the work stays visible until it is resolved. That is an unusually good fit for periodic reconciliation.
And the metric people optimise here is the wrong one. It is not milliseconds from review request to process start. It is how long an eligible PR stays unattended, including after partial failure. On that number, self-healing beats instant.
Two models is not one model twice
The review itself runs two models on one branch — Claude Opus 5 and OpenAI Codex over the same diff, independently, consolidated into a single comment. Strictly these are two independently configured review systems rather than two bare models: each brings its own harness, prompting and tooling, and Codex in particular wraps a model that can change under me. That distinction matters for what the pairing actually buys.
Run one model twice and you get two different answers. But randomness is not independence. The same model carries the same training distribution, the same architectural tendencies, the same learned conventions and the same blind spots into both runs. Sampling variation changes the path. It does not replace the map.
Different models may make different mistakes, but overlapping training data and similar harnesses can preserve correlated failures. Treat lower error correlation as a hypothesis to evaluate on representative defects, not a guarantee purchased by adding a second model.
Which means the consolidation step has real work to do. It must not flatten two reviews into a polite average, and it must not treat a majority vote as proof. Shared findings, unique findings and conflicts carry different weight and have to survive into the final comment with their provenance attached.
A second review system can produce a useful minority report, but that benefit must be tested. Compare findings against a labelled defect set and preserve disagreements; model diversity alone does not establish better malicious-code detection.
The transcript is part of the control surface
Design the evidence trail explicitly: repository diffs, exact reviewed commits, versioned inputs, tool results, review outputs and acceptance decisions. Ticket or chat comments can summarize progress, but summaries are not complete action logs.
Where audit requirements apply, the workflow must satisfy the applicable controls rather than merely produce more prose.
An uninstrumented review, human or automated, records a conclusion. Files were scanned, architecture was recalled, something was run locally, a judgement formed — and what survives is "LGTM" plus four comments.
An automated pipeline is simply much easier to force into leaving the path: which commit was reviewed, which inputs were read, which tool calls ran, what each reviewer returned, how the feedback changed the code. One caveat worth stating plainly: model prose is not a record of model reasoning. An explanation can be post-hoc or incomplete. What you can trust is the action trail — versioned inputs, tool calls, outputs, decisions — not the narration wrapped around it.
A human workflow can be instrumented to the same standard. In practice, under time pressure, it usually is not.
And more logging is not automatically more auditability. Records need stable identities, timestamps, model and policy versions, links to exact commit SHAs, access control, retention rules, and a hard line between a generated summary and the underlying action record. On sensitive data the accurate framing is minimisation and authorisation rather than an absolute: regulated data can be processed by systems explicitly approved to handle it, under the contractual, access, retention and security controls that come with that. A workflow can deliberately exclude sensitive data from prompts and logs; that is an implementation boundary, not a general statement about what the law permits.
Where this breaks
An autonomous review loop can fail while looking busy. These are five failure modes to test for.
And underneath all five, the oldest failure in automated controls: treating the existence of the gate as evidence that what passed through it is safe. A green check proves a configured process ran. It does not prove the process asked the right questions.
Humans still own the threat model, the critical-module boundary, the review policy, exception handling, model selection, access permissions and final acceptance of residual risk. They sample the output, challenge repeated agreement, study every defect that escapes, and rewrite the rules when production teaches something new. Moving the human out of the runtime does not move the responsibility out with them.
Evaluate the gate and the handoffs separately
Automating handoffs may reduce unattended queue time, but deeper review still costs wall-clock time and can produce false positives or repeated revisions. Compare those costs with accepted-result quality and escaped defects rather than assuming a free throughput improvement.
The design changes which layer humans operate at; its operational effect remains something to measure.
In the illustrative sequence, a PR appears, a sweep discovers it, independent reviews become one consolidated result, an implementer revises the code, and a later sweep checks the new head. Notifications improve visibility, but repository state remains the source of review eligibility.
A credible evaluation needs a baseline, comparable workloads, a stated observation window and defect outcomes. This article does not publish private rollout measurements or claim that throughput was unchanged.
The intended deliverable is an accepted change with a versioned record of the review process, subject to human judgment.