AI Adoption Beyond Hype and Rejection

A practical way to evaluate AI workflows: separate capability, value and permission, compare complete outcomes, and revise decisions as evidence changes.

AI Adoption Beyond Hype and Rejection — AI

An AI adoption debate becomes useful when it reaches a concrete piece of work. Can an assistant produce a patch that meets the requirements? Does it reduce the effort needed to get that patch through review? What access does it need, and which decisions should remain with the engineer?

Those questions can have different answers for different tasks. A team might accept help with test generation, restrict access to private repositories and reject autonomous deployment. Each decision needs its own evidence.

The useful position is provisional: define the workflow, test a claim, inspect the failures and change the decision when the evidence changes. Enthusiasm and skepticism can both contribute to that process.

Separate capability, value and permission

Capability: can this system do the task under the conditions that matter? A convincing example establishes that an outcome is possible. It does not establish how often the system will succeed, how it behaves on unfamiliar inputs or how reliably failures will be caught.

Value: does the complete workflow improve the outcome? Faster first drafts may still require more review, correction or coordination. Count the effort needed to reach an acceptable result, including work transferred to someone else.

Permission: is this an acceptable way to do the work? Data access, tool permissions, review responsibilities and recovery need explicit boundaries. A useful assistant does not automatically need permission to send messages, merge changes or operate production systems.

Keeping these questions separate makes disagreement actionable. A failed quality check calls for a technical change or a narrower task. An unacceptable data flow calls for a different configuration or no deployment. Neither requires an explanation of a colleague’s motives.

Match the evidence to the claim

A public benchmark can help identify a capability worth testing. A product demonstration can suggest a workflow. A user survey can describe what respondents say they do or experience. Before treating any of those as evidence of an engineering benefit, examine the actual task, comparison and success criterion.

One useful caution comes from METR’s July 2025 randomized study. Experienced open-source developers working on familiar repositories took longer with the early-2025 AI tools studied, even though they believed the tools had helped. METR explicitly limited the result to that setting; it was not evidence that AI slowed every developer or would continue to do so as tools changed.

In its February 2026 follow-up, METR reported that selection effects and difficulties measuring concurrent agent work made the new experiment an unreliable estimate of the current productivity effect. Its researchers thought developers were likely benefiting more than in early 2025, while cautioning that the data provided only weak evidence about the size of that change.

The practical lesson is to keep the scope attached to the result. A study describes particular people, tasks, tools and measurement choices. Use it to improve the questions in a local evaluation, then report the local evaluation with the same care.

Write the decision before running the pilot

Choose a recurring task with a result that can be inspected. “Adopt AI across engineering” is too broad for a first evaluation. “Draft regression tests for a reproducible bug, with no access beyond a disposable checkout” is specific enough to examine.

Write down the current workflow and its acceptance criteria. Identify what a reviewer must check, which failures would invalidate the result and what evidence would justify expanding the trial. Set those criteria before seeing the output, so a polished demonstration does not quietly become the standard.

NIST’s AI Risk Management Framework 1.0 offers a broader vocabulary through its Govern, Map, Measure and Manage functions. It is a voluntary framework for considering AI risks across a system’s lifecycle. The lightweight engineering pilot below is a proposed approach, not a claim of NIST certification or compliance.

Keep the pilot record small enough to use, but specific enough that another engineer could understand the decision:

  • Task and boundary: name the work, allowed inputs, available tools and actions that remain outside scope.
  • Baseline: describe how comparable work is completed today, including review and correction.
  • Acceptance: define observable quality requirements and failures that cannot be traded for speed.
  • Cost: record active human effort, elapsed time and tool charges separately.
  • Ownership: name who inspects the output, accepts it and can stop the trial.
  • Decision: record what would justify adoption, a narrower workflow, further testing or rejection.

Measure the complete workflow

Run the comparison on representative work, including ordinary cases and known difficult ones. Give participants enough practice that the test is not simply measuring first-time setup. Keep the acceptance criteria consistent across approaches.

Where feasible, assign comparable tasks to the different approaches before work begins. Repeating the exact same task can introduce a learning advantage, so document that limitation if repetition is unavoidable. A small convenience sample can guide a local decision, but it should not be presented as a universal productivity result.

Record failures and abandoned attempts as well as accepted output. If an engineer stops using the assistant halfway through, that is part of the workflow’s cost. If the assistant proposes more work than the baseline, distinguish added value from a faster completion of the original task.

With parallel agents, keep wall-clock time separate from active human time. Ten minutes of waiting can overlap other work; ten minutes of reviewing competing patches cannot be counted as unattended progress. Report both measures and explain how they were collected.

Quality needs more than a completion flag. For code, inspect requirements, regression behavior, readability and integration with the surrounding system. Passing tests is useful evidence within the tests’ coverage; it does not settle requirements they never exercised.

A concrete example: regression-test assistance

Consider an illustrative pilot in which an assistant drafts tests for a reported parsing bug. This is a proposed evaluation, not a description of measured results.

Give both approaches the same bug report, expected behavior and repository snapshot. The assistant works in an isolated checkout without production credentials. Its output is a proposed test change; an engineer retains responsibility for reviewing and accepting it.

A useful acceptance check is that the regression test fails against the buggy implementation and passes against the intended fix. Also inspect whether it fails for the right reason, covers the reported behavior and avoids encoding an incidental implementation detail.

Then record the effort needed to reach the accepted test: setup, writing or prompting, execution, review and correction. Log discarded suggestions. Check that the test suite remains reproducible and that unrelated behavior was not weakened to make the new test pass.

If the assistant produces useful cases but too much cleanup, a narrower trial might ask it only to propose edge cases. If the workflow improves end-to-end effort while meeting the quality checks, expand to another comparable task class. Neither result settles whether the assistant should edit application code or deploy a service; those are separate evaluations.

Make disagreement specific

Ask an advocate which outcome should improve and what result would change that expectation. Ask a skeptic which failure is unacceptable and how the trial could detect it. Apply the same standard to both answers.

Concerns about confidentiality, reliability, ownership or maintainability belong in the evaluation. So do proposed benefits such as clearer documentation, broader test coverage or reduced repetitive work. Translate each concern or benefit into an observable check where possible, and document what cannot yet be measured.

Some boundaries are requirements rather than averages to optimize. A workflow that exposes prohibited data does not become acceptable because its average completion time looks good. Conversely, a failed experiment on one task is a reason to revise that task’s decision, not a basis for judging every possible use.

Keep policy claims equally specific. Before saying that a rule helps or harms smaller teams, examine the actual requirement, who must satisfy it and what alternatives are available. An adoption pilot cannot establish a hidden motive behind public disagreement, and does not need one to produce a useful decision.

Choose an outcome, and a reason to revisit it

Adopt within the tested boundary. The workflow meets the agreed quality and access requirements, and the observed benefit warrants its cost. Record the supported task classes and retain the review steps used in the pilot.

Narrow the scope. A component is useful, but the complete workflow is unreliable or expensive to supervise. Keep the useful part and test the smaller claim.

Continue evaluating. The evidence is too limited or confounded to decide. State what is missing and the next test that would resolve it. Uncertainty should produce a question with an owner, not an indefinite rollout.

Do not deploy this workflow. It fails an essential requirement or offers no worthwhile improvement under the tested conditions. Preserve the evidence so the same proposal is not repeatedly debated without new information.

Date the decision and record the model, tool configuration and relevant constraints. Revisit it when those inputs change materially, when a failure exposes a missing check or when the task itself changes. A newer model is a reason to retest a claim, not proof that the earlier result has reversed.

AI adoption earns credibility through decisions that others can inspect: what was attempted, how it was evaluated, where it failed and why it is being used. Keep those decisions open to correction. The goal is useful, dependable engineering work.