Research synthesis
Why faster coding does not guarantee faster delivery
Research measures task time, perceived savings, pull requests, and production flow differently. Reading those measures together reveals where the constraint moves.
Written by Benedikt Stemmildt Founder & Co-CEO
Published: 2026-07-17
The studies are measuring different things
Developer productivity research often appears contradictory because the headline numbers answer different questions.
Task completion time measures how long one person needs for a bounded task. Self-reported time savings measure perception. Pull-request throughput measures accepted changes. CI activity measures attempts to integrate work. Main-branch success measures whether those attempts pass. None of these measures alone tells an organization whether customers received value sooner.
Comparisons need the population, task, measurement window, and quality guardrail beside every number.
What METR found, and what changed
METR’s early-2025 randomized study found that 16 experienced open-source developers took 19 percent longer when they could use the available products. The developers expected to be faster and still believed they had been faster after completing the work.
The setting matters. These developers worked in mature repositories they knew well. The tasks required context that small coding benchmarks usually remove.
METR started a larger follow-up with newer products. In February 2026 it reported that the new design could no longer produce a reliable current speed estimate. Developers increasingly refused tasks that prevented agent use, and concurrent agents made time measurement harder. The raw data suggested improvement, but METR called the evidence for its size weak.
The honest reading includes both results. Early products slowed this specialist group in one realistic setting. Newer products appear more useful, while the current effect remains difficult to estimate with the same method.
The delivery system can absorb local gains
DORA’s 2025 research describes AI-assisted development as an amplifier of the surrounding system. Teams still need a healthy platform, clear policies, user focus, and reliable feedback. Faster code production adds little when work waits for review, integration, a product decision, or a release window.
CircleCI’s 2026 telemetry makes the downstream constraint visible. Across more than 28 million workflows, average activity increased 59 percent year over year. The median team’s feature-branch activity increased while main-branch throughput fell. Main-branch success reached 70.8 percent, its lowest level in five years in that dataset.
The top-performing group behaved differently. Its change volume and main-branch throughput grew together. The report associates that result with validation and recovery capacity that kept pace with generation.
Adoption is not a result
DX reports 93 percent industry adoption in its Q1 2026 company sample. That makes users versus non-users a weak comparison. The more useful question is what changes within a team as usage deepens.
DX found uneven throughput and volatile quality. Some teams improved. Others saw more defects or weaker maintainability. Self-reported time savings were larger than longitudinal PR-throughput changes.
An adoption target can show that a product is present. It cannot establish that the organization delivers sooner or with more confidence.
A measurement set that survives scrutiny
Start with the value stream, not the product license. Record a baseline before changing the system.
Pair an output signal such as PR throughput with review time, rework, failure rate, and recovery time. Add developer confidence because a system can look healthy while the people operating it lose understanding. Name the business outcome separately.
The measures should explain where work waits and whether the constraint moved. That is more useful than searching for one universal productivity multiplier.
Two leading indicators from a practitioner interview
In a practitioner interview, Patrick Debois proposes two leading indicators for teams running coding agents. The first is human interventions per accepted change, segmented by risk or change type. A high-risk migration and a copy tweak carry different intervention counts, so a single average hides the pattern. The second is adoption and reuse of shared context, skills, and harness components. When one team’s setup spreads and gets reused, the investment compounds. When every team rebuilds from scratch, it does not.
Neither is an outcome, and neither is safe as a target on its own. Set interventions as a target and a team can ship unreviewed risk. Force reuse and it can standardize on the wrong pattern. Read both beside the flow, review, rework, failure, recovery, confidence, cost, and business-outcome measures above. Debois’s account points to where to look. It is one practitioner’s view, not independent evidence that these signals predict delivery.
Sources and limits
Independent research
Measuring the impact of early-2025 AI on experienced open-source developer productivity
- Method
- Randomized controlled trial with 16 experienced open-source developers completing 246 tasks in repositories they knew.
- Limits
- Early-2025 products, a small specialist population, and issue completion rather than organization-level delivery.
2025-07-10 | Verified: 2026-07-17
Independent research
We are changing our developer productivity experiment design
- Method
- METR's follow-up experiment and methodological review of late-2025 product data.
- Limits
- Selection effects and concurrent-agent use made the new speed estimate unreliable. METR describes the evidence as weak.
2026-02-24 | Verified: 2026-07-17
Survey
State of AI-assisted Software Development 2025
- Method
- Google DORA's cross-organization research on AI use and software-delivery systems.
- Limits
- Survey-based associations do not isolate one causal intervention and can differ from repository telemetry.
2025 | Verified: 2026-07-17
Telemetry report
The 2026 State of Software Delivery
- Method
- Analysis of more than 28 million CircleCI workflows across thousands of teams.
- Limits
- The data covers CircleCI users and CI activity. It does not observe product value or every part of delivery.
2026-02-18 | Verified: 2026-07-17
Industry report
AI-assisted engineering: Q1 impact report
- Method
- DX synthesis of survey and system data from more than 400 companies.
- Limits
- DX sells developer-intelligence products, cohorts differ, and several measures are self-reported.
2026-Q1 | Verified: 2026-07-17
Related articles
Test the constraint in your system
The Factory Readiness Assessment connects this research to your value stream, your data, and the next decision.