Skip to content
Insights

Evidence review

Software Factory evidence and its limits

Public implementations show that agent-run delivery systems can produce and ship software. Their numbers need source labels, scope, and careful limits.

Portrait of Benedikt Stemmildt

Written by Benedikt Stemmildt Founder & Co-CEO

Published: 2026-07-17

What public implementations establish

Public Factory-like implementations establish feasibility. Agents can work in parallel, open pull requests, run checks, respond to failures, deploy changes, and continue operating after merge.

They also show where people move. Engineers spend more time specifying intent, designing environments, improving feedback, and deciding which output deserves attention.

They do not establish one transferable productivity number. Most reports come from model or platform vendors, greenfield products, and teams that can build custom infrastructure around the agents.

OpenAI’s internal product

OpenAI reports building an internal beta from an empty repository with no manually written code. After five months, the repository had roughly one million lines and about 1,500 merged pull requests. The report gives an average of 3.5 pull requests per engineer per day and says hundreds of internal users used the product.

The most useful evidence is operational. The team had to make the application, logs, metrics, and development environment legible to agents. Human QA capacity became a constraint. Engineers worked on scaffolding and feedback instead of contributing code directly.

The report does not provide a controlled comparison for quality or business value. OpenAI has direct access to its models and a team selected for this experiment. The numbers show what this system produced, not what another organization should forecast.

Ona’s ten-day Factory

Ona reports 375 merged pull requests, more than 67,000 lines, 1,067 tests, and an 87 percent autonomous rate over ten days. Sixteen narrow automations covered planning, implementation, review, deployment, monitoring, and incident handling.

The architecture matters more than the code count. Ona did not run one general agent. It connected bounded automations with clear triggers and handoffs.

The experiment was short and greenfield. Repository activity does not establish useful product adoption, long-term changeability, or results in a regulated enterprise. Ona sells the platform used for the work.

Delivery research supplies the missing context

DORA and CircleCI examine broader populations rather than one Factory implementation. DORA associates better results with the surrounding organizational system. CircleCI’s telemetry shows that increased development activity can arrive with weaker main-branch flow and stability for the median team.

These sources explain why implementation reports need delivery measures beside output measures. A Factory can generate many accepted changes and still move the constraint into review, integration, recovery, or product decision-making.

How h&w uses this evidence

h&w treats external reports as evidence for the category and its mechanisms. They are not evidence of an h&w client outcome.

Our public client testimonials come from Training, Pilot, and Rollout engagements. Each one carries the matching product tag. The internal record keeps the narrower claim and limitations. No current testimonial belongs to Assessment. The internal Factory dashboard is an operational snapshot, not a client result or a live feed.

A client engagement needs its own baseline, named data source, and business outcome. External benchmarks can inform the questions. They cannot answer them for the client.

What stronger evidence would look like

Stronger Factory evidence would connect a defined value stream to delivery and business measures over time. It would preserve the baseline, show the intervention, record review and failure behavior, and state what else changed during the same period.

It would also publish negative or neutral results. A credible evidence base needs to show where a Factory failed to improve flow, where cost outweighed benefit, and where human coordination remained the binding constraint.

Sources and limits

Practitioner account

Harness engineering: leveraging Codex in an agent-first world

Method
OpenAI's account of an internal product built from an empty repository with Codex.
Limits
First-party report, no controlled comparison, new product, specialist team, and model-provider access.

2026-02-11 | Verified: 2026-07-17

Industry report

We built a software factory in 10 days

Method
Ona's public implementation report with repository, timing, test, and autonomy metrics.
Limits
Vendor-authored, ten-day greenfield experiment. Lines and pull requests do not establish customer value.

2026-04-30 | Verified: 2026-07-17

Survey

State of AI-assisted Software Development 2025

Method
Cross-organization research on AI use, platform quality, policies, and delivery conditions.
Limits
Survey associations and self-reported measures cannot prove that one Factory design caused an outcome.

2025 | Verified: 2026-07-17

Telemetry report

The 2026 State of Software Delivery

Method
Analysis of more than 28 million CI workflows across thousands of teams.
Limits
CircleCI-user telemetry covers integration activity and stability, not product outcomes.

2026-02-18 | Verified: 2026-07-17

Related articles

Test the constraint in your system

The Factory Readiness Assessment connects this research to your value stream, your data, and the next decision.

See the assessment