Skip to content
Insights

Engineering practice

Verification in a Software Factory

Once generation becomes cheap, Factory throughput depends on the system's ability to establish that a change is acceptable, safe, and worth shipping.

Portrait of Benedikt Stemmildt

Written by Benedikt Stemmildt Founder & Co-CEO

Published: 2026-07-17

Generation moves the constraint

Code generation can increase the number and size of changes entering a delivery system. Review, integration, and product decisions do not automatically gain the same capacity.

The result is a queue. Pull requests wait longer, checks run more often, and people receive more plausible output than they can judge. A Factory has to improve verification before it increases autonomy.

Deterministic gates establish fixed facts

Builds, type checks, linters, tests, schema checks, security scans, and policy rules answer bounded questions. The same input should receive the same verdict.

These gates work best when they protect a contract that existed before implementation. A test written after the change can encode the same misunderstanding as the code. Acceptance evidence should remain independent where the risk justifies it.

Useful independent evidence includes an existing API contract, a recorded behavior comparison, a golden output, a property that must hold, or an end-to-end scenario maintained outside the implementation task.

Agent review has a different job

Agent review can inspect intent, architecture, security, maintainability, and missing tests at a volume a single human cannot match. It remains probabilistic. Several reviewers can repeat the same mistaken assumption or flood the correction loop with broad suggestions.

A production review stage needs a shared severity vocabulary and bounded scope. Critical findings can block. Warnings can trigger rework. Questions should reach a person. Suggestions should not consume the same budget as a broken contract.

Review loops also need a stop condition. A loop that cannot converge after a small number of attempts has found uncertainty, not an invitation to spend tokens indefinitely.

Green does not mean correct

A change can satisfy every automated check and still solve the wrong problem. It can preserve safety properties while missing the useful behavior. It can pass unit and integration tests while introducing a product decision nobody intended.

People remain responsible for intent and for the residual judgment that the gates cannot encode. The Factory should surface the work that deserves that attention instead of asking people to read every generated line.

Gate by reversibility and impact

Not every change deserves the same process.

A local refactoring with no public contract is cheap to undo. A database schema, public API, permission boundary, or production-data migration becomes expensive once other systems depend on it. Human attention belongs where a wrong decision is hard to reverse.

This creates practical operating modes. Low-risk work can pass through automated gates and merge. Higher-risk work can require an approved plan, a reviewed contract, or a person at deployment. The organization owns those boundaries.

Make the Factory legible

Agents need access to the application state they are expected to evaluate. OpenAI describes making user interfaces, logs, metrics, and traces legible to Codex inside isolated worktrees. That allowed agents to reproduce failures and verify runtime behavior instead of stopping at source code.

Every run should leave evidence: the original intent, plan, changes, gate results, review findings, deployment state, and cost. A repeated failure can then become a durable rule or test.

Verification capacity sets the safe speed of the Factory. More agents help only when the evidence system can keep up.

Sources and limits

Practitioner account

Harness engineering: leveraging Codex in an agent-first world

Method
OpenAI's first-party description of agent legibility, review, testing, and observability in an internal product.
Limits
A model-vendor account from a greenfield product with custom infrastructure and a small specialist team.

2026-02-11 | Verified: 2026-07-17

Industry report

Effective harnesses for long-running agents

Method
Anthropic's engineering account of state, progress, and acceptance handling for long-running agents.
Limits
Implementation guidance from a model provider, not an independent outcome comparison.

2025-11-26 | Verified: 2026-07-17

Telemetry report

The 2026 State of Software Delivery

Method
Analysis of more than 28 million CI workflows across thousands of teams.
Limits
CI telemetry shows integration behavior, not whether a feature matches product intent.

2026-02-18 | Verified: 2026-07-17

Related articles

Test the constraint in your system

The Factory Readiness Assessment connects this research to your value stream, your data, and the next decision.

See the assessment