Definition
What is a Software Factory?
A Software Factory coordinates agents, context, deterministic gates, and human decisions across the path from intent to production.
Written by Benedikt Stemmildt Founder & Co-CEO
Published: 2026-07-17
A working definition
A Software Factory is the delivery system that produces software changes. It receives intent, turns that intent into bounded work, gives agents the context and environment to execute, and accepts output only after defined checks. People decide what should exist, own the constraints, and remain accountable for production.
The Factory is wider than a coding agent. A coding agent works on a task. The Factory coordinates the path from that task to a deployed and observed change.
This distinction matters because faster implementation does not remove planning, review, deployment, or operational responsibility. It changes where engineering attention is needed.
The parts of the system
The names vary across implementations, but the responsibilities remain recognizable.
An intake layer turns an idea, issue, or specification into work that can be evaluated. It preserves the original intent and records the constraints that must survive decomposition.
An orchestrator builds the execution plan, assigns bounded work, manages dependencies, and decides what happens after success or failure. It should not hide state inside one conversation. The plan and run state need to remain inspectable.
Workers implement narrowly scoped changes inside a prepared environment. The environment supplies repository context, commands, architecture rules, and access boundaries.
Validators apply deterministic checks and focused review. Builds, types, tests, security checks, behavior comparisons, and deployment health provide different forms of evidence. Failed evidence sends work back through a bounded correction loop.
A feedback layer records what happened after merge and deployment. A repeated correction can become a test, a repository rule, or a better intake constraint. This is how the Factory improves rather than repeating the same mistake at higher speed.
Agent, harness, pipeline, and Factory
A harness wraps one kind of agent work with context, commands, permissions, and feedback. A Factory composes several harnesses and routes work between them.
A pipeline moves a change through a defined sequence. A Factory includes pipelines, but it also observes the result and changes how future work runs.
An agent can be useful without either. The Factory becomes relevant when an organization wants repeatable delivery across repositories, teams, or lifecycle stages.
Where people stay responsible
People define intent and decide what is worth building. They own architecture boundaries, production risk, and the acceptance evidence for high-impact work.
Automation should expand where the evidence is strong. Database schemas, public contracts, sensitive data, and irreversible operations usually deserve tighter human gates than a local refactoring. The boundary should follow the cost of a wrong change, not enthusiasm for autonomy.
OpenAI’s implementation account describes engineers spending more time on environments, specifications, and feedback loops as code production moved to agents. Ona reports a similar shift toward narrow automations connected across the delivery lifecycle. These accounts show that the pattern can run. They do not establish that the same autonomy level belongs in every organization.
How to know whether it is working
Lines of code and agent activity describe volume. They do not establish delivery value.
A Factory needs a baseline and a named outcome for the selected value stream. Engineering measures can include PR throughput, change confidence, failure rate, review time, and recovery time. The business measure depends on the work: migration progress, lead time for a product decision, support cost, or another result the organization can inspect.
The first useful Factory is bounded. One value stream gives the organization enough reality to test context, gates, ownership, and measurement before it expands the system.
Sources and limits
Practitioner account
Harness engineering: leveraging Codex in an agent-first world
- Method
- OpenAI's first-party account of building and operating an internal product with Codex.
- Limits
- A first-party implementation report from a model vendor, based on a new internal product and an unusually capable team.
2026-02-11 | Verified: 2026-07-17
Industry report
We built a software factory in 10 days
- Method
- Ona's public ten-day implementation report, including repository and run metrics.
- Limits
- A short, vendor-authored experiment on a greenfield product. It does not establish results for established enterprise systems.
2026-04-30 | Verified: 2026-07-17
Industry report
Building effective agents
- Method
- Anthropic's account of agent patterns observed in production work with customers.
- Limits
- Pattern guidance from a model provider, not a controlled comparison of software-delivery outcomes.
2024-12-19 | Verified: 2026-07-17
Related articles
Test the constraint in your system
The Factory Readiness Assessment connects this research to your value stream, your data, and the next decision.