How to build an AI software factory: a step-by-step guide
Published October 2, 2026
Summary
A step-by-step guide to building an AI software factory: pick the first kind of work, make environments reproducible, write down how your team works, connect intake, set review gates, and measure before you scale. Written for engineering leads rolling coding agents out to a team.
Prerequisites
What to have in place first
A software factory amplifies the engineering system you already have. Agents do best in repositories where a new engineer could get productive quickly.
- A test suite that runs from one command and fails when something breaks.
- CI on every pull request, with branch protection so nothing merges without passing checks and a review.
- A repeatable setup: documented dependencies, seed data, and service startup.
- Written conventions agents can read, such as an AGENTS.md or CLAUDE.md file with build commands, code style, and what not to touch.
- Someone who owns the rollout and has time to review agent output for the first few weeks.
Step 1
Pick the first kind of work
Start narrow. Choose one kind of task that comes up often, is easy to verify, and is low risk if an agent gets it wrong.
- Avoid starting with greenfield features, changes across several services, or anything that needs a product decision.
| Kind of work | Why it fits | How you verify it |
|---|---|---|
| CI failures | Frequent and well defined; the failing job is the spec | The job passes and the fix explains the root cause |
| Review follow-ups | The reviewer's comments are the acceptance criteria | The reviewer approves without new comments |
| Dependency upgrades | Repetitive and mechanical | Builds and tests pass on the new version |
| Bug fixes with a reproduction | The reproduction defines done | The reproduction passes and a regression test is added |
| Dead code and stale flag cleanup | Low risk and easy to review | The diff only removes code and tests pass |
Step 2
Make the environment reproducible
An agent can only verify what it can run. Define the environment once and reuse it for every task.
- Repository and dependencies
- Clone, install, and build steps that run unattended. Cache whatever is slow.
- Services
- The databases, queues, and other services the tests need, started the same way every time, with seed data.
- Credentials and secrets
- Scoped, non-production credentials, synced from a secrets manager rather than pasted into prompts.
- Tools
- A browser for UI checks, MCP servers, and the CLIs the agent needs, such as your cloud provider or database client.
- Isolation
- One environment per task, so concurrent agents do not share ports, files, or state. Worktrees separate files; containers and VMs also separate processes.
Step 3
Write down how your team works
Agents follow written instructions far better than unwritten norms. Put them where every agent reads them, in the repository.
- Build, test, and lint commands, and which ones must pass before opening a pull request.
- Code conventions and patterns to reuse, with pointers to good examples.
- Boundaries: directories, migrations, or configuration that agents should not change without asking.
- What a good pull request looks like: a description, a linked issue, and screenshots for UI changes.
- Update the instructions whenever an agent repeats a mistake. This file is the factory's operating manual.
Step 4
Connect intake where work already starts
Do not ask people to learn a new place to file work. Route tasks from the tools they already use.
| Source | Example trigger | Good for |
|---|---|---|
| Issue tracker | Assign a Linear or GitHub issue to the agent | Small features and bug fixes |
| Chat | Mention the agent in a Slack thread | Quick fixes and investigations |
| Pull requests | A review comment asking for changes | Review follow-ups |
| CI | A failed build on the main branch | CI failures and flaky tests |
| Schedules | A weekly job | Dependency upgrades and cleanup |
| Webhooks | An alert from monitoring or another system | Incident investigation |
Step 5
Choose agents per kind of work
Different agents do better on different work, and the best one changes as models improve. Avoid wiring the factory to a single agent.
- Run the same five tasks through two or three agents and compare merge rate and review effort.
- Use model contracts you already hold where possible, such as an Anthropic, OpenAI, or Bedrock account.
- Record which agent and model produced each change, so you can compare them over time.
Step 6
Set review gates and limits
Gates keep a person accountable for what ships while agents do more of the work.
- Plan approval
- For larger tasks, review the agent's plan before it writes code.
- Verification before review
- Require tests and checks to pass, with screenshots or recordings for UI changes, before a person looks.
- Human merge
- Keep branch protection and a required reviewer on every agent pull request.
- Continue from feedback
- Let the agent address review comments in the same environment instead of starting over.
- Concurrency limits
- Cap how many agents run at once to what your reviewers can handle.
- Audit
- Keep a record of who started each task, from where, and what the agent did.
Step 7
Measure, then widen
Track a few numbers from the first week. When a kind of work merges reliably with little rework, add the next one, and move recurring work from one-off tasks to scheduled automations.
| Metric | What it tells you |
|---|---|
| Merge rate | The share of agent pull requests that merge. Low rates usually mean vague tasks or an incomplete environment. |
| Rework | How many review rounds a change needs before it merges. |
| Time to first review | Whether review capacity is keeping up with agent output. |
| Human interventions | How often someone had to step in mid-task. |
| Cost per merged change | Model and compute cost divided by the changes that actually merged. |
| Escaped defects | Bugs found after merge in agent-written code. |
Pitfalls
Common mistakes
Most failed rollouts trip on the same few things.
- Scaling agents before review capacity, so output piles up as unreviewed pull requests.
- Starting with vague tickets. Agents fill gaps with guesses, so write acceptance criteria first.
- Skipping the environment work. Agents that cannot run the tests produce plausible code that fails CI.
- Giving agents production credentials instead of scoped, non-production access.
- Counting pull requests opened instead of changes merged.
With Replicas
How the steps map to Replicas
Replicas provides the environment, intake, agent, and review layers, so the team can spend its time on the first kind of work rather than on infrastructure. Your existing CI/CD pipeline still decides what ships.
| Step | In Replicas |
|---|---|
| Environment | Hooks, variables, files, MCP servers, and skills configured once and inherited by every workspace; each task runs in its own Linux VM |
| Instructions | Your repository's instruction files, plus organization-level skills every agent receives |
| Intake | Slack, Linear, GitHub, GitLab, schedules, webhooks, and the API |
| Agents | Claude Code, Codex, Cursor, OpenCode, and other supported agents, chosen per task, on your own credentials where supported |
| Review | Live sessions, desktop and browser takeover, comments on diffs, and continuing from feedback in the same workspace |
| Measure | Organization analytics by person, agent, model, trigger, and cost |
| Governance | SCIM, roles, an audit log, and retention controls on enterprise plans |
FAQ
Building a software factory: questions
Getting started with Replicas
Start with one kind of work this week
Replicas includes a 14-day free trial with no credit card required. Connect a repository, define the environment once, and route your first kind of work, such as CI failures or review follow-ups, to an agent. Each task runs in its own cloud workspace and comes back as a pull request for your team to review.