Ben Holmes, a developer at Warp, joined Hugo Bowne-Anderson and Eleanor Berger for Factory Life to demonstrate Warp Factories, Warp’s platform for running and coordinating coding agents. It connects agents to a team’s repositories and tools, accepts work from messages, issues or scheduled jobs, and helps the team review results and test improvements.
The Warp team uses it to build Warp Terminal and Warp Factories itself. Ben started with a whiteboard, then opened the team’s working dashboards and examples. The demos covered how Warp routes work to agents, brings the results back for human review and improves the factory over time. They included scoring rubrics, proposed changes to instructions and benchmarks built from the team’s own tasks.
Hugo Bowne-Anderson and Eleanor Berger teach Build Your Agentic Software Factory: a practical course on directing agents, checking their work and turning an idea into a working application.
One Factory the Team Can Improve Together
The architecture Ben sketched starts with a foreman: the agent the team talks to when it wants to put work into the factory. It coordinates what happens next, doing a task itself or delegating parts to other agents. Warp’s example adds triage, implementation and code review. A request can move through research and a specification, implementation, then a few rounds of review before a pull request reaches a person.
This gives the team a single place to agree on the skills, integrations, permissions and model choices those agents use. Improvements to that setup can then reach everyone’s work. There is a clear place to investigate when the factory keeps making the same mistake.
At the time of the talk, the whiteboard assigned Opus 5.5 to the foreman and implementation, GLM 5.3 Flash to triage, and GPT-6 Sol to code review. Ben presented this as an arrangement to test against the team’s tasks. It illustrates using a cheaper model for research and a different model family to review the code. The model names will change; the decision is how to spend the team’s budget across the work the agents need to do.
Warp Factories defines the setup in code, including YAML configuration, so changes can go through version control and review. The dashboard presents that configuration and the runs it produces. The factory is something the team can change deliberately and test.
Give Agents the Environment Their Work Requires
A shared factory still needs somewhere to run its agents. Ben compared two approaches: a persistent server that receives the team’s work, perhaps with containers inside it, and separate environments started for individual agents.
Ben described a persistent-server pattern he had seen with OpenClaw and Hermes, sometimes with containers inside it. It avoids waiting for a machine to start and can run on hardware the team owns. Isolation between tasks needs attention, and one operating system may not suit every job.
Warp had a concrete reason to separate environments: its factory mostly ran on Linux, but work on an iPhone app needed macOS. The orchestration could stay on Linux while an implementation agent used a macOS runner.
Separate environments introduce deployment complexity and startup delays. The choice follows from the work: the operating systems, task isolation and response times the team needs. For a team of one or a few, Ben also suggested that a Mac mini tucked away in a cupboard could be enough.
Start with Background Tasks, Then Work with the Factory
Ben suggested introducing a factory gradually, starting with work that already has a trigger: investigating a CI failure, looking into an error alert, or checking the unassigned backlog each morning. Each automation connects a trigger to an agent run and an output, such as a pull request or a message.
This gives the team something to inspect while people continue using their existing coding tools. As confidence grows, the factory can take on more of the work people actively direct.
At Warp, Ben showed that interaction happening in shared Slack threads. Engineers can ask the agent to do work, discuss the result and see the same conversation. The factory also reports the cost of those interactions.
The iPhone example made this concrete. An agent built the app on a macOS runner, started a simulator and returned screenshots. Ben could inspect the interface and request changes without pulling the branch onto his laptop or starting Xcode himself. Once the visual result looked right, he could review the GitHub pull request.
The agent’s evidence becomes part of the review. A screenshot can answer a question about the interface; other checks are needed for behaviour it cannot show. The factory needs to return enough evidence for the person directing the work to judge the result.
Score Runs to Find Repeated Problems
Once work is flowing through the factory, the team needs to know how well it is running. Ben showed a dashboard that tracks productivity, spending and cost per pull request. It separates inference costs, the model calls billed in tokens, from compute costs, the machines running the agents’ environments.
Those totals tell the team where to look. They do not explain why a run took too long or needed several rounds of correction. Ben described how Warp investigates those problems with scoring agents: models that assess completed runs against written criteria.
Ben showed rubrics for efficiency and code quality. The code-quality configuration distinguishes work a reviewer would accept, work with insufficient evidence to judge, and work with a defect that needs fixing. Its instructions ask the scorer to evaluate the code diff against the repository’s standards and hold back when there is too little code to judge.
In Ben’s examples, efficiency scoring can expose problems even when the agent eventually finishes. Did it keep making failed tool calls? Did unclear testing instructions force it to work out how the tools were supposed to behave? How much human intervention did it need?
These are judgements against the team’s rubric. They give the team a way to find runs worth examining, alongside the code, tool calls and review history. The rubric itself needs care: if it rewards the wrong behaviour, its scores will point the factory in the wrong direction.
Turn Failures into Changes People Can Review
Warp’s self-improvement agent looks across scored runs for recurring problems and proposes changes. In response to an attendee’s question, Ben explained the timing: a scorer assesses completed work, while a separate scheduled job reads recent failures, their scores, notes and full conversations. He said once a day was probably the team’s current default for that improvement pass.
One example he opened was a draft pull request to improve branch-handling instructions: check whether a remote branch name already exists before pushing. Its description linked back to runs that prompted the proposal.
This was a proposed improvement. Ben did not show a measured saving from merging it. The example showed how the factory can turn a recurring issue into a reviewable change, with evidence a person can follow.
Ben also described a broader problem from his own setup: poorly written testing instructions and a troublesome test harness made agents waste time on failed calls and workarounds. Looking across conversations helped an improvement agent identify that pattern. An individual task could finish while the underlying problem kept adding cost to later tasks.
The changes here are to the factory’s instructions, skills and configuration. This is a feedback loop around the system running the agents. People still decide whether to accept the proposed changes.
Crediting Warp CEO Zach Lloyd’s phrasing, Ben called this factory engineering: improving the system the team relies on to build software. Product engineering continues to involve deciding what the product should do and how people should experience it. Factory engineering examines the cost, speed and quality of the process producing it.
Benchmark the Team’s Own Work
Ben’s final demonstration tackled a question that comes up whenever a new model arrives: would it improve this factory?
Warp builds benchmarks from its own completed tasks. An agent helps identify candidate conversations and estimate their difficulty, then a person narrows the selection. Each benchmark task records the repository state to start from, a prompt and success criteria. The factory can replay that work with a different configuration and score the result.
The chart Ben showed compared configurations on judge-rated correctness and average API cost. Each point represents a configuration tested on those tasks. He also switched the view to latency: the time needed to do the work. Those comparisons can inform a choice when a team cares about lower spending, faster responses or the strongest results on its selected tasks.
Ben was explicit that the results were specific to Warp’s setup and sample. A team building an iOS app may get different results from one doing front-end work or a specialised migration. The scores also depend on the success criteria and judge used in the benchmark.
Ben explained how to keep an experiment focused by changing only the implementation agent’s model. The other agents’ models stay fixed, so the team can compare the implementation model’s performance within the same system. Warp was also investigating whether different agent arrangements worked better: one agent doing everything, a foreman with a reviewer, or a larger assembly line. Ben did not present a settled winner among those arrangements.
Running these comparisons costs money. Ben noted that the team can reduce the number of candidate models, tasks or repeated runs to make the experiment manageable. The purpose is to get evidence for a decision the team needs to make.
Keep Discussion and Ownership Clear
In the Q&A, Eleanor asked how Ben handles work that can run in parallel alongside tasks that depend on one another. Ben tends to start with one implementation thread, then split emerging work into well-scoped issues. Those issues can feed an automation or become separate conversations.
That preserves an opportunity to discuss each piece. Ben was experimenting with ways for agents to recognise when a task needs further discussion before implementation. Delegating more work can also mean delegating more assumptions, so those questions matter.
Eleanor described planning the dependencies with an agent first, then using blocking relationships in Linear to tell the factory which tasks could start. Both approaches make the relationships between tasks explicit enough to direct the work.
Ownership stays with people too. Ben said Warp’s agents open pull requests under their own identity and include the source conversation and initiating person. For directly scoped work, that person reviews and merges; larger changes can involve another human reviewer.
Warp’s approach makes the factory’s operation available for the team to inspect and improve: what the agents are allowed to do, what evidence they return, which failures recur and whether a change improves the work. That gives the team a way to develop the factory alongside the product.
More information is available on the Warp Factories website, and the full session is available on Maven.
Build Your Agentic Software Factory on Maven
Hugo Bowne-Anderson and Eleanor Berger teach Build Your Agentic Software Factory on Maven, a practical course for builders who want to put these ideas into practice.
Over two weeks, participants set up a factory locally or in the cloud, give agents clear outcomes, context and constraints, and define the checks they need before trusting the agents’ work. They start with a conversational utility, extend the approach into a working web application, and learn how to publish or share what they have built.
The course is for builders with or without coding experience, working on ideas from their own work or life. Live workshops, office hours and project feedback help participants build a repeatable way to delegate, inspect results and keep improving their factory.
👉 Course details and enrolment on Maven.









