There is a big trend happening in the software industry about creating your own software factory. Inside all the fuzz of all non-realistic claims, the basis of what structure the argument of a software facture, for me, is what software development will look like in the following years.
At least for now, it won't be a fully autonomous factory of coordinated team of agents that will plan, implement and test the code end by end. Instead, software factory as the harness of the harness that will enable software developers to fully integrate background agents in their own company infrastucture to build code!
So we can define software factory as:
An open source environment allowing you to combine persistent coding agents, repository workspaces, issue intake, planning, implementation, and pull request review in a configurable SDLC with managed background enviroment.
I'm not the only one talking about this trend, there are some big players in the market that are surfing this wave.
- Ramp inspect
- OpenHand
- Mastra Factory
A software factory can be divided in two main building blocks. The first one is the methodology (and all the process that touches building reliable software) and the second one is the insfrastrucre around it (Llauching the sandboxes, giving a persistent filwsystem to the agent with durable states.....)
This blog is about the latter! As I was starting to tinker about how build it, I had some inspirations in the OpenHand Agent Server! Mostly because it was in python and second because it was very easy to undestand their codebase.
Architecture

What is the agent server? We are creating a simples envorment to run agents locally with it's own file system, sandbox and conversation management. A minimal API that let's your client manage states and sandboxes.
We can divide the project in 4 main parts. The first one is the Runtime Control Plane
Runtime Control Plane
The runtime control plane is responsible for managing all state related to building sandboxes. Its service does not involve dealing with agents or manipulating workspace files. Its business is creating and terminating Docker sandboxes.
It all starts when the client creates a remote workspace. After that, using the remote URL, it creates a brand-new sandbox with a server inside it, returning the session API key to the client.
Client SDK
The client SDK is an interface that other parts of the application can use, including the UI, backend, or a Python script. The interface begins by provisioning and starting the control plane and then creating a conversation that works as a remote client to communicate with the server inside the sandbox.
The client uses the exposed endpoints to control the events, conversations, and states inside the sandbox.
The Agent Server (my favorite part)
When the Docker container is up, it runs the gg.server FastAPI application. It is responsible for WebSocket and HTTP connections with the client, including authentication. The client can then control these main features:
- Conversation service
- LocalConversation
- The workspace and filesystem used by the agent
The server runs inside the Docker container, so the server and agent share the same environment.
The Agent
We are using the Pi “minimal” agent, which was very popular. However, the agent is written in TypeScript while our program is written in Python, so we needed an RPC adapter. The server sends Pi the user prompt through standard input as a JSONL RPC message. Pi returns structured JSONL responses through standard output.
A final overview looks like this:
- Client SDK → requests and uses a sandbox
- Control plane → creates and destroys the sandbox
- Agent server → hosts and records conversations in the sandbox
- Pi agent → performs one conversation run in the sandbox
Gaps and Limitations
As this is the beginning of a bigger implementation, it does not have persistent sessions, stronger guardrails, observability, or sufficient resources to control agent behavior across sessions.
Future Research
The next step for this project is to change the main language to TypeScript so we can build a more robust asynchronous application that is a better fit for most coding agents. Then, we can move toward provisioning virtual machines through companies such as Modal. Agents would then have their own repositories, computers, environments, and cohesive state to maintain long-running coding tasks in complex scenarios.
Sources
Why We Built Our Background Agent — Ramp ↗
Why Ramp Built Inspect — The Pragmatic Engineer ↗
How Ramp Built a Full-Context Background Coding Agent on Modal ↗