Evaluation-driven development in YAML files
Building a non-deterministic system is iterative tuning, and the knowledge it produces is fragile, it lives in your head and the chat scrollback, both...
|
May 23, 2026One agent building software has two problems. Neither is the model. It cannot specialize. The same context plans the work, writes the code, reviews its own output, and handles the commit. It does the middle two badly, because an agent is a poor critic of what it just produced. It also cannot remember. What it learned lives in one conversation. That conversation is gone the moment the session ends or the process crashes. Building with agents at any real scale means solving both. They are different problems. One needs coordination between specialized agents. The other needs a shared memory that grounds them.
Ghostbus is a small tmux-backed bus that gives each agent its own pane in one window, with a role. The setup I run has a builder that plans and writes, an auditor I call orange that reviews, and a cicd agent that owns the repo state, the commits, and the CI checks. Three processes, three contexts, side by side.
The first thing that buys is sight. You can watch which agent is working, which is waiting on input, and which command is still running. A single agent stuck in a loop is invisible until you go read its transcript. With three agents in three panes, the stuck one is obvious. Seeing them matters less than what they coordinate around.
The agents do not stay in sync by talking to each other. They stay in sync around a shared file. Before the work begins, the investigation is written down as a story, a small YAML document that lives in the repository:
The story the agents coordinate around
name: Builder copy architecture, a page reads as one voicestatus: activeproblem: > Section headlines drift in register across a page. The hero reads confident, the features tentative, the CTA generic.theory: > The copy step has no page-level voice pass. Derive one voice directive per page and thread it into every section prompt.refs: - packages/ai/src/shared/builder/stages/creative/copy/architecture.tsruns: - 2026-05-05T14:02 FAIL a6bc5142 # directive too weak, features still tentativeEvery agent reads that same file. It is the durable source of truth that no single agent's context owns, and it is what makes more than one agent possible at all. When three agents work one problem, they are not each holding a private copy of it in their own scrollback. They are all looking at the one story on disk. When one of them learns something, it writes it back where the others will see it.
Here is a single pass through the loop.
The builder picks up the story. It reads the problem, the current theory, and every run already recorded, starting from what has been tried rather than from a cold read of the code. It makes the change. Then it appends one line to the story, a run, with a timestamp, a verdict, the commit, and the signal that came out.
Then it asks for review, over the bus:
The review handoff, over the bus
ask_agent("orange", "Diff on builder-copy-architecture. Review it against the story and push back.")wait_for_reply(request_id, session_id="orange")The auditor is not a rubber stamp, which is why it is a separate agent. A large change carries a large context. A large context is where the ordinary mistakes hide, the ones obvious to a fresh reader and invisible to the agent that has lived inside the problem for an hour. This system is heavily structured and has design rules that matter. The fastest way to break one of them is to fix something else without noticing. So the auditor is positioned to disagree. It reads the same story, the theory the change was supposed to serve, and the diff. Its job is to argue the change, not wave it through.
That argument happens in the panes, over the bus:
The auditor pushes back
builder → orange One voice directive per page, threaded into every section prompt. The headlines cohere now. Diff is on the branch. Does it hold?
orange → builder It reads consistent, which was the problem. But you are running the directive over every section, including the ones bound to verbatim facts. A testimonial is a customer's words. A pricing row is a number. The design rule is that fact-bound sections do not get rewritten. You are breaking it to fix the register.
builder → orange The directive only touches headline register, not body facts. The quote itself is untouched.
orange → builder The quote, yes. The headline above it, no. A confident house-voice headline over a customer testimonial is you writing the customer's frame. Exempt the fact-bound sections from the voice pass and rerun the fixture. If the rest still coheres without them, it is a real fix. If it falls apart, the directive was never doing the work.
builder → orange Fair. Scoping it to unbound sections. Rerunning.Neither agent had the whole answer at the start. The builder had a change that looked right from inside the change. Orange had the design rule and a colder read. What resolved it was the exchange. A single agent would have committed the first version and moved on. The broken rule would have surfaced a week later, in a page that quietly rewrote a customer's words.
When they settle it, the builder appends the outcome as a run. The cicd agent, which watches the repo, commits the code and then appends the run in a second commit. The record lines up with the git history instead of being written up after the fact. It runs the tests. The builder has already opened the next story.
Agents crash and sessions close. By default that ends the work. A running agent's reasoning lives inside its session and nowhere else. When the process exits, it is gone. This loop does not have that failure. If the builder crashes in the middle of all that, nothing that mattered is lost, because none of it lived inside the builder. The story is on disk at its last committed state, runs and all. The next agent to pick it up opens it and resumes, grounded. The team's shared state survives any one of them.
The two tools do two different jobs. Ghostbus moves messages between agents while they run and keeps none of it once the panes close, which is the right design for coordination. The stories are the memory, files versioned with the code that outlive every session, which is the right design for grounding.
Neither is enough alone. A bus without the shared stories is a team that talks and forgets what it decided. The stories without a bus are a grounded record that nothing is actively working. Together they are a team of specialized agents that stays in sync and grounded, across crashes, across handoffs, across the end of any single session.
The bus, ghostbus | The stories, EDD | |
|---|---|---|
What it is | A live wire between running agents | Files versioned with the code |
How long it lasts | Nothing kept once the panes close | They outlive every session |
What it is for | Coordination | Grounding |
The review is not free. The builder stops and waits while the auditor reads. Running a big build through a second agent can double the engineering time, sometimes more. The cicd lane claws some of it back, staging commits and running tests and watching CI in parallel while the builder is already onto the next story. Local testing catches what would otherwise round-trip through CI. But that recovery is partial. Over a long build it does not come close to erasing the cost of the second pane.
I pay it anyway, because speed is not what this setup is for. Ghostbus and EDD are for long, big-context builds, the kind where one agent holds too much context to catch its own mistakes, and where no single human can keep all of it in view. On those, what has to be right is the output rather than the clock. Three sets of eyes carry the work, two agents and me. A solopreneur has no team standing behind the ship button, so the doubled hours are what trustworthy output costs. For anything I put my name on and ship alone, it is the trade I take every time.
Building with agents means a few that each do one thing, visible in their panes, coordinating over a bus, grounded in a record they all read and write. The leverage was never a bigger model. It is the wiring between small specialized agents, and the shared memory underneath them.
Remote, Pacific time, full-time or contract. Get in touch.