The Swiss Cheese Model

Jesse James Richard
|
|
11 min read
#Method
#AI & Agents
#Testing
#CI/CD
#Keystone

An agent writes most of the code in Giant Context, and it deploys to production where real customers depend on it. That only works if I can trust code I did not type, and no single check is enough to earn it. The trust comes from eleven checks stacked one behind another.

The frame is James Reason's Swiss cheese model, drawn up for accident analysis in aviation and medicine. Every defence a system has contains holes, and the holes move, because the system keeps changing. A hazard gets through when a hole in the first defence lines up with a hole in the second, and a hole in every one after it. Nothing in the model asks a defence to be perfect. It asks how many have to be crossed at once, because each one multiplies against the others and the probability of a failure crossing all of them falls to near zero.

Slices of Swiss cheese stacked front to back, each with holes in different places, so no straight path runs through all of them.
A failure has to pass through a hole in every one, all lined up at once.

The slices that constrain what gets written

The first three layers do not catch mistakes. They make whole classes of mistake difficult or impossible to write in the first place.

Slice one, strict design patterns. The codebase has firm, non-negotiable rules about how code is shaped, from naming to file layout to where a function is allowed to live. An agent generating code inside those rules cannot scatter logic into a component or invent a novel structure, because the structure is fixed and enforced. The hole is that structure is not logic. Code can sit in exactly the right place and do exactly the wrong thing.

Slice two, typed schemas that drive the features. The platform's critical surfaces are defined as typed schemas, and those schemas are the source everything else derives from. A field's type, its bounds and its validation live in one declaration, and code that violates it does not compile or does not pass validation. The hole is that a schema constrains the shape of data and not the logic acting on it. What it does remove is a whole category of error, the mismatched field, the wrong type, the missing required value.

Slice three, the critical machinery is generated. The data hooks, the API client, the query cache keys, the form validation and the route tables all come out of the generation step. The generating system decides what is in those files. An agent can edit one, and the next pass overwrites the edit, so the only durable way to change generated code is to change the schema it came from. The hole is that the generator can be wrong, and then it is wrong in every file it produced. But the generator is small, it is reviewed once, and it does not change per feature, so its holes are few and they stay still.

The slices that check what was written

The next two layers assume something did get written and try to catch what is wrong or missing.

Slice four, tests, forced and generated. No API route can exist until it has a passing test. The generation step supplies the scaffolding, the mocks and the factories, so what the agent writes is the assertion rather than the apparatus around it. The hole is that the assertion is still the agent's, and an assertion written from the same misunderstanding as the code will agree with it.

Slice five, checks for what a unit test cannot see. Some obligations are not attached to any one route. Every string a customer sees has to exist in three languages, and a feature can be complete and correct with its French and Spanish entries missing, because translating strings was a rule about the task rather than the task. No route test fails, because no route is wrong. A separate completeness check reads the locale files and goes red.

The completeness check going red

translation check: FAILED  missing (fr): apps.website.settings.newField  missing (es): apps.website.settings.newField

Slices one to four all examine the code that was written. This one asks whether the work is finished.

The slices that gate the release

Now the code is written, generated, and checked locally. Two layers stand between it and production.

Slice six, a pipeline that gates in more than one place. Continuous integration runs the full suite and the checks on every push to the shared branch, and runs them again at deploy.

The pipeline's job list

validate-secrets:detect-changes:test-ts-lint:test-ts-unit:test-ts-integration-api:test-ts-rbac:# …the full suite, twice through the pipeline

The hole is that the pipeline runs only the checks that exist, so every gap in slices four and five is a gap here too. It adds no new judgement. It adds repetition, which is worth having, because a bug now has to survive the same check twice and survive it in a second environment.

Slice seven, deploy itself. The build and deploy step occasionally fails, and when it does it is almost always a dependency or a lockfile rather than a logic error that reached this far.

The slices that catch what got through

No honest model stops at the release, because the premise is that something can get through. The last four layers watch what happens once the code is live.

Slice eight, the check at deploy. The moment a service deploys, a health job polls it until the exact version that just shipped is the version answering, per service across the fleet. Two failures die here. A revision that cannot come up healthy never takes traffic, so a bad deploy leaves production on the last good version. And a deploy that reports success while the old revision is still serving fails the version match. The hole is that the check asks the service whether it is up. It does not ask it for a page.

Slice nine, the continuous check. The deploy check runs once. This one runs on a short interval for as long as the service exists, recording whether it answered and how quickly. It catches what was fine at deploy and broke later, a dependency that went down, a service that slowed under load, a certificate that lapsed. The public status page is where that history is published. The hole is slice eight's hole on a timer. It polls the same health endpoint, so it knows the service is up and knows nothing about what the service returns to a reader.

Slice ten, production errors go to Brain. Anything that throws in production is captured and sent to Brain, the platform's own error-triage service, which investigates it and returns a root cause and a fix for me to confirm. Some failures reach production. Brain is what turns one of them into a fixed bug rather than a standing one.

Slice eleven, the fix arrives where I already work. Brain runs an MCP server, so the error, the diagnosis and the proposed fix land on my command line inside the tool I am already in. An error in production is only cheap if the distance between finding it and shipping the fix is short.

A vertical list of system health status indicators on a dark background.
Slice eight, at deploy. Each service confirmed on the version that just shipped.
A digital dashboard showing system status bars for API, Files, AI, Website, Router, and Brain services over a 90-day period.
Slice nine, on the interval after. The public status page and uptime history.
A screenshot of a software issue tracking interface showing a recommended fix and a code diff.
Slice ten. A production error captured by Brain, with root cause and a suggested fix.
A screenshot of a software interface showing a list of investigation steps with code-like commands.
Slice eleven. The full error manifest on the command line, where the fix gets applied.

Why the stack works when no slice does

None of these layers is a guarantee. Design patterns permit wrong logic in the right place. Schemas constrain shape and not logic. A generator's bug reaches every file it wrote. A forced test is only as good as the assertion inside it. Completeness checks cover the obligations someone thought to write a check for. The pipeline runs the checks that exist. Deploy catches dependency failures. The health checks, at deploy and on the interval after, ask whether a service is up and never ask it for a page. Brain files what throws. I can state the hole in every one of them, which is why I am willing to rely on the stack at all.

The stack holds because a failure has to find a hole in all eleven at the same time. It has to be structurally legal, schema-valid, untouched by the generator, agreed with by its own assertion, complete in every language, passed through the pipeline twice and in two environments, deployed onto a healthy revision, invisible to a check that never asks for a page, and either not throwing at all or throwing so rarely that no pattern forms. That is possible. It is not likely, and each layer multiplies against the one before it, which is what makes it reasonable to hand the typing to a machine and still ship to customers.

This is defence in depth, and it is the ordinary way to build anything that has to survive its own faults. No layer is asked to be complete. Each is asked to fail differently from the ones on either side of it, so what one lets through the next one stops. Building defensively costs more up front than trusting a single check, and it is what makes the difference between a system that is correct and a system that stays correct.

My errors come with fixes attached

#Method
#AI & Agents
#APIs & MCP

Production throws an error and the hour before the fix is the expensive part, the reproduce, the search, the reconstruction. Brain files each one as a...

Jesse James Richard

|

Mar 30, 2026
Read previous

Hiring an early engineer, or building something like this

Remote, Pacific time, full-time or contract. Get in touch.

Contact Jesse
Home
About
Contact
Sitemap
Privacy Policy
Terms of Service
Cookie Policy
The Swiss Cheese Model | Jesse James Richard