Assume the model call fails

Jesse James Richard
|
|
5 min read
#Method
#Architecture
#AI & Agents

Model infrastructure is unreliable, and not only at the edges. It is unreliable in the middle of an ordinary day. Calls time out. The service is overloaded and hands back a rate limit. A response arrives empty, or truncated, or well-formed and wrong. That is the ordinary condition of building on inference you do not run yourself, and it holds no matter which provider or model you point at. A pipeline that assumes every model call returns a clean, complete answer rests on an assumption that fails constantly.

And the failures do not stay where they happen.

A bad response does not stay put

An agentic pipeline is a chain of model calls, one classifying, the next composing, the next checking, each feeding the one after it. When a call in the middle returns something broken, the break travels downstream. A truncated section becomes a page with a hole in it. An empty payload becomes a default nobody chose. The dangerous case is the response that is broken but valid, a body that parses cleanly and says the wrong thing, because nothing rejects it and it rides several more stages before it shows up as damage that is hard to trace to its origin.

So a pipeline built on model calls needs three things wrapped around every one of them.

Retry absorbs the transient

Most model failures are transient. An overload passes. A rate limit lifts. A timeout turns out to have been a fluke. Retry with exponential backoff clears them, running the call again after a growing delay until an attempt lands. This is the first layer. Without it the pipeline dies the first minute the provider has a bad one.

But retry only fires when a call fails loudly. It trusts any response that signals success. Some responses signal success and are broken.

The audit rejects what retry trusts

The second layer is an audit. It is a gate that inspects each response and decides whether it is usable, complete, parseable, inside the contract the next stage expects, before anything downstream is allowed to trust it. A response that fails the gate is rejected rather than passed along and hoped over. The call runs again. Retry recovers from the failures the provider reports. The audit recovers from the failures the provider does not.

Trace so nothing dies in silence

Retry and audit keep the pipeline running. Neither of them tells you why a call failed. For that, every model call has to leave a trace, a record of what it did and how it ended, written whether the call succeeded or not, so that no failure disappears without evidence and any of them can be run to ground later. A pipeline that cannot explain its own failures can only recover from the ones it already understands.

Every call now emits one line:

The trace line

[LLM] Result: task=classify_subtype finish=MAX_TOKENS      input=2311 output=147 thinking=3820 blocked=none

Anomalous at WARN, ordinary at INFO. It records the finish reason and the token accounting on every call, including the ones that worked. Error-path logging is a reflex every engineer has. Success-path logging feels like noise until it is the only evidence that matters. For model work, the worst failures come back as a 200.

That trace surfaced every kind of failure the pipeline could produce, overloads and empty bodies and safety blocks and malformed payloads the parser had been repairing on its own. The quietest one justified the whole mechanism. A response would come back as a 200, parse cleanly, and be truncated, because the upgraded models spend part of the output budget on internal reasoning and a task tuned for the old budget ran out before it finished the answer. Nothing in the response looked wrong. The trace showed finish=MAX_TOKENS on a hundred-and-fifty-token answer against a four-thousand-token budget. The diagnosis was that one field.

0

input tokens

0

output tokens

0

thinking tokens

Retry, audit, trace

Without the trace it stays a vague complaint about an inconsistent model. With it, the fix took an afternoon. Each task's budget gained a reasoning level, low for classification and high for composition. Nothing ran out on thinking after that.

Together the three are one durable task. Trace makes sure whatever slips past the first two leaves a record instead of a mystery. A model call wrapped in all three either produces a result good enough to use or fails in a way you can see and explain. That is the unit an agentic pipeline has to be built from, because a pipeline is dozens of model calls deep and every one of them sits on the same unreliable infrastructure.

What you are building on

The model upgrade introduced the condition, and a model that changes under you is the standing arrangement of building on someone else's inference. The vendor will change behavior. The service will have a bad minute. A response will come back looking like success and be broken. A pipeline that handles only the failures announcing themselves as failures can survive only the failures it already knows.

So I assume the call fails. Every one of them retries the transient part with backoff, gets audited before anything trusts it, and leaves a trace so the failures that get through still leave something to find. It is more machinery around every call than I wanted to write. On inference I do not run, it is the only arrangement that has held.

Turning uploaded files into retrievable context

#Method
#Data
#AI & Agents

A large language model is brilliant and knows nothing about your business. Files are how the platform closes that gap. Every document, image, audio an...

Jesse James Richard

|

May 19, 2026
Read previous

Hiring an early engineer, or building something like this

Remote, Pacific time, full-time or contract. Get in touch.

Contact Jesse
Home
About
Contact
Sitemap
Privacy Policy
Terms of Service
Cookie Policy
Assume the model call fails | Jesse James Richard