Fractional CTO · · 4 min read
I hardened my pipeline against AI hallucinations. Then a CLI hallucinated.
The week I added an LLM safety gate to a file pipeline, the only data loss came from deterministic code: a CLI that silently refused an operation, a fallback that guessed an identity, and a verifier that faithfully verified the wrong object. A short postmortem, and the argument for auditing your boring layers.
Everyone hardening a pipeline with AI in it worries about the same thing: the model making something up. I spent a week on exactly that - a judge that vets inputs before they enter my personal library pipeline, a verifier that checks every file before it ships anywhere. Belt, braces.
Then the pipeline deleted the only copy of a file, and the model had nothing to do with it.
The four-line disaster
The flow was a replacement: swap a mediocre copy of a document for a better one. Download the new file, ingest it into the library, verify it, and only after a verified success remove the old copy. That last clause is the safety gate. I wrote it carefully.
# the replacement flow, as it actually ran
download -> new file ok, 89,557 bytes
ingested -> library id 16 <- the SAME id as the copy being replaced?!
verify -> ok <- verified the OLD file
removed old id 16 <- which was the only copy
$ library list
[]
Three deterministic components each did something defensible, and the composition destroyed data:
The CLI refused silently. The library tool (calibredb, Calibre’s command line) declines to add a document whose title and author already exist. No error, no exit code worth the name, just a polite sentence on stdout and no id. From the caller’s side, “I added nothing” and “I added something but you failed to parse the id” look identical.
The fallback guessed. My wrapper’s recovery path for “couldn’t parse the new id” was to take the max id in the library. Reasonable on the happy path it was written for. On the refusal path, it attributed the existing document to the new ingest. The function returned a number, the number was real, and it pointed at the wrong object.
The safety gate anchored to the lie. The verifier looked up id 16, found a perfectly good file, and passed it. Every check it ran was correct. It was checking the thing that was about to be deleted, on behalf of the thing that never got added. Then the “remove the old copy after a verified success” step did exactly what it was told.
The fix is boring. The lesson is not.
The repair took minutes: pass the flag that permits duplicates, throw instead of guessing when no id comes back, and refuse removal when the “new” id equals the id being replaced. Belt, braces, and a third strap that checks whether the belt is lying.
The lesson is the general shape, because this failure is everywhere once you see it:
- Silent refusal reported as success. Any tool that declines an operation on stdout instead of an exit code turns every caller into a parser, and every parser has a fallback.
- Fallbacks that guess identity. “Take the max id,” “use the latest snapshot,” “assume the previous deploy” - each one is fine until the day the primary path fails for a different reason than the fallback was written for.
- Safety gates anchored to unverified identity. A verification step is only as good as its answer to “verify what, exactly?” Mine verified a file thoroughly and never questioned whether it was the right file.
Swap the nouns and this is a deploy pipeline promoting the wrong artifact because the build id was reused, a backup rotation pruning the only good snapshot because “latest” resolved wrong, a queue consumer acking a message it never processed. The AI in your stack did not invent this class of bug. It just got you to finally look for it - in the one place you were already suspicious.
What an architecture review is actually for
When I review a system, hallucinating models are on the checklist, but they are the easy part - everyone is already scared of them, so they are already fenced. The findings that matter are almost always in the seams between deterministic components: the return value nobody validates, the fallback nobody re-reads after the happy path changes, the gate that trusts an identity it never checked. Those seams do not announce themselves, because every component in isolation is behaving as documented.
Audit your boring layers like you audit your model outputs. The model at least has the decency to be famous for lying.
This is build-in-public from devil.services. The longer story of the pipeline this happened in - the judge, the verifier, and an eleven-model benchmark - is in Regex can’t count pages.