devil.services Field notes

AI-assisted dev · · 13 min read

Regex can't count pages

My ebook pipeline's top-ranked pick for Atomic Habits was a 115-page abridged edition with perfect metadata. So I put an LLM judge in the pipeline, benchmarked eleven models on ten fixed candidates with five runs each, and watched the only real data loss of the week come from deterministic code instead.

My ebook pipeline ranks download candidates with sensible deterministic rules: EPUB beats AZW3 beats PDF, preferred language, exact-ISBN bonus. For Atomic Habits those rules put a 5 MB English EPUB at the top of the list - correct title, correct author, correct publisher. Its metadata also said 115 pages, for a book whose real editions run around 300.

Candidate no. 1, ranked best by the picker index

Atomic Habits: An Easy & Proven Way to Build Good Habits & Break Bad Ones

Clear, James

publisher
Penguin, 2019
format
EPUB, 5 MB, English
pages
115
expected
~306
Abridged judge, 0.90 conf

The catalog card my pipeline liked best. Right title, right author, real publisher, plausible size. It is not the book - it’s a third of the book.

No regex catches that. A page count is only wrong relative to what the book should be, and “should” lives in world knowledge, not in the row. Same for “Shortcut Edition” knockoffs whose phrasing mutates faster than my blocklist, for workbook companions with the real author’s name in the title, and for a Tamil translation of the right book. The row is honest. The interpretation needs a reader.

This is the same wall filmoteka hit, when Radarr’s release parser silently dropped a fifth of real releases, every miss a foreign or alternate title, and a small GLM model recovered all of them. The fix there became a pattern I now trust: the LLM outputs judgments; policy stays in code. The model never downloads, deletes, or delivers anything. It fills in a verdict column, and deterministic code decides what a verdict is worth.

Two gates in an old pipeline

The pipeline itself is plain: resolve a request against OpenLibrary, search shadow-library indexes, download the best edition, normalize it to EPUB in Calibre, deliver it to my iPad and Kindle. It got two new stages this week.

Click a step. The two the model sees are the only places one is consulted; every arrow between them is deterministic code.
A fuzzy request ("that pitching book by Enns") becomes canonical metadata: title, author, ISBN. OpenLibrary and Google Books, no model.

The judge sees the whole candidate list in one call - title, authors, publisher, year, format, size, page count, language - and returns strict JSON: real, summary, abridged, wrong-book, wrong-language, or unsure, plus a scan-risk guess and a confidence. Code then rejects confident non-real verdicts and sorts the rest. Cost: one call per search.

The verifier runs after ingest, before anything leaves for a device. Deterministic checks first: word count against a floor, mojibake ratio, and a scan for conversion artifacts. Then the model reads three text samples - start, middle, end - and answers: right book, complete, readable?

The artifact scan earned its keep on day one, in places body-text extraction can’t even see. A PDF-to-EPUB conversion had baked my local download path into the per-chapter <title> tags - sixteen times, invisible in the extracted text, visible in a reader’s page header. And a “clean” EPUB carried a shadow-library stamp baked into its OPF description metadata, invisible in any reading view. So the verifier greps the raw zip entries too, not just the prose.

The layer that actually ate a book

Here’s the part I didn’t expect to write. While swapping a PDF-conversion copy for a nicer native EPUB, the pipeline deleted the only copy of the book. Not the model’s fault - the model was right at every step. The culprit was calibredb.

# the replacement flow, before the fix
  download -> epub  87 kB   ok, 89,557 bytes
  ingested -> Calibre id 16     <- same id as the book being replaced?!
  verify -> ok (23,243 words)   <- verified the OLD file
  removed old Calibre id 16     <- which was the only copy

$ calibredb list -s "Pitching"
  []

calibredb add silently refuses to add a book whose title and author already exist - no error, no id, just a polite message on stdout. My wrapper’s fallback for “couldn’t parse the new id” was to take the max id in the library. That attributed the existing book to the new ingest, the verifier dutifully verified that existing (fine) file, and the “remove the old copy after a verified success” step removed everything. The safety gate worked; the layer under it lied about identity.

The fix is boring and absolute: pass --duplicates, throw instead of guessing when no id comes back, and refuse removal when the “new” id equals the id being replaced. But the lesson is the good kind of embarrassing: I spent the week hardening against hallucinating models, and the near-disaster came from deterministic plumbing asserting something false with total confidence. Audit your boring layers like you audit your model outputs. This incident got a standalone postmortem, because the pattern deserves more air than a section in an ebook story.

Eleven models, ten catalog cards, five runs each

Which model should sit in the judge seat? Cost turns out to be a non-question: a judgment is ~2,000 tokens, so even the priciest candidate here judges a book for well under a cent. What matters is accuracy and latency - the judge runs inline in search, so every extra second is felt.

I froze the ten-candidate list into a fixed benchmark - real edition, abridged edition, workbook, summary-mill product, three translations, a probable scan, a wrong-book decoy, and a junk-titled-but-real row - and ran the production judge across eleven models, five repeats each, temperature 0 where allowed, every run receipted to JSONL.

model runs real EPUBabridged 115pworkbookIndonesianInstareadTamilscan PDFSpanishdecoyjunk title
haiku-4.5 5x ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
ox-alpha 5x ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
glm-5.3 11x ✓ 10/11 ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
minimax-m3 5x ✓ 4/5 ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
deepseek-v4-flash 5x ✓ 3/5 ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
deepseek-v4-pro 5x ✓ 1/5 ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
gemini-3.7-flash 5x ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
glm-5 10x ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
glm-4.6 10x ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
kimi-k3 5x ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
qwen3-max 4x ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ 2/4 2/4
Share of runs each model judged each candidate correctly. Hover a cell for the verdicts it actually gave. One column carries nearly all the disagreement: the abridged edition.

Nine of the ten cases are a solved problem: workbooks, summary-mill titles, translations, and the decoy get caught by everything from a $0.09 model to a $3 one. The whole benchmark collapses into a single burning column: who can tell that 115 pages is not the book?

0% 50% 100% 5s 10s 20s 40s 60s median judge latency (log scale) haiku-4.5: abridged catch 5/5, median 6.5s, $1.00 / 5.00 per M tokens, 5 runs haiku-4.5 ox-alpha: abridged catch 5/5, median 29.6s, $0 / 0 per M tokens, 5 runs ox-alpha glm-5.3: abridged catch 10/11, median 62.9s, $1.40 / 4.40 per M tokens, 11 runs glm-5.3 minimax-m3: abridged catch 4/5, median 9.2s, $0.30 / 1.20 per M tokens, 5 runs minimax-m3 deepseek-v4-flash: abridged catch 3/5, median 27s, $0.09 / 0.17 per M tokens, 5 runs deepseek-v4-flash deepseek-v4-pro: abridged catch 1/5, median 35.2s, $0.58 / 1.16 per M tokens, 5 runs deepseek-v4-pro gemini-3.7-flash: abridged catch 0/5, median 6.9s, $0.38 / 1.88 per M tokens, 5 runs gemini-3.7-flash glm-5: abridged catch 0/10, median 7.2s, $0.60 / 1.92 per M tokens, 10 runs glm-5 glm-4.6: abridged catch 0/10, median 10.2s, $0.60 / 2.20 per M tokens, 10 runs glm-4.6 kimi-k3: abridged catch 0/5, median 25.2s, $3.00 / 15.00 per M tokens, 5 runs kimi-k3 qwen3-max: abridged catch 0/4, median 11.3s, $0.78 / 3.90 per M tokens, 4 runs qwen3-max
Abridged-edition catch rate against median judge latency. Top-left is where you want to live. One model lives there.

Findings, in the order they surprised me:

Claude Haiku 4.5 won. 10/10 on every run, including the abridged catch five times out of five - and it was also the fastest thing I tested, at a 6.5 s median. I expected to trade accuracy against latency; the smallest Anthropic model declined the trade.

A free mystery model tied it on accuracy. OpenRouter’s cloaked stealth/ox-alpha also went 10/10 across the board at ~30 s. Whoever is behind the curtain, their model can count pages. It’s a preview that can vanish or change identity any week, so it’s a data point, not a default - and on a cloaked free model you should assume prompts are logged.

Reasoning helps exactly once. GLM-5.3, which won’t run with thinking off, caught the abridged edition in 10 of 11 runs - at a minute per judgment and behind a rate limit that made repeated runs miserable. Its cheaper, faster siblings (GLM-5, GLM-4.6) plateaued at a flat, perfectly consistent 9/10: twenty runs, twenty times “real” for the abridged edition.

Price bought nothing here. The most expensive model in the bench, at $3/$15 per million tokens, matched the sub-dollar models: flat 9/10, abridged edition missed every time. Meanwhile MiniMax M3 at $0.30/$1.20 caught it four times out of five at 9 s, and DeepSeek v4-flash - nine cents per million input tokens - caught it three of five. The correlation between price and judgment on this task is roughly zero.

Only one model fell for the decoy. Qwen3-max, alone in the field, judged “Atomic Habits of Highly Effective Teens: A Teen’s Adaptation Inspired by James Clear” to be the real book in half its runs. Every other model - including the free ones - saw through it.

Five providers, one JSON judgment

The dirty secret of the benchmark is that getting the same request to run everywhere was harder than writing the judge. All I need is OpenAI-compatible chat completions, JSON mode, temperature 0. In practice:

One provider’s newest models reject thinking: disabled outright - reasoning is mandatory, and you ask for less of it with effort: "low". Another pins temperature to exactly 0.6 with thinking off and exactly 1 with it on, and 400s anything else. OpenRouter returned 402s that looked like an auth problem but were actually a pre-authorization problem: send no max_tokens and it reserves credit for the model’s full 65K worst case. And one provider enforces a quota on its reasoning tier that no backoff inside a single evening could outlast.

The client that survived is a small ladder: try JSON-mode + thinking-off; on a 4xx, retry with JSON-mode + low-effort thinking; then bare. Strip temperature the first time an endpoint complains about it, remember what worked per model, back off on 429s, and always send an honest max_tokens. None of this is clever. All of it was necessary.

What actually shipped

<1¢
judge + verify cost per book, on any model tested
~55¢
for the whole benchmark: 66 receipted runs across eleven models
0 books
lost to the LLM. The one deletion was deterministic code lying about an id.

A judgment is ~2,000 tokens. At personal-library volume, model price is noise; latency and accuracy are the whole decision.

The pipeline currently judges with GLM-5 - fastest of the 9/10 plateau, on an account with actual credit in it - and the moment the OpenRouter balance is funded, the judge seat goes to Haiku 4.5: one env var, no code. The verifier holds delivery on failure either way, and everything the gate does is logged to JSONL, so the next post about it can have receipts instead of anecdotes.

What this benchmark is not

Small and deliberately hard: ten hand-built candidates around one book, so read the direction, not the third decimal. Five repeats per model (eleven for GLM-5.3 across two rounds, four for Qwen3-max after provider errors), temperature 0 where the provider allows it, the production judge prompt unmodified. Latency is the median of full judge calls, ten candidates per call. Prices are per million tokens, input/output, as listed on 2026-08-25. The abridged case is one candidate, not a distribution of abridgments; a different fake could reshuffle the middle of the field. The receipts (per-run verdicts, JSONL) are what I trust, and they’re small enough to rerun in an afternoon.

The pattern holds up for the second domain in a row. Give the model the one column deterministic code can’t fill - is this actually the thing? - and keep every consequence in code you can read. Regex still can’t count pages. Now something in the pipeline can.

This is build-in-public from devil.services, one field note from the biblioteka build. The next one has thirty days of quality-gate receipts from real acquisitions, including the knockoffs it turned away.