devil.services Field notes

Tags /

AI-assisted dev

4 posts

· AI-assisted dev · 15 min

I benchmarked four ways to give an agent a browser

Playwright, chrome-devtools-mcp, ego lite and the Claude Chrome extension on the same pages. Measured latency, tokens, fidelity and price. All three scriptable harnesses passed all five tasks; the worst one costs $308 more per thousand runs for the same answers.

  • ai-dev
  • browser-automation
  • playwright
  • mcp

· AI-assisted dev · 13 min

Regex can't count pages

My ebook pipeline's top-ranked pick for Atomic Habits was a 115-page abridged edition with perfect metadata. So I put an LLM judge in the pipeline, benchmarked eleven models on ten fixed candidates with five runs each, and watched the only real data loss of the week come from deterministic code instead.

  • ai-dev
  • evals
  • llm
  • biblioteka

· Fractional CTO · 4 min

I hardened my pipeline against AI hallucinations. Then a CLI hallucinated.

The week I added an LLM safety gate to a file pipeline, the only data loss came from deterministic code: a CLI that silently refused an operation, a fallback that guessed an identity, and a verifier that faithfully verified the wrong object. A short postmortem, and the argument for auditing your boring layers.

  • postmortem
  • reliability
  • architecture
  • ai-dev

· AI-assisted dev · 11 min

A regex skipped a fifth of my films for being foreign. A $0 model read every one.

filmoteka's parser silently dropped 22% of real releases, every one multilingual. I benchmarked four GLM tiers against it on 115 real titles, five runs each. Every tier recovered all of them; only the paid ones did it without ever grabbing the wrong film.

  • ai-dev
  • evals
  • llm
  • glm