AI-assisted dev Ā· Ā· 11 min read
A regex skipped a fifth of my films for being foreign. A $0 model read every one.
filmoteka's parser silently dropped 22% of real releases, every one multilingual. I benchmarked four GLM tiers against it on 115 real titles, five runs each. Every tier recovered all of them; only the paid ones did it without ever grabbing the wrong film.
filmoteka, my media-automation MCP, kept telling me films were unavailable when they were sitting right there on the indexers. The parser inside Radarr reads each release title with a regular expression, and a regex cannot tell that ДолŃŃŠøŃ is Solaris. So it binned the release as Unknown Movie and reported nothing found. On 115 real release titles it missed 22%. Every single miss was a foreign or alternate title. Then I put four GLM model tiers on the same job. Every one of them recovered all 22%, on every run. The free one did too, and it also grabbed the wrong film on four runs out of five. The tier that costs about a dollar per thousand searches never wobbled.
- 22%
- of real releases the regex silently dropped
- 100%
- recovered by every GLM tier, on all five runs
- $1.03
- per 1,000 searches for the cheapest tier that never wobbles
Ten films, 115 real release titles off the live indexers, five runs per model. Not a synthetic set.
Regex reads names. It cannot read meaning.
Radarr and Sonarr decide whether a release is the film you asked for by parsing its title. For a clean scene release, Heat.1995.1080p.BluRay.x264, that works perfectly. It falls over the moment a title carries an alternate title, a transliteration, or a different alphabet. And half of what my indexers return does exactly that.
Here is what the parser walked away from. Every one of these is genuinely the film I wanted:
The Hyena La iena (1997) SD h264 Ita Eng- the film is titled La Iena at homeSolaris Solyaris (1972) AC3 5.1 ITA 1.0 RUS 1080p- transliterated, not translatedBrother / ŠŃŠ°Ń / Brat (1997 Russian Audio, MULTi15 Subtitles)- three alphabets, one filmHeat La sfida (1995) 1080p H265 BluRay Rip ita eng- the Italian release nameThe Matrix 1 4 Pack 1999 2021 REMASTERED- a pack that contains the film, parsed as a mess
A person reads all five in a second. The regex reads Solyaris, La sfida, ŠŃŠ°Ń and shrugs. It is not broken. It is doing exactly what a regex does, which is match patterns, not understand them.
How I tested it, so the numbers mean something
I picked ten films spanning the difficulty range: Soviet cinema, alt-titled European thrillers, plain English controls, and a couple of deliberate traps. For each, I pulled the real candidate releases straight off the live indexers. That is 115 real titles, 106 of them genuine matches, plus 9 hard negatives the indexers themselves handed me: a German film mistaken for Brother, a discography, a soundtrack, a TV series.
Each system judges every candidate match or no-match. The two mistakes are not equal. A miss is an annoyance: a film I do not get, and go looking for by hand. A wrong grab is worse: filmoteka downloads the wrong film and files it under the right name, and I do not find out until I press play. So I care about precision more than recall, even though recall is where the regex bleeds.
The baseline is Radarrās own parser, its real /parse endpoint plus a title-and-year match, not a strawman I built. Against it, four GLM tiers from z.ai, each judging a filmās whole candidate list in one call, temperature 0, reasoning off. And I ran every model five times, which turned out to matter more than anything else here.
The prompt
No few-shot examples, no chain-of-thought, one paragraph. This is the whole instruction:
You match torrent release titles to a target movie for a media library. Titles
are often multilingual, with alternate/transliterated titles, director names,
edition tags, audio-track lists and other clutter. Mark match=true ONLY if the
release is the SAME FILM as the target (same work, matching year and identity) -
a pack/collection that CONTAINS the target counts as a match. Reject different
films that merely share a word, soundtracks, TV series, and other movies. Return
STRICT JSON: {"results":[{"index":int,"match":bool,"confidence":number}]} with
one entry per candidate.
Each film is one call. The user message is the target plus its candidate list:
TARGET: {"title":"The Hyena","year":1997,"original_title":"La Iena","language":"Italian"}
CANDIDATES:
0. The Hyena La iena (1997) SD h264 Ita Eng-MIRCrew
1. La Iena - The Hyena 1997 DVD5 ITA
2. La iena / The hyena [1997, DVDRip] AVO
Settings: temperature: 0, thinking: disabled, response_format: json_object. The plainness is deliberate, and it is also the honest caveat: this is the first prompt I wrote, not a tuned one. A better prompt would probably steady the small models. I have not written it, because the point already stands without it.
The results
| System | Recall | Wrong grabs (5 runs) | Perfect runs | Cost / 1k |
|---|---|---|---|---|
| Radarr regex | 78% | 0 | deterministic | $0 |
| GLM-4.5-flash (free) | 100% | 0 to 2 | 1 / 5 | $0 |
| GLM-4.5-air | 100% | 0 to 2 | 2 / 5 | $0.45 |
| GLM-4.6 | 100% | 0 | 5 / 5 | $1.03 |
| GLM-5.2 | 100% | 0 | 5 / 5 | $1.55 |
Recall is solved everywhere: every tier recovers the whole multilingual fifth. Consistency, the share of five runs with zero wrong grabs, is where the cheap tiers wobble. The regex is perfectly consistent but blind; the free model sees everything and stays clean one run in five.
Two things fall straight out of that table. Recall is solved, completely and stably: every GLM tier recovered all 22% the regex dropped, on all five runs, with zero variance. The recovery is not luck. The wobble is entirely in the wrong grabs, and only in the two cheap models.
The regex, for its part, never grabs the wrong film and never will, because it only matches what it can already read. Perfect precision, blind to a fifth of the catalogue. That is the trade the whole post is about.
The free model is a coin flip on the hard case
Both small tiers, flash and air, kept tripping on one release:
Geschwister - Kardesler (1997) 1080p WEBRip x264
matched to Russian Brother. Geschwister is German for āsiblingsā; KardeÅler is Turkish for ābrothers.ā The model saw the meaning of brotherhood, saw the year 1997, and called it a match. It is a different film. That is the failure mode you buy when you swap pattern-matching for meaning: the model understands the title, so it can be wrong about what the title means.
Here is the part that nearly got me. The first time I ran the benchmark, flash scored a clean 100%, zero wrong grabs, and I started writing this up as āthe free model is flawless.ā Then I ran it again. Two wrong grabs. Same prompt, same temperature 0, same 115 titles. So I ran every model five times instead of once.
flash grabbed the wrong film on four runs out of five. air on three. Temperature 0 does not mean deterministic, it just narrows the distribution, and on the one genuinely hard decoy the small models land on both sides of the line. GLM-4.6 and 5.2 rejected the German film on all five runs, no exceptions. That is the real gap between free and a dollar a thousand: not accuracy on a good day, but whether you can trust a single run at all. I almost published the good day.
Which one Iād actually run
GLM-4.6. It is the cheapest tier that is stable: perfect recall and perfect precision on all five runs, it read the decoy the free models kept missing, and it costs about a dollar per thousand searches with reasoning off. At the volume a home library does, that rounds to pennies a month, and identical titles cache after the first look.
GLM-5.2 is just as stable and it was the fastest at 3.3 seconds a call. But it costs 50% more than 4.6 for no accuracy gain here. Its million-token context and deeper reasoning are headroom this job does not use. The most capable model was the wrong pick, which is the general lesson: match the model to the task, not to the leaderboard.
The free tier is not off the table, but read the finding correctly. Its recall is rock solid, so it never misses a film. Its precision is a coin flip on the hard cases, so it occasionally grabs an extra wrong one. That is a fine trade only if you gate low-confidence matches behind a year-and-identity check before anything reaches the downloader. Do that and glm-4.5-flash gets you the whole recall win for nothing. Skip it and you are one unlucky run from the wrong film on your shelf.
What this benchmark is not
Small and deliberately hard: 10 films, 115 candidates, 9 negatives. The precision numbers are coarse at that size, so read the direction, not the third decimal. Ground truth is my own labelling, and a couple of genuinely ambiguous strings (a bare āBrother 1997ā, franchise packs) I labelled to the most notable film. Five runs per model at temperature 0, one prompt, the plain untuned one above. No Western-model cross-check yet, which is the next thing I want to run: the same 115 titles against Claude Haiku and a small GPT, because if a free Chinese model already clears the recall bar, the interesting question is who else does, how stably, and for how much.
The takeaway
The pipeline was never short on films. It was short on reading. For years the honest fix for this in the self-hosting world has been to hand-maintain regex title-cleaning rules per indexer, which is exactly the kind of unpaid janitorial work that never ends. A small language model does the whole recall job for a dollar a thousand, and the free one does it for nothing if you put a guardrail on its precision.
The fix is a thin matcher that sits in front of the parser: let the model read the title, and keep the regex for the clean releases it already gets right. I have benchmarked it, not yet wired it into the live pipeline, and I will not pretend otherwise. If you self-host the *arr stack and you have ever watched it swear a film does not exist while it seeds on three indexers, that is the bug, and this is the shape of the fix. The benchmark and the labelled set are small enough to rerun in an afternoon. I would start by measuring your own miss rate before you spend a cent, because you might be surprised how much of your catalogue is quietly foreign.
This is build-in-public from devil.services, one field note from the filmoteka build. The next one is the shim that puts this in the pipeline for real, and the run where I find out whether a tuned prompt makes the free model trustworthy. And the pattern already has a second domain: in Regex canāt count pages the same judge idea moves to ebooks, where the thing regex canāt read is a page count.