My 9-Reviewer AI Gate Failed Articles It Scored 9.1

2026年8月26日2 次浏览来源:Dev.to阅读原文

Nine LLM reviewers score every article I publish.

The gate is simple: zero blockers, and a mean score above my floor.

Yesterday it failed an article it had scored 9.1.

Then it failed a second one three separate times.

Neither reviewer ever disagreed with the writing — all three failures were the harness misreading its own reviewers.

A perfect score, filed as a blocker The Checklist reviewer returned 10.0 and wrote this under : That is a clean bill of health.

It failed the gate.

The parser dropped lines meaning "nothing here" with .

The prompt asks reviewers to explain why nothing is wrong, so they write the explanation instead of the word — and "no biography" is not "none".

Why it survives review: the regex is correct for the string it was written against.

Nobody writes a test for the sentence a model didn't emit.

The fix wasn't a longer regex.

Guessing every phrasing of "nothing is wrong" is unbounded.

The prompts already state the contract — fully in scope means and — so a 10 that also reports a blocker is self-contradictory: Demoted, not deleted.

If a reviewer ever scores 10 on something genuinely broken, the text still reaches a human.

The word that made good reviews look like outages My runner treats a quota message as a transport failure, so a billing error can never be scored as a bad article.

The pattern included a bare .

It was tested against the reviewer's entire output.

My corpus is security articles.

Reviewers say "rate limit" constantly: Three finished reviews, scores visible in the text that got thrown away.

That article lost 4 of 9 reviewers and failed two batch runs before I looked.

It reads as an account problem, which is exactly why it survived.

The tell separates the two cleanly: a real quota banner is the only thing on stdout.

A review that merely discusses limits has a line sitting right there.

An excluded reviewer voting anyway One reviewer is informational — its rubric weights brand fit 40%, which structurally caps anything outside that niche.

It was excluded from the gate score.

It was not excluded from blockers.

So it vetoed through the back door — and what it files under is its own arithmetic: Four "blockers" on an article it had just called original and reproducible.

A half-applied exclusion is not an exclusion — and a weighted composite is exactly the shape that hides one, because the arithmetic looks like a finding.

The pattern Every one of these is the harness misreading agreement as disagreement, and all three fail in the same direction: quietly, toward rejection.

A false pass is loud — something bad ships and you see it.

A false fail looks exactly like a strict gate doing its job, so it can run for weeks while you assume the work is merely not good enough yet.

It is the same asymmetry that makes a leaderboard wrong in the flattering direction, and the same reason measurement bias survives longest when it agrees with what you expected.

If you run an LLM as a judge, assert on the disagreements it cannot logically have: a perfect score with a blocker, a reviewer excluded from the score that still blocks, a transport error carrying a parsed result.

Those are contradictions, and contradictions are testable without predicting a single word the model will say.

Both directions.

A one-sided test would have passed on all three bugs — my earlier fix for the opposite failure is what introduced the second one.

That is the same lesson ground truth taught me that unit tests could not: a test written from the failure you already know about only ever proves you fixed that one.

More on the tooling behind this at github.com/ofri-peretz/eslint.

The three defects above were found on 2026-08-11 across 39 gated articles.

What's the last false negative you found in your own tooling — and how long had it been running?

分享
Baike.dev

baike.dev helps you discover great languages, frameworks, databases, DevOps and cloud-native tools.

Quick links

About

Contribute

Found a great developer tool? Share it with the community.

Submit a tool
© 2026 baike.dev Developer EncyclopediaUpdated daily · Discover great developer tools