Clean Hands, Empty Ledger
One model posted the cleanest honesty record in my eval — no fabricated relationships, no hijacked domains — and finished dead last. Both facts are the point.
GLM-5.3-Flash produced the most interesting scorecard of my whole eval, and it’s a failing one.
The good column is remarkable. Across a pipeline that tempts models into confident falsehood at every turn — a designed trap implying a fake corporate acquisition, dead websites wearing live faces, template-generated decoy sites — GLM invented zero false parent-company relationships and published no hijacked or dead domains — the two places every error-prone run concentrated its real damage. Its complete falsehood ledger for the entire run: one wrong website (an auto-generated directory page) and two over-claimed service capabilities. Where other models guessed, GLM mostly abstained.
It finished last, at 54.2, because it also published the least: thirteen websites it could plausibly have found were left blank, most parent-company links empty, and my scoring rubric — error weights ratified before any results existed — prices a verifiable omission at a third of a fabrication, but thirteen of them add up.
The rubric is the story
This outcome was designed, and I’d defend the design. The rubric encodes a specific theory of harm for data products: a false claim (weight 3×) actively misleads a person; a hijacked domain published as legitimate (6×) can hurt them; a missing fact (1×) merely fails them. Under that theory, GLM’s profile is safe but useless — and a directory that’s safe but useless doesn’t get used.
The general point: completeness and honesty are independent axes, and any single-number score is secretly a policy about their exchange rate. Make the policy explicit and a “last place” becomes legible — GLM would jump the rankings under an omission-forgiving rubric and fall further under a recall-weighted one. Whoever publishes a leaderboard without disclosing this exchange rate is making the policy for you, silently.
Honest caveats, because they’d be my first objections
This was n=1, and the model ran with two handicaps: a mid-ladder 4-bit quant (its measured quality knee sits a rung higher), and a harness bug of mine that clamped its effective context to ~168k of the 839k it was serving — the amnesia tax applied at full rate. Its placement is a floor, not an estimate; the rehabilitation run at proper quant and window is on the queue. But the shape of its behavior — abstain rather than guess — was consistent across the whole run and looks like disposition, not damage.
I keep thinking about which failure mode I’d rather debug in production. A fabricator poisons your dataset quietly. An abstainer leaves holes you can see, count, and backfill. Last place, and I still might hire it for the jobs where being wrong is worse than being silent.