REALLYCUTEPUPPIES.AI
Standing results · updated 2026-09-07 · n=3 per configuration

The Gauntlet

Six model configurations ran an identical long-horizon enrichment pipeline three times each — real browser, native context, hundreds of judgment calls — scored against human-adjudicated ground truth with error weights fixed before any results existed.


Quality ranking

606570758085Claude Sonnet 576.7Qwen flash Q8 (self-hosted)75.7Claude Opus 572.8Claude Opus 4.871.4Qwen 27B BF16 (self-hosted)70.5Claude Opus 4.668.3● mean — min-to-max of 3 runs
Fig. 1 — Score = 100 − Σ weighted errors (weights 1 / 0.2 / 3 / 6, ratified pre-results). Red = self-hosted open weights, blue = Anthropic API. The headline: a self-hosted model one point behind the field's best.
#ConfigurationMeanRunsTime/runCost/run
1Claude Sonnet 5 (API)76.781.2 / 77.2 / 71.61.4–1.9 h$56–85
2Qwen flash Q8 (self-hosted, native 262k)75.778.4 / 75.8 / 73.03.1–7.4 h≈C$0.46–1.12 elec.†
3Claude Opus 5 (API)72.874.0 / 73.0 / 71.41.2–1.4 h$85–254
4Claude Opus 4.8 (API)71.476.6 / 72.2 / 65.41.0–1.1 h$55–84
5Qwen 27B BF16 (self-hosted, native 262k)70.578.4 / 70.6 / 62.64.3–6.9 h≈C$0.64–1.04 elec.†
6Claude Opus 4.6 (API)68.370.4 / 70.2 / 64.41.3–2.1 h$24–47

† Electricity model: ~1.0 kW at the wall × BC residential Step-2 ≈C$0.15/kWh × productive hours, for a workstation whose VRAM fits both configurations; the model reproduces this lab's independently measured single-GPU run cost. As-tested rented-GPU figures ($2.75/GPU·h) ran $12–41/run.

Reading it as a buyer

Best value, if you can self-host: the flash Q8 configuration — one point off the field's best quality at effectively zero marginal cost, with the second-tightest replica spread, and its best run declined the eval's designed trap.

Best quality, API-only: Sonnet 5 — the ceiling of the field and 2× the wall-clock speed, at 19–28¢ per record. Its own 9.6-point spread means a bad roll lands mid-pack.

Not recommended for this workload: the 27B (a 15.8-point lottery between identical runs — see the n=3 flip); Opus 4.8/4.6, outscored at ~50× the marginal cost; Opus 5 only where its exceptional 2.6-point consistency outranks both cost and peak quality.

Results by error category

Configurationomittedjunk keepssites missedwrong sitesfalse linksfalse flagsterritoryscore
Sonnet 51.33.00.30.00.01.05.776.7
flash Q81.33.02.30.70.70.72.075.7
Opus 50.34.72.70.00.01.74.072.8
Opus 4.80.72.31.31.01.00.36.371.4
27B BF163.03.71.30.71.01.04.770.5
Opus 4.62.32.73.00.71.00.00.368.3

Mean errors per run within each configuration, vs operator-ratified ground truth. Omissions/inclusions ×1, field-level under-claims ×0.2, verified-false assertions ×3, hijacked domains published ×6. The wrong-site and false-link columns are dominated by one designed acquisition trap, which has fooled 12 of 23 scored runs at every precision and context size.

The workload, and its rules

300 frozen map-platform listings — one category of local physical-service business, one US state, full of real noise: duplicates, look-alikes, defunct operators. Each run filters imposters, untangles brands and acquisitions into parent companies, verifies every live website (catching dead, hijacked, and template-generated fakes), extracts attributes with cited evidence, maps service territories, and publishes a relational dataset behind an integrity gate. Its hard rules shape every number above: judgments must come from model reasoning (regex and fuzzy matching are forbidden for decisions), every published fact needs evidence or must be null, and the only network access is a real headful browser the model drives itself, one throttled page at a time.

Disclosures

Scoring used Claude-based judging with a data-only collector fleet — Claude models are contestants, and the judge shares a vendor with half the field. One task domain, one US state, anonymized. Launches that ended without publishing were resumed or re-run with state verified; wall-clock and cost figures exclude infrastructure- and operator-caused segments (all classified from transcripts — autopsy here). A GLM-5.3-Flash run also completed the workload but is excluded from this ranking (n=1, ran under two handicaps); its scorecard is its own story.