REALLYCUTEPUPPIES.AI
A home LLM lab, measured

Field notes from running frontier-class evals on my own silicon.

One real production workload. Twenty-five instrumented runs across two open-weights families and four generations of Claude. Mining GPUs, thermal forensics, and receipts for every number. These are the notes, posted raw.

Headline result: a self-hosted model finished 1.0 point behind Claude Sonnet 5 on a judgment-heavy enrichment pipeline, at about a dollar of electricity per run. See the full board →

Hardware · Part III

The No-Op That Wasn't

The chassis RGB lives in SMM firmware with no Linux driver in existence. A write designed to change nothing turned the lights off — and seven confident beliefs died in 27 hours.

Hardware · Part II

Five Millimetres

A 3.5-slot GPU leaves 5 mm above the next slot. Every riser on the market needs 8–15. A tour of CAD teardowns, a vendor with no drawings, and the fix that made it all moot.

Hardware · Part II

The Temperature nvidia-smi Won't Show You

My 3090 sits 3 mm above a 4090. Under sustained load its core read a comfortable 70°C — while the memory junction, invisible to standard tooling, ran 94. Getting that number took reverse-engineering a monitoring app's shared memory.

Eval · Part II

Clean Hands, Empty Ledger

One model posted the cleanest honesty record in my eval — no fabricated relationships, no hijacked domains — and finished dead last. Both facts are the point.

Eval · Part II

The Amnesia Tax

85% of my agents' tokens were spent rebuilding memory they'd been forced to throw away. One environment variable and a native context window made runs 7× faster — and better.

Eval · Part I

The Harness Beat the Model

Identical weights ranked last in one agent stack and first in another. Most leaderboard deltas I measured live in the plumbing, not the parameters.