The No-Op That Wasn't
The chassis RGB lives in SMM firmware with no Linux driver in existence. A write designed to change nothing turned the lights off — and seven confident beliefs died in 27 hours.
One real production workload. Twenty-five instrumented runs across two open-weights families and four generations of Claude. Mining GPUs, thermal forensics, and receipts for every number. These are the notes, posted raw.
The chassis RGB lives in SMM firmware with no Linux driver in existence. A write designed to change nothing turned the lights off — and seven confident beliefs died in 27 hours.
The capstone: the full pipeline on a 4090+3090 pair — faster than two of three rented datacenter runs, for 68 cents of electricity. Plus the +17% environment variable and the wall that stopped a bigger model.
ATX 3.1 quietly gutted PCIe 8-pin counts — verified unit by unit. Plus the mining-card power socket that looks exactly like the connector that would destroy it.
A 3.5-slot GPU leaves 5 mm above the next slot. Every riser on the market needs 8–15. A tour of CAD teardowns, a vendor with no drawings, and the fix that made it all moot.
My 3090 sits 3 mm above a 4090. Under sustained load its core read a comfortable 70°C — while the memory junction, invisible to standard tooling, ran 94. Getting that number took reverse-engineering a monitoring app's shared memory.
One model posted the cleanest honesty record in my eval — no fabricated relationships, no hijacked domains — and finished dead last. Both facts are the point.
Two configurations swapped rank when third replicas landed — one of them spanning 15.8 points between identical runs. Single-run model comparisons measure the dice.
Transcript autopsy of every run that stopped without finishing: a model that mimicked its own memory system's voice, a promised file that never appeared, and the harness semantics that let both count as 'done.'
85% of my agents' tokens were spent rebuilding memory they'd been forced to throw away. One environment variable and a native context window made runs 7× faster — and better.
A third of Qwen's flash model is a predictive n-gram table that never does a matmul. It quantizes as a step function, offloads to RAM for free, and even runs from SSD.
A price-history chart mixed transaction data with remembered guesses. The web-sourced 'correction' was worse than the guess. Only sold listings survived.
The CMP 170HX is A100-class silicon that unlocks to 64 GB of HBM2e. Buying two meant mapping a market that is part bargain, part minefield.
Identical weights ranked last in one agent stack and first in another. Most leaderboard deltas I measured live in the plumbing, not the parameters.