27 Apr 2026
RTX 3090 — returned it after testing local LLMs
Bought an RTX 3090, spent a few days pulling open-weight coding models and wiring them into local agent stacks. Ultimately returned the card.
The models looked promising — local models claim frontier-adjacent benchmarks. Reality on a single consumer GPU is different. You're running quantised (Quantized — compressed) versions of smaller models, and on multi-file context and harder tasks the gap to frontier shows up fast. Maybe 70% of daily prompts are fine. The other 30% — the ones that matter — aren't.
Then there's the harness. Cloud tools have years of polish on tool use, error recovery, and context management. Local stacks don't yet. So you're losing on both axes — slightly weaker model, noticeably weaker harness.
This makes the cloud subscriptions actually seem reasonable - I get frontier models plus tested agent loops. Expensive GPU + electricity costs + my time got me a worse experience on the tasks that matter.
Local intelligence is closer than I expected. Still not close enough.