27 Apr 2026

RTX 3090 — returned it after testing local LLMs

Bought an RTX 3090, spent a few days pulling open-weight coding models and wiring them into local agent stacks. Ultimately returned the card.

The models looked promising — local models claim frontier-adjacent benchmarks. Reality on a single consumer GPU is different. You're running quantised (Quantized — compressed) versions of smaller models, and on multi-file context and harder tasks the gap to frontier shows up fast. Maybe 70% of daily prompts are fine. The other 30% — the ones that matter — aren't.

Then there's the harness. Cloud tools have years of polish on tool use, error recovery, and context management. Local stacks don't yet. So you're losing on both axes — slightly weaker model, noticeably weaker harness.

This makes the cloud subscriptions actually seem reasonable - I get frontier models plus tested agent loops. Expensive GPU + electricity costs + my time got me a worse experience on the tasks that matter.

Local intelligence is closer than I expected. Still not close enough.

LocalLLMLLMOpenSourceRTX3090AICodingSelfHosted

Originally on LinkedIn

newer AI-generated UGC — authenticity rantolder De-Googling my life — local-first philosophy