Jensen huang doesn't want you to know you can run a local llm called kimi k3 with 2.8 trillion…
Jensen huang doesn't want you to know you can run a local llm called kimi k3 with 2.8 trillion parameters on the old $500 nvidia card in your closet for $10 instead of $200 opus 5 or $200 chatgpt pro one pip install and every open source llm runs on a gaming gpu: pip install…
INSANE. Qwen 3.8 27B is now running locally on an RTX 4060 with just 8GB VRAM.
64,000 token context window using Unsloth's new IQ4_XS quant, only 14.6GB on disk. Prefill hits 150 tokens/sec, decode at 5 tokens/sec via native MTP. Just 25 GPU layers offloaded to stay inside
Grok 4.6 came out swinging this year — xAI pushed hard on reasoning benchmarks and it's showing.
But Opus 5 still has that edge in depth and nuance for complex tasks. Hard to call honestly. Depends on what you're building. Which one won YOU over in that test? 👀
Real usage data from someone in my chat
Opus 5 high — 95.2M tokens = 72% of plan Grok 4.6 extra high — 113M tokens = 5.7% of plan Roughly 20x, measured on actual work. If you can't afford the frontier tier, you can still build almost anything. You just can't build something
You have reached the end of the archive
All of Claude Opus 5