Claude Opus 5 xhigh: ~28m · ~170K tokens · ~$4.60 Qwen 3.8 Max: ~35m · ~50K · ~$0.60
GPT-5.6 Sol: ~19m · ~22.5K · ~$0.68 Grok 4.6 xhigh: 23:48 · 113K · $0.68
This should not be possible on a $300 GPU.
Qwen 3.8 27B is running locally on an RTX 4060. 8GB VRAM. 64K context. 14.6GB on disk. ~150 tok/s prefill. ~5 tok/s decode. And somehow it stays inside those 8GB. No 4090. No $10,000 workstation. No cloud bill quietly eating your
Exploring building my own bespoke sovereign agent middleware + front-end.
✅ Agent-agnostic: agent teammates can be any mix of Hermes, OpenClaw, Codex, LangChain, etc. ✅ Model-agnostic: any OpenAI-compatible endpoint ✅ Open standards: ACP, AG-UI, MCP, A2A (eventually) ✅
Qwen 3.8 27B running smooth on iPhone.
No cloud required for this level of coding help.
You have reached the end of the archive
All of qwen38