This should not be possible on a $300 GPU.
Qwen 3.8 27B is running locally on an RTX 4060. 8GB VRAM. 64K context. 14.6GB on disk. ~150 tok/s prefill. ~5 tok/s decode. And somehow it stays inside those 8GB. No 4090. No $10,000 workstation. No cloud bill quietly eating your
Exploring building my own bespoke sovereign agent middleware + front-end.
✅ Agent-agnostic: agent teammates can be any mix of Hermes, OpenClaw, Codex, LangChain, etc. ✅ Model-agnostic: any OpenAI-compatible endpoint ✅ Open standards: ACP, AG-UI, MCP, A2A (eventually) ✅
Qwen 3.8 27B running smooth on iPhone.
No cloud required for this level of coding help.
A full Call of Duty-style FPS, built from scratch by Qwen 3.8 2.4T and runs on locally
You have reached the end of the archive
All of qwen38