FRAMEWIREIndonesiaUpdated Aug 29Live wire
0:00 / 0:00

You can now run a legit 27B parameter model completely offline on a normal consumer PC🚀

Qwen 3.8 27B. Download LM Studio (free) → search “Qwen 3.8” → official version from the Qwen team on Hugging Face. They have 4-bit, 5-bit, 6-bit and 8-bit quantizations. Higher bits =

MassimoAug 28
0:00 / 0:00

DeepSeek V4 Flash • Qwen 3.8 Flash

GLM 5.3 Flash • Gemini 3.7 Flash

Fabiano FirmoAug 291
0:00 / 0:00

262K context. On a 16GB RTX 5070 Ti.

🤯 Qwen 3.8 27B Q3 hits ~25 tok/s while an adaptive llama.cpp fork streams KV cache between RAM ↔ VRAM. Stock llama.cpp starts thrashing around ~120K context. Same consumer GPU. 2x+ the usable context. This could be huge for local LLMs.

SachinAug 291
0:00 / 0:00

The underlying Qwen 3.8 Max model is huge

2.4 trillion total parameters. But only around 95 billion are activated for an individual task. That mixture-of-experts approach lets it pull in the relevant parts of the model instead of activating everything.

Julian Goldie SEOAug 291