My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model.
O Monday I thought 140 was the ceiling. I was wrong. mr_r0b0t built a different engine. I wanted to push the card again to see if I can squeeze more out of it. The headline number moved 30%, the like-for-like number moved
I ran Qwen 3.8 Flash Next on SSD and it's fast!
Using atomic_chat_hq AD-IQ4XS 85GB size Model weight to 24GB VRAM + 32GB RAM and the n-gram stream directly from SSD Result: 22 tok/sec (14 tok/sec @ 76k context) Output is still amazing since we're using 4-bit quant See
No WiFi. No servers. No internet.
A builder just paired Qwen 3.8 27B with Cerebras chips to make an offline "browser" that doesn't load pages — it hallucinates them live, at 2,000 tokens/sec. Type any site. Pick any year. It invents a page that fits — 1999 web, 2045 web,
My kid asked for a riddle monster to share with friends.
Here is Ridley! Tool Stack: Totem LLM: qwen 3.8 27B ComfyUI with ltx_io All locally running AI! Try it: Demo Site: You can make your own here:
You have reached the end of the archive
All of qwen38