And now I also have a swarm running Qwen 3.8 27b in my nvidia RTX 5090.
I like where this is going so far. Pretty locked in. gm
Ornith-1.5-35B-A3B-oQ8e-mtp still strong in speed with omlx on M3 Ultra even with 150K context: 45 t/s
Surely it's not at the same level of Qwen 3.8 27B honestly. To try getting same results I'm constantly nudging, steering and 🤬 at it.
2 x 4090s running Qwen 3.8 27B at a total of over 300 tok/s.
8 concurrent sessions with vLLM, Qwen3.8-27B-INT8-W8A16-MTP, 128k context. Over 1 million tokens per hour for my AI agents to use. Without paying for subs or APIs. Without giving my data away to AI companies.
If you're testing Qwen 3.8 Max, don't waste it on tiny prompts.
Use it where the long context actually matters. Try these: 1. Feed it your full company knowledge base. 2. Give it several SOPs at once. 3. Add your past customer conversations. 4. Include common objections and
You have reached the end of the archive
All of qwen38