Ornith-1.5-35B-A3B-oQ8e-mtp still strong in speed with omlx on M3 Ultra even with 150K context: 45 t/s
Surely it's not at the same level of Qwen 3.8 27B honestly. To try getting same results I'm constantly nudging, steering and 🤬 at it.
2 x 4090s running Qwen 3.8 27B at a total of over 300 tok/s.
8 concurrent sessions with vLLM, Qwen3.8-27B-INT8-W8A16-MTP, 128k context. Over 1 million tokens per hour for my AI agents to use. Without paying for subs or APIs. Without giving my data away to AI companies.
If you're testing Qwen 3.8 Max, don't waste it on tiny prompts.
Use it where the long context actually matters. Try these: 1. Feed it your full company knowledge base. 2. Give it several SOPs at once. 3. Add your past customer conversations. 4. Include common objections and
GPT turned six product photos into an actual brand.
I gave GPT, Claude, Qwen, and Grok the same six images and the exact same ecommerce brief. GPT-5.6 Sol: ~11m · ~19K tokens · ~$0.56 Grok 4.6 xhigh: 16:34 · 114K · $0.68 Qwen 3.8 Max: ~25m · ~45K · ~$0.60 Claude Opus 5 xhigh:
You have reached the end of the archive
All of qwen38