Oneshot qwen 3.8 27b q4 on mac m5 (8 mins) and the second one on rtx4090 (just under 3 mins) -- same model
Prompt borrowed from: (thanks!)
Qwen 3.8 27B is an open artificial intelligence model that you can run directly on your device without a subscription and without paying per use.
What is interesting about it is that its performance has become very close to powerful and expensive models such as the Claude Opus 4.6, and in some tests it has surpassed it. It also supports a huge context that reaches 262 thousand tokens, and is designed for programming and implementation…
Qwen 3.8 27B performance bench on 1 x RTX 6000.
After lots of crashes, found n3 the sweet spot for spec-on. Had it working as high as n9 at 275 tok/s but killed concurrency and random crashes during load. Will have eval v3 and impossible task results soon + Spark/3090 numbers.
You have reached the end of the archive
All of qwen38