Kawaii Invader made with Qwen 3.8 27B ROCmFP4 Quant.
This came out looking good.
OpenAI charging $15/M tokens while Qwen 3.8 27B hits 65 tok/s on a single 4090.
Inference costs are collapsing. How long until the cloud dependency becomes optional?
Qwen 3.8 - 27B
It was finally released so I tried it right away. Is it the setting that is unusually slow (9t/s)? Is the default size (8bit quantization model) mismatched to 128GB/m5max? In any case, I quickly shared the data as it was interesting. 1st LMstudio Simple Reasoning 2nd ollama+claudecode collaboration inference large…
You have reached the end of the archive
All of qwen38