Hello! Thank you for this.
I have tested it on an RTX A6000 and the speed of the output is very impressive. I have done a quick comparison with Ollama Qwen 3.8:27B and GPT-5.6 Luna using Open WebUI. For the sake of keeping digital footprint low, my posts will auto-delete soon.
This is the Pelican Test using Qwen 3.8 W4A16 Flash Next running locally.
Looks pretty good to me. What do you think?
Qwen 3.8 Flash Next ROCmi4 built this with OMP on AMD Strix Halo and I am so happy to make a fully working quant of this.
This works in Hermes + OMP and other coding harnesses. It built this retro cube game PLUS enabled it over my tailscale for me. It's a very capable model.
Have you ever had the AI refuse to do what you asked it to do?
It happens: ChatGPT, Gemini and Claude have internal rules that block risky requests. But the risk is not always real, it happens in an analysis of the competition or in a text that talks about health.
You have reached the end of the archive
All of qwen38