GLM 5.3 Flash + Qwen 3.8 Flash are FREE right now 😳
You can also try: - Hy3 - MiMo 2.5 no idea how long the free access lasts try them here
Qwen-3.8-Flash-Next and GLM are awesome, which one is stronger?
I won’t run the benchmark today. I will make a cyber review: just let them fight. These two actually represent two very different routes. Qwen-3.8-Flash-Next has begun to enter the scope where ordinary people can deploy, quantify, and toss the runtime locally. GLM Niulai is a completely different kind of monster:
I just ran Qwen 3.8 flash next through my benchmark....
Terrible. Use GLM 5.3 flash instead
Qwen 3.8 Flash Next (125B-A6B) now runs natively in mlx-serve on Apple Silicon.
No Python, Zig + Metal. (Preview Release) M4 Max, 4-bit pack (~75 GB resident): - ~70 tok/s serial, ~98 tok/s with MTP when tested using `npx llmprobe` - prefill ~700 tok/s, sparse attention past 2k
You have reached the end of the archive
All of qwen38