I had two of the best local models build the same floating tree
Glm 5.3 flash at 320b and qwen 3.8 flash next at 125b, and despite the size gap look where both landed glm 5.3 flash: 320b moe with 18b active, nvidia's nvfp4 weights split over 2x dgx spark, served with vllm
Qwen 3.8 27B + TensorFold built a game for my Hardball 3D challenge in 3h 14m.
Of the three builds, it’s the most playable, even more so than the GPT-6 Luna baseline. Decode · 90th percentile: 42.37 tok/s Prefill · mean: 160.55 tok/s Run details:
Paris Nocturne - by Qwen 3.8 Flash Next at 3bpw no thinking loop no break one shot prompt 3 hour working...
This complexity, the thinking process almost same as intelligent with Opus 5.5 or Fable 5.. look at the shading sky, the horizon, the baloon, ferrieswheel, the river, the
[AI News] October 5, 2026 Early morning top news ・Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
You have reached the end of the archive
All of qwen38