I let 16 AI agents build the Colosseum in Minecraft.
8 communicate via a messaging board as a team. 8 others work solo in parallel. The team's WORST agent beat the BEST solo agent. (0.72 vs. 0.58). All using Qwen 3.8 27b via 5090's. The era for Swarm engineering has begun.
I unlocked the full MTP potential in this new update of cafe-llama.cpp, it doubles decoding speed.
Qwen 3.8 Flash Next is over 50t/s, it was 25t/s Qwen 3.6 35B Q8 100t/s was 40t/s Qwen 3.8 27B 80t/s was 40t/s recommended --spec-draft-n-max 4
Local model test on my DGX Spark: Qwen 3.8 Flash vs. a V8 engine.
One HTML file. Code-only geometry. 8 pistons in firing order, exploded view, X-ray mode, touch orbit on iPhone. No cloud. Full prompt in the first comment.
Wtf Turns out Qwen 3.8 27B can make and edit videos like Opus 5.5
Literally, why did no one think to try this before?
You have reached the end of the archive
All of qwen38