I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX Spark.
Cruz shipped a custom ExLlamaV3 fork and a native engine path. I ran the same model on the same Spark: 102.6 max, 80 tok/s on code, 70 tok/s on prose. My previous benchmark and
There's been a lot of hype about Ternary Bonsai2 but fw demonstrations on how it actually compares to the full 16 bit version of Qwen 3.8 27B.
Since I have my own quantized version of that model thatI'm testing, ~12gb file size, I decided to throw Bonsai2 in there also to get a
Ahora probe la diferencia entre Qwen 3.8 27B local contra GLM 5.3 ambos en Droid
Se nota una diferencia de calidad y de poder seguir mas de cerca las imagenes y la atencion al detalle. Claramente GLM 5.3 esta por delante en calidad, aunque demoro mucho mas en construir el juego.
Qwen 3.8 Omni Flash can WATCH your videos, not just read the transcript.
That means AI can finally understand what was said AND what happened on screen at the same time. Here’s how I’d test it: → Open the Qwen chat app. → Pick Qwen 3.8 Omni Flash. → Upload a video. →
You have reached the end of the archive
All of qwen38