The gap in token usage v/s output for low/medium/xHigh thinking on Qwen 3.8 Flash Next (and many models similarly) is mind boggling.
I gave the same Voxel Pagoda Garden prompt to all 3 (The EXL3 4.05 bpw variant on my dgx spark) and the outcome was interesting, yet re-affirming
What a crazy week in AI!
🚀 Gemini 3.8 Live Qwen 3.8 Omni Flash Qwen 3.8 Live Translate Bonsai 2 27B Needle 3 Jev Laya Nimble Occamy ZGCM Meridian R2T2 Jing Dao Dream RSI MiniMax Code & more! Watch the full recap:
Ran FlappyBench on Qwen3.8-Omni-Flash, DeepSeek-V4.1-Flash, and Gemini 3.8 Flash with the same design prompt.
🔹 Qwen 3.8 Omni Flash: 9/10 · $0.013 · smooth gameplay and cheaper 🔹 DS V4.1 Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 Gemini 3.8 Flash: 5/10 · $0.205 ·
Running Qwen 3.8 Flash Next NVFP4 on one RTX 6000 Pro 96GB.
$2/hr if you need some quick work done
You have reached the end of the archive
All of qwen38