Watch the progress of frontier models in bringing Ancient Rome to life on Arena, from Q4 2025 to now.
Scores for Claude Sonnet 5.5 by AnthropicAI and GPT-6.1 by OpenAI are coming soon. Real-world tasks from our global community of users power the Arena leaderboards. Head to
I'm testing the operation of the Qwen 3.8 Flash next IQ3
I'm testing the operation of the Qwen 3.8 Flash next IQ3, which I think has completed 90% optimization (making a list of models waiting to be served), and waiting for the Oai Devday announcement. Is this how Noah felt as he nailed the ark?
I tested Bonsai 2 27B
They claim it retains 98.2% performance of the original Qwen 3.8 27B but my results suggest otherwise. It can handle basic tasks, but anything beyond simple HTML gives unusable outputs. Don't get me wrong, it's already impressive compressing a 5.9GB model
Qwen 3.8 Next Flash + OMP
ONE SHOT - Full blown 3d scene editor. What is this sorcery o_O Assets made locally by HY3d2.1 + Qwen next + Qwen Image It's too much for me ;) I need to sit and think. ps. You can also test/walk on the stage
You have reached the end of the archive
All of qwen38