Alibaba just quietly shipped an AI that watches video like a person — start to finish.
No transcripts. No screenshots. No doing half the work yourself. It's called Qwen 3.8 Omni Flash. The word that matters is "omni." Most models read text. Some can look at a picture.
You: need complete answers to test the system.
Model: "Sorry, I can't help with that." 🙃 Qwen 3.8 27B Uncensored: fully uncensored, 131K context, chat/coding/agent, cache read Rp. 2,160.
Qwen 3.8 omni realtime Japanese voice is
We ran Gemini 3.7 Flash, Gemini 3.8 Flash and self-hosted Qwen3.8-27B through the same grounded video QA pipeline on 300 Perception Test questions.
Gemini wins end to end, Qwen lands at 84% of its HOTA. The biggest lever was not the model at all. Write-up:
You have reached the end of the archive
All of qwen38