Alibaba's Qwen 3.8 Omni Flash watches video the way you do
It sees the screen and hears every word at once. Not a transcript. Not screenshots. The actual video. Here's what that unlocks: → Someone says "look at this" and points. The model ties the words to the picture like a
Drop an hour of meeting video and get action items plus edit plans, not another summary dump.
Qwen 3.8 Omni Flash from Alibaba_Qwen takes text, image, audio, and video in. Text out. About 1M context. Plans tasks, calls tools, and drives creative audio-video workflows. Big jump
People underestimate what you can do with local models.
This is Qwen 3.8-27B running on an M5 Max building a 3D racing game end to end.
This new Chinese AI reads 100 hours of your calls and meetings in one shot.
Alibaba just dropped Qwen 3.8-Omni-Flash. Most AI reads text. Your business runs on calls, videos, and recordings. This one takes all of it at once: → Text, images, audio, and video in the same
You have reached the end of the archive
All of qwen38