We ran Gemini 3.7 Flash, Gemini 3.8 Flash and self-hosted Qwen3.8-27B through the same grounded video QA pipeline on 300 Perception Test questions.
Gemini wins end to end, Qwen lands at 84% of its HOTA. The biggest lever was not the model at all. Write-up:
Qwen 3.8 27b via Tesseract Server
What a phenomenal model that can be run on a Mac
Try out cache-to-cache here
ONE DGX Spark Working with MiaAI_lab recipe of Qwen 3.8 flash next, but you can use any model Simple test of an essay about the research paper:
No single-stream headline today, so here's the part builders actually argue about
Two boxes running the same Gemma 4 12B IT QAT at 4-bit, and the spread between them is 3%. 29 tok/s on Strix Halo (Q4_K_XL). 28 tok/s on DGX Spark (Q4_K_M). Same model, different machines, and the
You have reached the end of the archive
All of qwen38