GLM 5.3 Max vs Fable 5 vs Qwen 3.8 Max vs Grok 4.6
Gemini 3.7 Flash vs Qwen 3.8 vs
DeepSeek V4 Pro vs GLM 5.3
This is how Qwen 3.8 27B looks at around 30 tok/s in mtplx
I will try a smaller quant cause it's 8bit and it grows to around 35 gb in ram usage
Stop paying API visual context taxes to the cloud cartel.
I just benched Qwen3.8 27B with vision on a single NVIDIA RTX 4090 (24GB). I fed it a massive 30,000 token payload (28.5k text + 1 hi res image) to see where the physical VRAM and throughput ceilings are with and without
You have reached the end of the archive
All of qwen38