KAT Coder V2.5 Dev (Q4_K_M) - 162 tokens/sec decode - Single RTX 4090 (24GB VRAM) - 8500+ tokens/sec…
KAT Coder V2.5 Dev (Q4_K_M) - 162 tokens/sec decode - Single RTX 4090 (24GB VRAM) - 8500+ tokens/sec prefill - 250k context - almost fully in VRAM - llama.cpp After force feeding this 4090 a 91GB behemoth (DeepSeek V4 Flash 0731) 2 days ago and watching it choke through 12 t/s,…
Today is the day! Qwen 3.8 27B open weights.
I asked Hermes Studio to create a desert island using threejs 🏝️ For me who can't travel abroad during summer vacation lol
Go's Qwen-3.8-MAX used up the 4-hour limit twice in no time and finished halfway, so I had Codex's GPT 5.6 Luna help me with the final adjustments.
GLM 5.3 Max vs Qwen 3.8 Max vs Grok 4.6 vs Opus 5
GLM 5.3 Max is actually holding up really well here. My ranking for this run: Opus 5 > Qwen 3.8 Max > GLM 5.3 Max > Grok 4.6 Definitely competitive now. GLM is getting scary close.
You have reached the end of the archive
All of qwen38