Gemini 3.7 Flash High vs Claude Opus 5 High vs Grok 4.6 High vs Qwen 3.8 27B
Mia will there be a GLM 5.3 flash for Nvidia 5090 RTX 24 GB VRAM Users ?
Like a modded 27B Version in EXL3 ? Also i get stunning results with the Qwen 3.8 27B with 5090RTX and max context window i get 90-100 tokens per second.
"Hey codex, make my LLM run 5x faster"
Introducing Tama 🐣 Give your agent a computer with GPUs Sync your code, spin up boxes, run experiments in parallel I pointed Codex at Qwen 3.8 and tested 5 spec decoding methods in parallel, and made it 5x faster than default vLLM
Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this…
Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this model but I just feel like my current setup is the smarter one.
You have reached the end of the archive
All of qwen38