Mia will there be a GLM 5.3 flash for Nvidia 5090 RTX 24 GB VRAM Users ?
Like a modded 27B Version in EXL3 ? Also i get stunning results with the Qwen 3.8 27B with 5090RTX and max context window i get 90-100 tokens per second.
"Hey codex, make my LLM run 5x faster"
Introducing Tama 🐣 Give your agent a computer with GPUs Sync your code, spin up boxes, run experiments in parallel I pointed Codex at Qwen 3.8 and tested 5 spec decoding methods in parallel, and made it 5x faster than default vLLM
Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this…
Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this model but I just feel like my current setup is the smarter one.
This is the cutting edge period!
Omarchy first Then any other agent harness of your choice. My rank? Hermes Leave openclaw alone they are not serious download qwen 3.8 8b flash model and run locally
You have reached the end of the archive
All of qwen38