FRAMEWIREIndonesiaUpdated Sep 2Live wire
0:00 / 0:00

"Hey codex, make my LLM run 5x faster"

Introducing Tama 🐣 Give your agent a computer with GPUs Sync your code, spin up boxes, run experiments in parallel I pointed Codex at Qwen 3.8 and tested 5 spec decoding methods in parallel, and made it 5x faster than default vLLM

Eli MernitSep 23
0:00 / 0:00

Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this…

Quad 3090s Hitting 90 tok/s + on Qwen 3.8 Flash on llama.cp | Definitely wouldn't mind running this model but I just feel like my current setup is the smarter one.

Tech2WildSep 21
0:00 / 0:00

This is the cutting edge period!

Omarchy first Then any other agent harness of your choice. My rank? Hermes Leave openclaw alone they are not serious download qwen 3.8 8b flash model and run locally

Marvel 🏆Sep 2
0:00 / 0:00

Anthropic said in the early morning that Fable 5.1 saves 45% on long tasks, and Artificial Analysis…

Anthropic said in the early morning that Fable 5.1 saves 45% on long tasks, and Artificial Analysis said in the middle of the night that each task is 20% more expensive. Neither side lied. It is smarter and can eat more tokens. It saves your time, not your bills. On the same day, Qwen 3.8 Max reached the top of Code Arena, with official $2 entry and $6 exit.

雨哥向前冲Sep 2