Qwen 3.8 27B at 56tps; on 9 year old GPU btw
Nvidia V100 32GB ~$650 on EBay right now! Using Dflash 2; disabling the ECC adds some more speed too! Thinking and prose is a bit slower, but 56-63 tps in code gen! MTP runs faster for prose vs DFlash2 but slower sustained code
Lowk impressed i gave a local Qwen 3.8 my nix repo
Asked it to simplify whatever it can this was the result after 2 hours
Perplexity just put its full AI agent on your computer.
And the security model might be more important than the AI model itself. What changed: → The model, orchestrator, tools, reasoning, and file access can run locally → Launch model: Qwen 3.8 27B, plus Perplexity’s
This mysterious app was created autonomously by Clio Agent 3 Beta without any instructions.
Base model is Qwen 3.8 27B
You have reached the end of the archive
All of qwen38