DeepSeek v4.1 flash, Qwen 3.8 flash, or GLM 5.3 flash 🤯
You can just hallucinate the entire internet with Qwen 3.8 27b running at 2,000 tokens/second?
Part 2 of turning cerebras + Alibaba_Qwen 3.8 27B into an OS: built an offline browser with zero network calls and mounted it directly the JIT ubuntu desktop. no wifi. no scraping.
Qwen 3.8 Max Full COURSE 1 HOUR (Build & Automate Anything)
Qwen 3.8 Flash Next with llama.cpp on a Macbook M5 24 GB.
3/4 tok/s - reasoning off - 2k context With these numbers it is definitely unusable especially for the prefilling phase. With a larger context it will takes hours. I think I reached the limit also with this model.
You have reached the end of the archive
All of qwen38