You can just hallucinate the entire internet with Qwen 3.8 27b running at 2,000 tokens/second?
Part 2 of turning cerebras + Alibaba_Qwen 3.8 27B into an OS: built an offline browser with zero network calls and mounted it directly the JIT ubuntu desktop. no wifi. no scraping.
Qwen 3.8 Max Full COURSE 1 HOUR (Build & Automate Anything)
Qwen 3.8 Flash Next with llama.cpp on a Macbook M5 24 GB.
3/4 tok/s - reasoning off - 2k context With these numbers it is definitely unusable especially for the prefilling phase. With a larger context it will takes hours. I think I reached the limit also with this model.
Qwen 3.8 Flash Next is hitting my MacBook.
That’s the sound of local AI when you try to force one of the best AI models to fit on a tiny machine with 24 GB of RAM
You have reached the end of the archive
All of qwen38