Qwen 3.8 Max Full COURSE 1 HOUR (Build & Automate Anything)
Qwen 3.8 Flash Next with llama.cpp on a Macbook M5 24 GB.
3/4 tok/s - reasoning off - 2k context With these numbers it is definitely unusable especially for the prefilling phase. With a larger context it will takes hours. I think I reached the limit also with this model.
Qwen 3.8 Flash Next is hitting my MacBook.
That’s the sound of local AI when you try to force one of the best AI models to fit on a tiny machine with 24 GB of RAM
Today I deployed a Qwen 3.8 27b quantitative version locally.
What do you think of this speed? Can it be used? 😂
You have reached the end of the archive
All of qwen38