Qwen 3.8 Flash Next with llama.cpp on a Macbook M5 24 GB.
3/4 tok/s - reasoning off - 2k context With these numbers it is definitely unusable especially for the prefilling phase. With a larger context it will takes hours. I think I reached the limit also with this model.
Qwen 3.8 Flash Next is hitting my MacBook.
That’s the sound of local AI when you try to force one of the best AI models to fit on a tiny machine with 24 GB of RAM
Today I deployed a Qwen 3.8 27b quantitative version locally.
What do you think of this speed? Can it be used? 😂
PERPLEXITY: The Portable Computer now runs on Windows, agents and models work locally on its own RTX card.
The prerequisite is 24 GB VRAM, PPLX 27B calculates locally based on Qwen 3.8 27B. There are also local MCP servers and scheduled tasks. Who client data so far
You have reached the end of the archive
All of qwen38