Running Qwen 3.8 Flash Next on a 12GB Android phone - 2 tok/s
Qwen 3.8 Flash just hit Atomic Agent.
A FREE local agent - $0 per token, runs on your machine. No cloud, no API key. Inference is local GBNF grammar constraints force valid tool calls from any model KV-cache reuses compute across steps. SQLite stores searchable session
Qwen3.8-Flash-Next from Alibaba_Qwen has day-0 support in Atomic Agent!
125B main model, 51B N-gram embeddings, 6B activated per token and 62.5 on SWE-bench Pro. We gave it a folder and it checked every file to sort them by content. Total: 25 steps + 230K tokens ≈ $0.07
3.8 Flash Next runs perfectly on my Android phone (12GB RAM)
As you know, my Bigmoeonedge project enables running massive models on edge devices - such as a mid-range Android phone with 12GB of RAM. Following DeepSeek and various other models, Qwen 3.8 Flash Next
You have reached the end of the archive
All of qwen38