AI Highlights of the Week: GLM-5.3 is launched on Zhipu, DeepSeek V4 Pro is heavily updated, and Google launches Gemini 3.7 Flash
BigMoeOnEdge: streaming MoE inference on Android, built on llama.cpp (ggml_org).
DeepSeek V4 Flash IQ2_M (92 GB) running on a 12 GB midrange phone at ~1 tok/s. Qwen3.6-35B-A3B up to 7-8 tok/s, Gemma 26B and Qwen3-30B in the same tok/s range. CPU only. Experts stream from UFS
Xianyu is the largest AI hotspot in China and the most profitable self-media platform in China🙃 The previous one is MiniMax H3
这次是DeepSeek Harness Ps: 现在还有人蹲Mac mini M4 吗?🙂
DeepSeek Harness finished the same coding task almost 3× faster than Claude Code.
But Claude still produced the better result. Julian Goldie and Kasra gave both systems the same one-shot prompt: build an animated accounting website and a working Tetris game. The results: 00:41
You have reached the end of the archive
All of deepseek