Gaming notebooks can now directly run cutting-edge large models, with official weights, without extreme quantification.
Qwen3.6 35B, 4060 notebook with 8G video memory, 39 tok/s. DeepSeek-V4-Flash 284B, 5090 desktop, 22-25 tok/s. GLM-5.2 753B, one PRO 6000, 15 tok/s. Claude Code and Codex are direct local prostitutes, and the electricity bills are cheaper than API.
MoE compute is sparse. MoE memory is not.
A token activates 2 of Mixtral's 8 experts (25% of FFN weights), or 8 of DeepSeek-V3's 256 (~5%). But every expert must live SOMEWHERE, and no single GPU holds them all. Expert parallelism inverts the usual rule: experts stay put, spread
This AI can literally rewrite itself, and it changes everything!
DeepSeek just released a free tool that builds new features on the fly based on your prompts. Just ask for a tool that doesn't exist, and the system creates it instantly. It runs 100% locally on your computer
Deepseek just dropped a vision model that rivals GPT-4O
It reads images fast and barely uses any tokens.
You have reached the end of the archive
All of deepseek