MoE compute is sparse. MoE memory is not.
A token activates 2 of Mixtral's 8 experts (25% of FFN weights), or 8 of DeepSeek-V3's 256 (~5%). But every expert must live SOMEWHERE, and no single GPU holds them all. Expert parallelism inverts the usual rule: experts stay put, spread
This AI can literally rewrite itself, and it changes everything!
DeepSeek just released a free tool that builds new features on the fly based on your prompts. Just ask for a tool that doesn't exist, and the system creates it instantly. It runs 100% locally on your computer
Deepseek just dropped a vision model that rivals GPT-4O
It reads images fast and barely uses any tokens.
Some DeepSeek plug-in recommendations that I use myself: 1.
Wrapping DeepSeek Harness for desktop 2. Change DeepSeek theme color and font size 3. Stick the information we sent to the top of the dialog box 4. Sidebar tool similar to CodeX 5.
You have reached the end of the archive
All of deepseek