Andrej karpathy could have charged $2,000 for this course.
He put it on YouTube. The full training stack. Tokenization. Neural network internals. Hallucinations. Tool use. Reinforcement learning. RLHF. DeepSeek. AlphaGo. 3 hours of the most comprehensive LLM education that
284B-class DeepSeek-V4.
Two 24GB 3090s. 18.26 tokens/s decode. The routed experts live in system RAM. The GPUs keep attention. That is the author’s own `llama-sweep-bench` on a hybrid `--cpu-moe` box — not an H100 rack, not a Discord screenshot. 🆕 ik_llama.cpp
DeepSeek-V4.1-Flash's performance metrics suggest superior reasoning and efficiency
But the "psyop" narrative ignores the technical rigor of open-source contributions. The real panic should be about the potential for misuse, not the model itself.
One app just put GPT-6, Claude, Gemini, Grok, Kimi and DeepSeek behind a single search bar with zero token limit
Someone opened their laptop and typed "search models" like they were browsing a menu, except the menu had every frontier lab on it at once. Scroll the dropdown and
You have reached the end of the archive
All of deepseek