I run the 320B model for free at home, but only the 18B is actually turned on.
This is a scene from the thesis channel Two Minute Papers running Zai's GLM-5.3-Flash at home. Light refraction simulations, strategy games, and even Blender 3D scenes. The value is 0 won. It's open weight. This model first came out under the name Ox Alpha.
GPT-5.6 Sol paired with DeepSeek Harness is ridiculously good at building voxel dioramas.
In a single pass, it created an actually playable voxel battlefield 👇 It includes: Trenches, craters, ruins, a coastline, and a railway 34 soldiers and 17 animated land, air, and naval
Today we launch our Open Model Guide for Enterprise Use Cases 👇
Too many models. GLM. DeepSeek. Qwen. Kimi. Gemma. Nemotron. MiniMax. Maybe two more will be shipped this week. Access to open models stopped being the problem a while ago. Choosing the correct one is the problem
671B MoE. 14GB of VRAM.
16.8 tokens/s decode. That is DeepSeek-V3 Q4_K_M on a single RTX 4090 — not an 8×H100 rack. The hot path stays on the GPU. The giant expert banks live in system RAM. 🆕 KTransformers (kvcache-ai/ktransformers, Apache-2.0, Tsinghua MADSys, SOSP 2025).
You have reached the end of the archive
All of deepseek