FreeToken lets gaming GPUs serve frontier models at interactive speed using official checkpoints with no extreme quantization required.
→ Qwen3.6 35B on an RTX 4060 laptop at 39 tokens per second → DeepSeek V4 Flash 284B on an RTX 5090 at 22 to 25 tokens per second → GLM-5.2
Had the opportunity to dive into specular decoding, one of my favorite topics
While introducing the DSpark paper, published by researchers at DeepSeek, at the machine learning reading group earlier this week.
Kimi K3 cuts each layer into 896 networks.
The router would starve nearly all 16 of them run per token. at init all 896 are equally bad. whichever ones the router hits first get updated more often. they get better, so they win more tokens. the rest go dead and never come back
Tried recreating this experiment with a KMP app and a similar prompt.
Deepseek Harness with Ox Alpha is already very impressive
You have reached the end of the archive
All of deepseek