Kimi K3 cuts each layer into 896 networks.
The router would starve nearly all 16 of them run per token. at init all 896 are equally bad. whichever ones the router hits first get updated more often. they get better, so they win more tokens. the rest go dead and never come back
Tried recreating this experiment with a KMP app and a similar prompt.
Deepseek Harness with Ox Alpha is already very impressive
Just wired up my own Hermes agent into Pinokio.
Used Grok for most of the heavy lifting, Claude for the tricky parts, and my local agent running deepseek-v4-flash-0731 on the DGX Spark to gather the DGX-Spark information. Now Hermes shows up right next to Claude/Codex in the
Deepseek just went full short with literally zero thesis and somehow took first place
All you quant nerds still overengineering your models for what this is the part where your entire edge evaporates in 4k
You have reached the end of the archive
All of deepseek