Fruit-fly connectome optimized inference for DeepSeek-V4.1-Flash on 2× NVIDIA GB300 DGX Stations!
It impressively hits 56k tok/s prefill and 201.7 tok/s c1 decode. Well done little dude.
Which flash model wins?
DeepSeek v4.1 flash, Qwen 3.8 flash, or GLM 5.3 flash 🤯
Okay, I guess I didn't waste all that money after all
Deepseek V4.1 Flash 304 tok/s code, 135 tok/s prose Zero optimization, will have it max optimized by the morning.
A maze game completed in one go using DeepSeek V4.1 Flash
It took less than half an hour to complete the in-depth search.
You have reached the end of the archive
All of deepseek