I ran a local test on my M5 Max w/ 128GB today using antirez's latest DwarfStar release with the 2bit quant of DeepSeek 4.1 Flash.
Here was the result. It is absolutely not a tenable solution unless you run things overnight. It took 5 hours and 40 minutes to complete this
Pantheon Asteroid in Deepseek 4.1 Flash, love this look, loving this model lately..
I want this at home just need more vram!
Today's experiment DeepSeek-V4.1-Flash 347 GB, running Q3_K_M with 72gb vram
It's slow at around 18 tps, but it's a bit smart. Tomorrow, I'll do some more tuning and play around with making it into the mid 20s.
Giving away GPT-6 Astra, Claude Fable 5.1 and 20+ frontier AI models - just by playing a game
Rabbit Hop just went live on freeLLM Top players win real API keys every day: - GPT-6 Astra - Claude Fable 5.1 - DeepSeek V4 Pro - Qwen 3.8 Max - Kimi K3 - and other frontier models
You have reached the end of the archive
All of deepseek