FRAMEWIREIndonesiaUpdated Sep 14Live wire
0:00 / 0:00

Deepseek v4.1 Flash on 24GB Macbook.

I think I reached my limit, 1.36 tok/s. It is far from being usable but the feeling of having a SOTA model on my consumer machine is awesome. I started from 0.74 tok/s and almost double it. It was a great journey.

Marco FranzonSep 1447
0:00 / 0:00

I'm using Codex as an orchestrator on Herdr, assigning implementation to DS and Agy, and reviewing the responses without permission.

メリ男Sep 14
0:00 / 0:00

510 GB model. 128 GB machine.

256k context, thinking on. DeepSeek-V4.1-Flash on one DGX Spark, no smaller model, no cloud. The trick: the model only uses 6 of 384 experts per token. So I tell the box what it should be good at, and it keeps only the experts that job actually

0xBakeerSep 1491
0:00 / 0:00

Which one is better, DeepSeek V4.1 Flash or GLM 5.3 Flash?

More cost-effective? I did a test to help you make a better decision on which model to choose.

塔斯海TasihiSep 14