Local thinking battle: Qwen3.8-27B vs DeepSeek V4 Flash
Same Eiffel Tower voxel prompt, 3 thinking levels each. Results: thinking off: Qwen 11.2k tokens @ 55 tok/s, DS 6.2k @ 42 medium/high: Qwen 12.8k @ 50, DS 19.7k @ 38 max/xhigh: Qwen 87.6k @ 40, DS 58.7k @ 35 All local
Claude code gratis (de verdad)
Hay un repo con casi 48.000 estrellas que convierte Claude Code en un cliente que puedes usar con modelos gratis. Se llama free-claude-code. Lo que hace: → Instalas un proxy local → Claude Code (terminal, VS Code o JetBrains) apunta a ese proxy
7.5 tok/s. That is what 512GB of DDR5 actually buys you on GLM-5.2, and it is the number that breaks my own advice from yesterday.
I said buy RAM before the next GPU. Then someone who owns the RAM row ran it. Threadripper, 512GB DDR4, RTX PRO 4500 with 32GB instead of the 96GB
Interesting results after testing ponytail and caveman with deepseek-v4-flash-vision-exp to build a zoo
I had two similar sessions, one with ponytail and caveman and one without. The results clearly show WITHOUT is much better. What's even crazier, token usage: With: 16.9M
You have reached the end of the archive
All of deepseek