I'm optimizing tensor parallel execution across two MacBook M5 Max systems using RDMA.
This is how fast DeepSeek v4 Flash Exp Vision can run in MXFP4 (full precision) right now *without* DSpark, so this is the min speed you get in practice.
There is a repository on GitHub called "free-claude-code" with over 49,000 stars.
However, if you accept the name as it is, it's easy to end up with something that's not what you expected, so I tried to organize the system. What this project does is "Claude
This one is one I've been working with locally with DeepSeek-V4-Flash-Vision-Exp but with no spec.
Kimi k3, deepseek v4, gemma 4, glm 5.3-flash, and qwen 3.8 all run on your own machine through unsloth
A free desktop app built by the han brothers that makes $200/mo chatgpt pro and $200/mo claude max optional one install and any open source model runs behind claude code,
You have reached the end of the archive
All of deepseek