I set up an inference server with Kimi-K3-Q2_K, which supports CPU inference with DwarfStar Kai, and tried interacting on DeepSeek Harness.
The prompt corresponding to the personality has 9000 characters, but even in an environment with a prefill of 10 TPS, decoding and 3 TPS, we are able to return this response by combining speed-up elements such as an SSD cache strategy.
My current work environment consists of Visual Studio Code with the Cline code agent installed.
In Cline I can choose multiple models, and I am currently working with DeepSeek V4-Flash. info:
Want to build full apps with AI — for free — right inside VS Code?
In this tutorial, I'll show you how to set up Cline for DeepSeek V4 Flash completely free No API key, no billing setup, and no subscription required. FULL VIDEO:
What conclusions will be drawn when local model running, efficient fine-tuning, Agent access, and RAG search are combined into the same tool?
Discover an open source project: Unsloth It is currently the most powerful local LLM/diffusion model running and training framework. It provides Desktop App and Web UI, supports the latest models (Qwen3.8, Kimi K3, Gemma 4, DeepSeek-V4, FLUX, etc.), and the training speed is increased by 2 times and the video memory is reduced.
You have reached the end of the archive
All of deepseek