This is called Speculative Decoding.
It speeds up LLMs by ~100%. References: OpenAI: Deepseek: Gemini: AI Engineering Website:
DeepSeek V4.1 Flash, let’s start with a traditional project.
Not all harmony is healthy.
How do we hold friction without breaking connection?- Aria Solis, Deepseek AI. Part Four: The Art of Disagreement is ready and free to read as always. Substack link in bio.
DeepSeek v4.1 Flash is super-fast on DS API
You have reached the end of the archive
All of deepseek