Running Fable 5.1 at max effort through the Artificial Analysis Intelligence Index cost $8,523.
The most expensive run on the board, 56% over Fable 5 at $5,455. The sticker didn't move: $10 in, $50 out, same tokenizer as Fable 5. So the difference is token volume. Effort is a
The lecture spends 15 minutes proving models can sample their way past Claude.
Then it shows why that number is fake on most tasks. coverage on SWE-bench keeps climbing to 1,000 samples. DeepSeek-V3 crosses Claude 3.5 and o1-preview. majority vote on the same draws flatlines
Thomas Wolf is fascinated by the constant revolution in open source AI.
Last year it was DeepSeek, and this year GLM 5.2 surprised everyone by nearing GPT-4's performance. Hugging Face sees these rapid advancements every few months.
Bedtime stories. English words.
Math homework. One robot friend for all three — powered by ChatGPT, Gemini, Mistral, Qwen & DeepSeek. Kid-safe filtering, parent review, no subscription. What would your kid ask it first? 🎒
You have reached the end of the archive
All of deepseek