Frontier LLM Coding Benchmark - Double Pendulum Challenge: Fable 5.1 x GLM 5.3 x GPT-6 Astra x Qwen 3.8
Effort: Max & Ultra Task: In this task, you create a cinematic, navigation-oriented Canvas 2D that simulates the full nonlinear double pendulum physics with real RK4 integration from the model.
Ran at 30 toks for most of the time.
Prefill was insane. Love that I can just have Qwen 3.8 Flash Next at home. It definitely runs a bit hot. Might get a tiny fan. Follow MiaAI_lab the GOAT
I’m honestly shocked.
Qwen 3.8 Flash on one home Spark. First three one-shots: pirate ship, Silly String, NES racer. Equal to — if not better than — my Gemini 3.8 Flash clips from the other day. Open-weight. MiaAI_lab recipe, optimized for a single Spark. You’d think a
Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, and Qwen 3.8 Max
You have reached the end of the archive
All of qwen38