FRAMEWIREIndonesiaUpdated Aug 16Live wire
0:00 / 0:00

Gemini 3.7 Flash

Coding instructions for physical simulation of water droplets falling on water surface in both Python and HTML5, one correction request (installation of quantitative change bar) Thinking for less than a few seconds, spitting out code for about 1 second Running video on Google Colab When I tried the same prompt with Qwen 3.8 27B Q8_0, it consumed a lot of tokens and stopped midway.

hiro_pismoAug 151
0:00 / 0:00

Qwen 3.8-27B (Golden Shore) vs Opus 4.6 (Coastal World)

Same exact prompt, both were given only one shot to see what they would make.

DanielAug 151
0:00 / 0:00

Update on the free community Qwen 3.8-27B endpoint (was getting a bit too slow)

2x H200 replicas (+1 replica) now use speculative decoding (70 → 126 tok/s, measured) default thinking is now medium when not set (xhigh burns an enormous amount of tokens)

Victor MAug 1519
0:00 / 0:00

Been using Qwen 3.8 27B (Q4) locally on 64GB of VRAM.

Here is the verdict: SLOW 18 tps with ZERO system prompt to process and that degrades significantly with a harness system prompt and as the context window grows. RIP if you have to compact. I had it implement this PRD

Burke HollandAug 1556