FRAMEWIREIndonesiaUpdated Aug 17Live wire
0:00 / 0:00

Stop paying API visual context taxes to the cloud cartel.

I just benched Qwen3.8 27B with vision on a single NVIDIA RTX 4090 (24GB). I fed it a massive 30,000 token payload (28.5k text + 1 hi res image) to see where the physical VRAM and throughput ceilings are with and without

AlokAug 1715
0:00 / 0:00

Qwen 3.8 27b

Prompts: give me an essay on ontologies now create a website that will explain this in threeejs animation Running locally: Qwen3.8-27B UD-Q8_K_XL 220K context 2× RTX 3090 Up to 42.26 tok/s decode (~39.3 average) Hermes Recipe: Thanks

lazybutaiAug 17
0:00 / 0:00

Qwen3.8-27B is really on the level of Opus 4.6 Max!

I just created this test with one simple prompt. Qwen ran for about one minute on my MacBook. Even faster than Claude. Left side is Opus 4.6 Max. Right side is Qwen 3.8. Very impressive for 18GB!

Tobias WupperfeldAug 173
0:00 / 0:00

I tested again Gemini 3.7 Flash vs Qwen 3.8 vs Grok 4.6 vs GLM 5.3

Same prompt: a family having lunch in 3D. Qwen 3.8 wins again for me on quality. but Gemini finished in under 2 mins, the others took 6+. Which one wins for you?

Tauhid IQAug 173