FRAMEWIREIndonesiaUpdated Aug 17Live wire
0:00 / 0:00

This is how Qwen 3.8 27B looks at around 30 tok/s in mtplx

I will try a smaller quant cause it's 8bit and it grows to around 35 gb in ram usage

Chris WAug 171
0:00 / 0:00

Stop paying API visual context taxes to the cloud cartel.

I just benched Qwen3.8 27B with vision on a single NVIDIA RTX 4090 (24GB). I fed it a massive 30,000 token payload (28.5k text + 1 hi res image) to see where the physical VRAM and throughput ceilings are with and without

AlokAug 1715
0:00 / 0:00

Qwen 3.8 27b

Prompts: give me an essay on ontologies now create a website that will explain this in threeejs animation Running locally: Qwen3.8-27B UD-Q8_K_XL 220K context 2× RTX 3090 Up to 42.26 tok/s decode (~39.3 average) Hermes Recipe: Thanks

lazybutaiAug 17
0:00 / 0:00

Qwen3.8-27B is really on the level of Opus 4.6 Max!

I just created this test with one simple prompt. Qwen ran for about one minute on my MacBook. Even faster than Claude. Left side is Opus 4.6 Max. Right side is Qwen 3.8. Very impressive for 18GB!

Tobias WupperfeldAug 173