Constantly amazed at what you can do at home with *significant* investment (but also not boat money).
My own agent/harness running entirely on my own hardware. It's the Qwen 3.8 Flash model running on one Spark, plus all the voice models running on a 3090. It's plugged into OMP,
Testing out Jev. I am putting Jev against Qwen 3.8 27B in a chess game.
So far I keep getting no winners. But Jev's speed is uncomparable.
I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX Spark.
Cruz shipped a custom ExLlamaV3 fork and a native engine path. I ran the same model on the same Spark: 102.6 max, 80 tok/s on code, 70 tok/s on prose. My previous benchmark and
There's been a lot of hype about Ternary Bonsai2 but fw demonstrations on how it actually compares to the full 16 bit version of Qwen 3.8 27B.
Since I have my own quantized version of that model thatI'm testing, ~12gb file size, I decided to throw Bonsai2 in there also to get a
You have reached the end of the archive
All of qwen38