Like any distilled model, Ternary Bonsai 2 27B inherits refusals and censorship from its teacher model: Qwen 3.8 27B.
Fortunately, there’s a very simple and elegant way to reduce or perhaps even completely eliminate them. 🧵
Your US stock position still carries risk even after your broker closes.
So I built Tesrune for the Bitget_AI S2 Hackathon. Tesrune is a dark-hours AI trading desk for US equities. You add the stocks you're holding. While your broker is closed, Tesrune watches relevant news,
Talking about Local Benchmaxxing - here is my 2 year old Intel i9 64 GB machine with Nvidia RTX 4090 running Qwen 3.8 Flash next at 30 t/s.
Happily using opencode at my Macbook Pro using my Older Intel Machine as an API endpoint. Quant : AD-3.84bpw-IQ4_XS-M64
Found a sweet spot between quality and speed for Qwen 3.8 Flash on 2x3090s @ 3.5bpw, running some tests
Backend is exllamav3, running pretty sweet @ around 105 tk/s average (92 tk/s prose/reasoning, 105 tk/s coding and 120 tk/s file editing) and 1635 tk/s prefill
You have reached the end of the archive
All of qwen38