Like any distilled model, Ternary Bonsai 2 27B inherits refusals and censorship from its teacher model: Qwen 3.8 27B.
Fortunately, there’s a very simple and elegant way to reduce or perhaps even completely eliminate them. 🧵
Fortunately, there’s a very simple and elegant way to reduce or perhaps even completely eliminate them. 🧵
So I built Tesrune for the Bitget_AI S2 Hackathon. Tesrune is a dark-hours AI trading desk for US equities. You add the stocks you're holding.
Happily using opencode at my Macbook Pro using my Older Intel Machine as an API endpoint. Quant : AD-3.84bpw-IQ4_XS-M64
Backend is exllamav3, running pretty sweet @ around 105 tk/s average (92 tk/s prose/reasoning, 105 tk/s coding and 120 tk/s file editing) and 1635…
All FREE: OpenRouter: DeepSeek paths, Qwen 3.8 27B, GLM 5.2, Ling 3.0 Flash VL, Laguna S 2.1, Nemotron 3.5 Lightning / 3 Ultra: OpenCode Zen:
Want the SOP? DM me. 💬
Qwen 3.8 Live Translate from Alibaba_Qwen: about 2.3s average lag (down from 2.8s), 60 input languages, spoken output in 29.
It sees the screen and hears every word at once. Not a transcript. Not screenshots. The actual video.
Qwen 3.8 Omni Flash from Alibaba_Qwen takes text, image, audio, and video in. Text out. About 1M context.
This is Qwen 3.8-27B running on an M5 Max building a 3D racing game end to end.
Alibaba just dropped Qwen 3.8-Omni-Flash. Most AI reads text. Your business runs on calls, videos, and recordings.
Hmm, well, maybe this could protect my car. Strong opener. 😂
Qwen 3.8 27B ran mostly unsupervised, with no human-written CUDA and only ~12 nudges across the entire experiment.
AND its FREE and OpenSource?? Real time demo using Qwen 3.8 27b, hosted locally.
This is blowing my mind honestly This is Unreal Tournament 99 on two real Windows 98 PCs. Each player is driven by a language model.
Anyone who can place a headline in front of it is writing into the context of the thing that places your orders.
🌀 I gave 4 fast models the same Currents cover. Same prompt. One shot via aimlapi.
My own agent/harness running entirely on my own hardware. It's the Qwen 3.8 Flash model running on one Spark, plus all the voice models running on a…