You asked how to run ego (lite) with a local model.
So we pulled Qwen 3.8 with Ollama, hooked it up to OpenCode, and pointed it at Airbnb Tokyo. A CSV with listings, prices, and ratings comes back. No API key. No cloud.
I'm not a dev nor an engineer guy.
Just an enthusiast with some compute. It has the overthinking flaw, but it made the best pagoda in my pc to the day 💪 Im now downloading the swift 1.5 for less thinking You should test it. Is fast and it runs at the same speed as qwen 3.8 27b
Strata qwen 3.8 flash on a 3090 card and I'll add a Pic of sys specs.
Local llm. Code focus. Only like 35gig used included is the OS for the machine, Ubuntu. So this can run while I am running over things for sure, easily. Aug is around 75 to 85tps tokens per sec. Strata is a
Testing Qwen 3.8 Flash IQ3_S GSQ RCO on 2× RTX 5060 Ti 16GB with Strata, running a 256K context window.
I’m testing decode speed, prompt processing, expert-cache behavior, and how well Strata’s multi-GPU layer split scales across 32GB total VRAM. Results coming soon!
You have reached the end of the archive
All of qwen38