FRAMEWIREIndonesiaUpdated Aug 14Live wire
0:00 / 0:00

DeepSeek V4 Flash succeeded in running locally on AMD Strix Halo (Ryzen AI Max+ 395 / 128GB)

The chat UI is about 15 t/s, and when I have a conversation it takes more than a minute (👇 that part is cut in the video) Lightweight Atomic Agent when using agents

いにしえ@AI Director / Creator / Engineer|Will OldgramAug 14
0:00 / 0:00

In actual testing, GLM-5.3 outperformed DeepSeek V4 Pro. This generation of GLM model has very unique capabilities and is really worth using.

It feels like every model factory is beginning to find its own unique path.

塔斯海TasihiAug 141
0:00 / 0:00

DeepSeek-SSD is on GitHub

An SSD streaming engine to run DeepSeek V4 Flash on Windows

ManyGlueAug 14
0:00 / 0:00

Is DeepSeek V4 Pro overturned? Many people started to curse.

But when I read DeepSeek’s technical report today, I found out why the video memory didn’t explode when 1 million tokens were inserted into the context. This report has a total of 58 pages, and I only focused on two numbers: the amount of single-token inference calculations and the KV cache. When V4 Pro has 1 million tokens, the calculation amount of single token inference is about 27% of that of V3, KV…

左海有个洲Aug 142