FRAMEWIREIndonesiaUpdated Sep 11Live wire
0:00 / 0:00

T1 is a new 122B-parameter MoE agent that runs real Linux shell tasks for 300+ steps—trained with RL, not just next-token prediction.

The secret: dense per-assertion rewards, token-faithful trajectories (TITO), and expert-mask replay (R³) to finally stabilize massive MoE RL. On

yesnoerrorSep 115
0:00 / 0:00

DeepSeek V4.1 Flash is here, caching requirements have been cut to a quarter, will Agent costs plummet?

A giant with 552B parameters, the input only activates 8B and the output is 16B. The efficiency of this asymmetric structure is incredible. The benchmark test comprehensively surpasses the old flagship V4 Pro. The official plans to take off the old model directly.

奶牛叔Sep 112