Glean 拾遗
Recent picks

2picks · chronological

07-30

From GPT2 to Kimi3: A 22,580x Scale-Up with Architectural Evolution

This worklog traces the architectural evolution from GPT-2 (2019) to KimiK3 (2026) with code snippets and diagrams. The central claim: progress is not just scale but innovations in state management and retrieval. It covers KV cache bottlenecks, linear attention's fixed-state trade-off, DeltaNet's precise overwriting via delta rule, Gated DeltaNet adding forgetting, KDA/Kimi Linear with per-channel gating, and finally KimiK3's hybrid of KDA and MLA layers, MoE, SiTU activation, and blockwise AttentionRes (AttnRes) for selective depth-wise residual access. Suitable for engineers interested in LLM internals.

x.com · 27 min · DeltaNet · Kimi K3 · Linear Attention
07-17

Kimi K3: Open 2.8T Frontier Model for Long-Horizon Coding and Knowledge Work

Moonshot AI releases Kimi K3, a 2.8T-parameter open model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), activating 16 out of 896 experts with a reported 2.5× scaling efficiency improvement over K2. It supports native vision and a 1M-token context window. While generally trailing top proprietary models like Claude Fable 5 and GPT 5.6 Sol, K3 achieves competitive scores on coding, knowledge work, and reasoning benchmarks. The post details case studies: GPU kernel optimization, a from-scratch Triton-like compiler (MiniTriton), 3D open-world game development, autonomous chip design (48-hour run), and rapid scientific research reproduction. K3 is available now via Kimi.com, Kimi Work, Kimi Code, and API; full weights open-sourced by July 27, 2026. Recommended for AI engineers, agent developers, and researchers needing long-horizon agentic capabilities.

www.kimi.com · 19 min · Agent Engineering · Coding · Kimi K3