moe

4 articles

← All topics

Moonshot AI releases Kimi K3 — a 2.8T open-weights MoE model with million-token context
moonshot-aikimikimi-k3open-weightsmoe+3

Kimi K3 is a 2.8T-parameter open-weights MoE model built on Kimi Delta Attention and Attention Residuals, with 1M-token context and native vision. Weights to follow by July 27.

Mistral opens Leanstral 1.5: 6B-active Apache-2.0 Lean 4 prover
mistralleanstralproof-engineeringformal-verificationlean-4+5

Apache-2.0 Leanstral 1.5 (119B/6B MoE) saturates miniF2F, hits 587/672 PutnamBench, sets SOTA on FATE-H and FATE-X, and finds 5 unknown bugs in 57 Rust repos — at ~$4 per problem.

Mistral releases Leanstral 1.5: 119B/6B-active Apache-2.0 prover saturates miniF2F and finds 5 unknown Rust bugs at $4 per problem
open-sourceopen-weightsformal-verificationlean-4mistral+18

Apache-2.0 119B/6B-active MoE prover from Mistral saturates miniF2F, hits SOTA on FATE-H (87%) and FATE-X (34%), and uncovers 5 previously unknown Rust bugs across 57 repos — all for about $4 per problem.

DiffusionGemma: 1,000+ tokens/sec open-weights text gen
google-deepmindgemmadiffusiongemmatext-diffusionopen-weights+10

Google DeepMind's June 10, 2026 release: a 26B/3.8B text-diffusion model denoising 256 tokens in parallel; ~4x faster than AR on a single H100, Apache 2.0, explicitly experimental.