multimodal

4 articles

← All topics

DeepSeek ships V4-Flash-Vision-Exp — its first multimodal model, matching V4-Flash on text, adding image input
deepseekdeepseek-v4-flashvision-expmultimodalimage-input+2

DeepSeek's experimental vision model matches V4-Flash on text tasks and adds image input at the same pricing; new Files API lets you upload once and reuse across requests.

Meta releases Muse Glimmer 30B — an Apache-2.0 agentic multimodal model for 24GB consumer GPUs
metamuse-glimmeropen-weightsmultimodalagentic-ai+1

Meta's Muse Glimmer is a 30B Apache-2.0 multimodal model distilled from Muse Spark that fits a 24GB consumer GPU with 4-bit quantization. Architecture, benchmarks, and DFlash.

Moonshot AI releases Kimi K3 — a 2.8T open-weights MoE model with million-token context
moonshot-aikimikimi-k3open-weightsmoe+3

Kimi K3 is a 2.8T-parameter open-weights MoE model built on Kimi Delta Attention and Attention Residuals, with 1M-token context and native vision. Weights to follow by July 27.

Mistral releases Leanstral 1.5: 119B/6B-active Apache-2.0 prover saturates miniF2F and finds 5 unknown Rust bugs at $4 per problem
open-sourceopen-weightsformal-verificationlean-4mistral+18

Apache-2.0 119B/6B-active MoE prover from Mistral saturates miniF2F, hits SOTA on FATE-H (87%) and FATE-X (34%), and uncovers 5 previously unknown Rust bugs across 57 repos — all for about $4 per problem.