AI Matters
Why it matters
Thinking Machines Lab releases Inkling-Small multimodal model
MarkTechPost Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
Thinking Machines Lab released Inkling-Small, a 12-billion active parameter mixture-of-experts model designed to run on a single Nvidia B300 GPU.
Why it matters
This open-weights model delivers high-end multimodal performance while being small enough to run efficiently on a single graphics card.
Latest brief · Aug 3, 2026