The dispatch
Models Jul 15 · 3-min read · updated Jul 15

Thinking Machines drops Inkling: ~1T-param open multimodal model

An open-weight ~1T model that natively ingests text, image, and audio, with day-0 support in transformers, SGLang, and llama.cpp, plus routing on Vercel AI Gateway. An NVFP4 build (~600GB VRAM) or llama.cpp quants make it self-hostable; try HF Inference Providers first.

Source: Hugging Face blog

← All recent updates

Get the ones that matter, weekly.

Stay on top of what shipped in AI. Weekly update.

Unsubscribe anytime
~/subscribe $