RT by @NVIDIAAI: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support!…
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - NVIDIA is amplifying Thinking Machines' release of Inkling-Small, an open-weights multimodal MoE model that reportedly matches the larger Inkling at a quarter of the size, with Day 0 vLLM support optimized for NVIDIA Blackwell. It matters because a frontier-class open-weights model with same-day serving support lowers the barrier to deploying multimodal agentic systems.
- 276B total parameters with 12B active (Mixture-of-Experts), native text, image, and audio input, and a 1M-token context window.
- Architecture mirrors the larger Inkling, so it reuses the existing vLLM code path — enabling Day 0 support and Blackwell-targeted optimization.
- Positioned for agentic/tool-use systems, coding assistants, and RAG; features controllable reasoning effort.
- Full weights released; fine-tuning offered via Thinking Machines' Tinker, with text/image/audio chat in Tinker Playground. Note: performance parity with Inkling is a vendor claim, with no benchmark numbers given in this post.
Sources (1)
RT by @NVIDIAAI: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support!…
TL;DR - NVIDIA is amplifying Thinking Machines' release of Inkling-Small, an open-weights multimodal MoE model that reportedly matches the larger Inkling at a quarter of the size, with Day 0 vLLM support optimized for NVIDIA Blackwell. It matters because a frontier-class open-weights model with same-day serving support lowers the barrier to deploying multimodal agentic systems.
- 276B total parameters with 12B active (Mixture-of-Experts), native text, image, and audio input, and a 1M-token context window.
- Architecture mirrors the larger Inkling, so it reuses the existing vLLM code path — enabling Day 0 support and Blackwell-targeted optimization.
- Positioned for agentic/tool-use systems, coding assistants, and RAG; features controllable reasoning effort.
- Full weights released; fine-tuning offered via Thinking Machines' Tinker, with text/image/audio chat in Tinker Playground. Note: performance parity with Inkling is a vendor claim, with no benchmark numbers given in this post.