🛰️ Daily AI Frontier
‹ back to 2026-08-03

RT by @NVIDIAAI: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support!…

Industry & News LLMs & Foundation Models

Ranking

Overall 64
Content 70
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for RT by @NVIDIAAI: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support!…

Merged summary

TL;DR - NVIDIA is amplifying Thinking Machines' release of Inkling-Small, an open-weights multimodal MoE model that reportedly matches the larger Inkling at a quarter of the size, with Day 0 vLLM support optimized for NVIDIA Blackwell. It matters because a frontier-class open-weights model with same-day serving support lowers the barrier to deploying multimodal agentic systems.

  • 276B total parameters with 12B active (Mixture-of-Experts), native text, image, and audio input, and a 1M-token context window.
  • Architecture mirrors the larger Inkling, so it reuses the existing vLLM code path — enabling Day 0 support and Blackwell-targeted optimization.
  • Positioned for agentic/tool-use systems, coding assistants, and RAG; features controllable reasoning effort.
  • Full weights released; fine-tuning offered via Thinking Machines' Tinker, with text/image/audio chat in Tinker Playground. Note: performance parity with Inkling is a vendor claim, with no benchmark numbers given in this post.

Sources (1)

RT by @NVIDIAAI: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support!…

@vllm_project 2026-07-30
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-02 14:29:38.085705 UTC

TL;DR - NVIDIA is amplifying Thinking Machines' release of Inkling-Small, an open-weights multimodal MoE model that reportedly matches the larger Inkling at a quarter of the size, with Day 0 vLLM support optimized for NVIDIA Blackwell. It matters because a frontier-class open-weights model with same-day serving support lowers the barrier to deploying multimodal agentic systems.

  • 276B total parameters with 12B active (Mixture-of-Experts), native text, image, and audio input, and a 1M-token context window.
  • Architecture mirrors the larger Inkling, so it reuses the existing vLLM code path — enabling Day 0 support and Blackwell-targeted optimization.
  • Positioned for agentic/tool-use systems, coding assistants, and RAG; features controllable reasoning effort.
  • Full weights released; fine-tuning offered via Thinking Machines' Tinker, with text/image/audio chat in Tinker Playground. Note: performance parity with Inkling is a vendor claim, with no benchmark numbers given in this post.
item →