🛰️ Daily AI Frontier
‹ back to 2026-08-03

Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning…

Industry & News Multimodal & Generative

Ranking

Overall 64
Content 70
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning…

Merged summary

TL;DR - Thinking Machines released Inkling-Small, an open-weights multimodal MoE model (276B total / 12B active) that reportedly matches the larger Inkling at a quarter of its size, and NVIDIA is amplifying it with an NVFP4 checkpoint for fine-tuning on NeMo/DGX Station.

  • Mixture-of-Experts architecture: 276B total parameters with only 12B active per token, targeting Inkling-level quality at ~1/4 the size.
  • Natively reasons over text, images, and audio, with controllable/variable "thinking effort" as a user-tunable knob.
  • Full weights are open; an NVFP4-quantized checkpoint is published on Hugging Face for low-precision inference and fine-tuning on NVIDIA NeMo + DGX Station.
  • Available to fine-tune via Thinking Machines' Tinker platform and to try interactively in Tinker Playground; no benchmark numbers are given in the post beyond the "comparable performance" claim.

Sources (1)

Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning…

@NVIDIAAI 2026-07-30
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-02 14:29:37.677982 UTC

TL;DR - Thinking Machines released Inkling-Small, an open-weights multimodal MoE model (276B total / 12B active) that reportedly matches the larger Inkling at a quarter of its size, and NVIDIA is amplifying it with an NVFP4 checkpoint for fine-tuning on NeMo/DGX Station.

  • Mixture-of-Experts architecture: 276B total parameters with only 12B active per token, targeting Inkling-level quality at ~1/4 the size.
  • Natively reasons over text, images, and audio, with controllable/variable "thinking effort" as a user-tunable knob.
  • Full weights are open; an NVFP4-quantized checkpoint is published on Hugging Face for low-precision inference and fine-tuning on NVIDIA NeMo + DGX Station.
  • Available to fine-tune via Thinking Machines' Tinker platform and to try interactively in Tinker Playground; no benchmark numbers are given in the post beyond the "comparable performance" claim.
item →