🛰️ Daily AI Frontier
‹ back to 2026-08-02

RT by @NVIDIAAI: Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains. At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks. Run Inkling-small with SGLang, and customize it with Miles.

Industry & News Multimodal & Generative

Ranking

Overall 78
Content 90
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for RT by @NVIDIAAI: Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains. At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks. Run Inkling-small with SGLang, and customize it with Miles.

Merged summary

TL;DR - Thinking Machines released Inkling-Small, an open-weight multimodal MoE model claiming comparable performance to Inkling at one-quarter the size. Its smaller active footprint makes inference and reinforcement-learning customization more practical.

  • 276B total parameters with 12B active, supporting controllable reasoning across text, image, and audio.
  • Full weights are available, with LoRA and full-parameter training supported through Tinker and Miles.
  • SGLang reportedly reaches 648 tokens/s with DSpark and 288 tokens/s without it on eight NVIDIA B200 GPUs using NVFP4 and TP8.
  • The release claims parity with the larger Inkling model and better results on some unspecified benchmarks.

Sources (1)

RT by @NVIDIAAI: Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains. At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks. Run Inkling-small with SGLang, and customize it with Miles.

@lmsysorg 2026-07-30
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-01 14:18:36.252836 UTC

TL;DR - Thinking Machines released Inkling-Small, an open-weight multimodal MoE model claiming comparable performance to Inkling at one-quarter the size. Its smaller active footprint makes inference and reinforcement-learning customization more practical.

  • 276B total parameters with 12B active, supporting controllable reasoning across text, image, and audio.
  • Full weights are available, with LoRA and full-parameter training supported through Tinker and Miles.
  • SGLang reportedly reaches 648 tokens/s with DSpark and 288 tokens/s without it on eight NVIDIA B200 GPUs using NVFP4 and TP8.
  • The release claims parity with the larger Inkling model and better results on some unspecified benchmarks.
item →