🛰️ Daily AI Frontier
‹ back to 2026-08-03

RT by @huggingface: Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total…

Industry & News LLMs & Foundation Models

Ranking

Overall 64
Content 70
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for RT by @huggingface: Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total…

Merged summary

TL;DR - Thinking Machines released Inkling Small, a sparse Mixture-of-Experts model (12B active / 276B total params) shipped in NVFP4 format, with Hugging Face publishing a companion blog on benchmarks, performance, and deployment. It matters because a smaller-active-parameter model reportedly beats the larger Inkling on coding, suggesting gains from sparsity and low-precision formats rather than raw scale.

  • Architecture is sparse MoE: 276B total parameters with only 12B active per token, targeting large-model quality at small-model inference cost.
  • Distributed in NVFP4 (NVIDIA 4-bit floating point), a quantized format aimed at efficient serving on modern NVIDIA hardware.
  • Claimed to outperform the larger Inkling model on coding tasks — a counter-intuitive result, though no specific benchmark numbers are given in this post.
  • Weights are available via a Hugging Face collection, with a HF blog covering benchmarks and deployment; details beyond the announcement are thin here and would need the linked blog to verify.

Sources (1)

RT by @huggingface: Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total…

@mervenoyann 2026-07-30
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-02 14:29:37.882107 UTC

TL;DR - Thinking Machines released Inkling Small, a sparse Mixture-of-Experts model (12B active / 276B total params) shipped in NVFP4 format, with Hugging Face publishing a companion blog on benchmarks, performance, and deployment. It matters because a smaller-active-parameter model reportedly beats the larger Inkling on coding, suggesting gains from sparsity and low-precision formats rather than raw scale.

  • Architecture is sparse MoE: 276B total parameters with only 12B active per token, targeting large-model quality at small-model inference cost.
  • Distributed in NVFP4 (NVIDIA 4-bit floating point), a quantized format aimed at efficient serving on modern NVIDIA hardware.
  • Claimed to outperform the larger Inkling model on coding tasks — a counter-intuitive result, though no specific benchmark numbers are given in this post.
  • Weights are available via a Hugging Face collection, with a HF blog covering benchmarks and deployment; details beyond the announcement are thin here and would need the linked blog to verify.
item →