RT by @huggingface: Baseten is now an official inference provider on @huggingface 🤗 Run Kimi K3…
TL;DR - Hugging Face announced Baseten as an official inference provider, letting users run hosted open-weight models like Kimi K3, DeepSeek V4 Flash, and GLM-5.2 directly from HF model pages. It matters because it further reduces friction between discovering an open model on the Hub and serving it in production without self-hosting GPUs.
- Baseten joins Hugging Face's inference provider program, exposing its serving stack behind the Hub's unified model-page inference UI.
- Named supported models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 — all large open-weight LLMs typically too heavy for casual self-hosting.
- Authentication is via an existing Hugging Face token, so any HF-compatible client/harness can route requests to Baseten without separate provider credentials or SDK changes.
- Announcement-level content only (a promo post plus link to huggingface.co/blog/baseten); no latency, throughput, pricing, or benchmark figures were provided.