RT by @_akhaliq: Baseten is now an official inference provider on @huggingface 🤗 Run Kimi K3…
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR — Baseten has joined Hugging Face's inference provider program as an official provider, letting users run hosted open-weight models directly from HF model pages. It matters because it further reduces friction between discovering a large open model on the Hub and serving it in production without self-managed GPU infrastructure.
- Baseten's serving stack is exposed behind the Hub's unified model-page inference UI, so inference can be triggered natively without separate Baseten account plumbing.
- Authentication uses an existing Hugging Face token, meaning any HF-compatible client or harness can route requests to Baseten without separate provider credentials or SDK changes.
- Named supported models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 — all large open-weight LLMs typically too heavy for casual self-hosting.
- Content is announcement-level only (a short promo post plus a link to huggingface.co/blog/baseten); no latency, throughput, pricing, or benchmark figures are provided.
Both sources are near-identical announcements; @huggingface frames it as an addition to its provider program, while @_akhaliq emphasizes the end-user workflow of running models from model pages.
Sources (2)
RT by @_akhaliq: Baseten is now an official inference provider on @huggingface 🤗 Run Kimi K3…
TL;DR - Baseten has been added as an official inference provider on Hugging Face, letting users run hosted open-weight models directly from model pages or via an HF token. It matters because it lowers the friction of serving large open models without self-managed GPU infrastructure.
- Integration is surfaced natively on Hugging Face model pages, so inference can be triggered without separate Baseten account plumbing.
- Authentication works with an existing Hugging Face token, making Baseten usable from third-party harnesses and existing tooling.
- Named available models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 — all large open-weight LLMs.
- Content is a short announcement post; no latency, throughput, pricing, or benchmark details are provided beyond the linked HF blog.
RT by @huggingface: Baseten is now an official inference provider on @huggingface 🤗 Run Kimi K3…
TL;DR - Hugging Face announced Baseten as an official inference provider, letting users run hosted open-weight models like Kimi K3, DeepSeek V4 Flash, and GLM-5.2 directly from HF model pages. It matters because it further reduces friction between discovering an open model on the Hub and serving it in production without self-hosting GPUs.
- Baseten joins Hugging Face's inference provider program, exposing its serving stack behind the Hub's unified model-page inference UI.
- Named supported models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 — all large open-weight LLMs typically too heavy for casual self-hosting.
- Authentication is via an existing Hugging Face token, so any HF-compatible client/harness can route requests to Baseten without separate provider credentials or SDK changes.
- Announcement-level content only (a promo post plus link to huggingface.co/blog/baseten); no latency, throughput, pricing, or benchmark figures were provided.