RT by @_akhaliq: Baseten is now an official inference provider on @huggingface 🤗 Run Kimi K3…
TL;DR - Baseten has been added as an official inference provider on Hugging Face, letting users run hosted open-weight models directly from model pages or via an HF token. It matters because it lowers the friction of serving large open models without self-managed GPU infrastructure.
- Integration is surfaced natively on Hugging Face model pages, so inference can be triggered without separate Baseten account plumbing.
- Authentication works with an existing Hugging Face token, making Baseten usable from third-party harnesses and existing tooling.
- Named available models include Kimi K3, DeepSeek V4 Flash, and GLM-5.2 — all large open-weight LLMs.
- Content is a short announcement post; no latency, throughput, pricing, or benchmark details are provided beyond the linked HF blog.