🛰️ Daily AI Frontier
‹ back to 2026-08-26

硅谷今日最热具身模型!不用后训练,看一遍就学会

量子位 Embodied AI henry 2026-08-26
Representative image for 硅谷今日最热具身模型!不用后训练,看一遍就学会

TL;DR - Skild AI released S1, a robot foundation model that learns new, long-horizon tasks from a single human demonstration video without fine-tuning or post-training. Its reported results suggest in-context learning could substantially reduce the real-world data cost of teaching robots new skills.

  • S1 uses video demonstrations as prompts while keeping model weights fixed, supporting previously unseen tasks lasting up to 10 minutes.
  • Demonstrations include making pancakes and coffee, repotting plants, and assembling equipment, with multi-step execution and recovery from errors.
  • On out-of-distribution tasks, S1 reportedly achieved 66% success versus 9% for a language-prompted vision-language-action model.
  • A conventionally post-trained model required about 380 demonstrations to match S1’s one-shot 66% result, though 2,000 demonstrations raised it to 86%.

view merged work →