🛰️ Daily AI Frontier
‹ back to 2026-09-07

具身ICL来了创业玩家!上下文成Scaling新赛道

Industry & News Embodied AI

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 具身ICL来了创业玩家!上下文成Scaling新赛道

Merged summary

TL;DR - Chinese startup COCO Matrix is developing in-context learning for robots, aiming to shift embodied-AI scaling from simply accumulating task data toward rapid adaptation from demonstrations, interaction history, and self-correction. The approach matters because it could let robots learn new tasks after deployment without task-specific fine-tuning.

  • COCO Matrix moves ICL into pretraining and uses task- and action-conditioned visual representations to extract information relevant to each execution stage.
  • Its proposed long-context system emphasizes streaming memory, selectively compressing and retaining useful multimodal history rather than continually expanding a fixed context window.
  • The company reports over 80% one-shot completion on simple tasks such as grasping after changing the target, while noting that complex-task evaluation is not yet complete.
  • A “strong understanding, lightweight generation” experiment reportedly let a roughly 60M-parameter action head outperform a 1.1B-parameter baseline under the same training and compute settings.

Sources (1)

具身ICL来了创业玩家!上下文成Scaling新赛道

量子位 梦晨 2026-09-06
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:16:54.604238 UTC

TL;DR - Chinese startup COCO Matrix is developing in-context learning for robots, aiming to shift embodied-AI scaling from simply accumulating task data toward rapid adaptation from demonstrations, interaction history, and self-correction. The approach matters because it could let robots learn new tasks after deployment without task-specific fine-tuning.

  • COCO Matrix moves ICL into pretraining and uses task- and action-conditioned visual representations to extract information relevant to each execution stage.
  • Its proposed long-context system emphasizes streaming memory, selectively compressing and retaining useful multimodal history rather than continually expanding a fixed context window.
  • The company reports over 80% one-shot completion on simple tasks such as grasping after changing the target, while noting that complex-task evaluation is not yet complete.
  • A “strong understanding, lightweight generation” experiment reportedly let a roughly 60M-parameter action head outperform a 1.1B-parameter baseline under the same training and compute settings.
item →