🛰️ Daily AI Frontier
‹ back to 2026-09-08

UniMate: One Unified Model to Animate Diverse Skeletons

Research Multimodal & Generative

Ranking

Overall 81
Content 90
Popularity 62

Observed public metrics from 1 member.

Representative image for UniMate: One Unified Model to Animate Diverse Skeletons

Merged summary

TL;DR - UniMate is a topology-aware diffusion transformer that generates text-guided motion for arbitrary rigged 3D skeletons without per-skeleton retraining or test-time optimization. It aims to remove topology-specific constraints from learned animation and enable zero-shot motion transfer across diverse asset types.

  • Encodes skeletal structure through graph-aware attention biases, graph-Laplacian spectral rotary embeddings, and a global rest-pose topology conditioner.
  • Trains on UniML3D, a curated dataset of 13,006 text-paired motion sequences spanning animals, articulated objects, and varied skeletal topologies.
  • Reportedly outperforms existing baselines in motion quality, generalization, and efficiency.
  • Supports zero-shot cross-topology transfer, motion in-betweening and expansion, and text-guided editing.

Sources (1)

UniMate: One Unified Model to Animate Diverse Skeletons

arXiv cs.CV Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai, Chaoyue Song, Adam Finkelstein, Szymon Rusinkiewicz 2026-09-04 arXiv:2609.05415
Public signals Hugging Face upvotes 16
Providers: Hugging Face · Upvotes 16 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:22:59.449213 UTC

TL;DR - UniMate is a topology-aware diffusion transformer that generates text-guided motion for arbitrary rigged 3D skeletons without per-skeleton retraining or test-time optimization. It aims to remove topology-specific constraints from learned animation and enable zero-shot motion transfer across diverse asset types.

  • Encodes skeletal structure through graph-aware attention biases, graph-Laplacian spectral rotary embeddings, and a global rest-pose topology conditioner.
  • Trains on UniML3D, a curated dataset of 13,006 text-paired motion sequences spanning animals, articulated objects, and varied skeletal topologies.
  • Reportedly outperforms existing baselines in motion quality, generalization, and efficiency.
  • Supports zero-shot cross-topology transfer, motion in-betweening and expansion, and text-guided editing.
item →