Text-to-Motion Leaderboard

HumanML3D official test split, selected GT caption protocol. Public rows report semantic retrieval metrics, universal SMPL-22 TMR metrics, and joint-level physical quality metrics. uTMR metrics use canonicalized SMPL-22 joints66 at 30 fps. Paper-only rows are kept separate when no released checkpoint can be evaluated through the shared SMPL-22 bridge.

Task: Text-to-Motion Dataset: HumanML3D official test MS samples: 4042 uTMR samples: 4034 Captions: selected full-clip GT
Current Public Snapshot Updated from verified internal metric JSONs. PRISM 1.0 is the no-KT/no-KAFS baseline rerun with pad360/crop; PRISM KAFS cfg5 uses the epoch-12 new-VAE checkpoint.
Best normalized uTMR FID
-
Generated methods only
Best MS R@1
-
Generated methods only
Best uTMR R@3
-
Generated methods only
Lowest Foot Slide
-
Generated methods only

Method Comparison

Best Second GT reference
Radar methods, up to four

Generated-method ranking

Normalized profile

100 is the best generated-method value on each axis; GT is excluded.

All-case motion comparison

4,042 selected-caption cases · lazy loaded

MS = MotionStreamer-272 evaluator. uTMR = universal SMPL-22 joints66 TMR evaluator. Only per-sample L2-normalized FID is displayed and ranked; historical raw-embedding FID values remain hidden until recomputed. Lower normalized FID and MM are better; higher R-Precision is better. Diversity is a reference statistic and is not ranked.