HumanML3D official test split, selected GT caption protocol. Public rows report
semantic retrieval metrics, universal SMPL-22 TMR metrics, and joint-level physical quality metrics.
uTMR metrics use canonicalized SMPL-22 joints66 at 30 fps. Paper-only rows are kept
separate when no released checkpoint can be evaluated through the shared SMPL-22 bridge.
Current Public Snapshot
Updated from verified internal metric JSONs. PRISM 1.0 is the no-KT/no-KAFS baseline rerun with pad360/crop; PRISM KAFS cfg5 uses the epoch-12 new-VAE checkpoint.
Best normalized uTMR FID
-
Generated methods only
Best MS R@1
-
Generated methods only
Best uTMR R@3
-
Generated methods only
Lowest Foot Slide
-
Generated methods only
Method Comparison
BestSecondGT reference
Radar methods, up to four
Generated-method ranking
Normalized profile
100 is the best generated-method value on each axis; GT is excluded.
All-case motion comparison
4,042 selected-caption cases · lazy loaded
MS = MotionStreamer-272 evaluator. uTMR = universal SMPL-22 joints66 TMR evaluator. Only per-sample L2-normalized FID is displayed and ranked; historical raw-embedding FID values remain hidden until recomputed. Lower normalized FID and MM are better; higher R-Precision is better. Diversity is a reference statistic and is not ranked.
Slide, Float, and Jitter use the shared SMPL-22 joint-level physical metric implementation. PoseQ is the MBench NRDF pose-quality score. Dynamic is compared with the GT reference rather than minimized.
Paper benchmark rows reproduce metrics reported by the original papers when the released checkpoint is not available in the shared SMPL-22 evaluation protocol. They are not used for generated-method ranking.
Calibration rows quantify legacy SMPL-to-HML263 conversion loss and are not ranked as generation methods. uTMR sanity is tracked by the GT row in the semantic table.