Motion Reconstruction · HumanML3D

Paired autoencoder reconstruction quality across semantic, geometric, and physical views. Every method is evaluated through its declared native representation or an explicitly named common bridge.

HumanML3D test 4,042 motions SMPL-22 geometry Normalized reconstruction FID Physical diagnostics
Current public snapshotLoading verified results
Evaluated autoencoders
9
Complete geometry population for every method
Best reconstruction FID
Generated methods only
Lowest PA-MPJPE
Generated methods only
Geometry population
4,042
Paired HumanML3D test reconstructions

Method comparison

Semantic, geometry, and physical metrics share one generated-method view.

Best generatedSecond generatedGT reference

Reconstruction FID

Lower is better; GT is excluded.

Normalized profile

100 is the best generated value on each axis; GT is excluded.

All-case motion comparison

Inspect every paired HumanML3D reconstruction beside its GT motion in one synchronized 3D view.

4,042 cases · 10 views
GT + all nine benchmarked tokenizersSelected captions · synchronized playback

Complete results

GT is pinned first; lower is better for every ranked reconstruction and physical metric.

3 generated + GT
Semantic clips shorter than 60 frames are omitted only from learned reconstruction FID.Geometry metrics use all 4,042 motions.
Protocol details

Semantic reconstruction

rFID and paired embedding L2 use the MotionStreamer HumanML3D reconstruction embedding protocol.

Geometry reconstruction

MPJPE, PA-MPJPE, and MPJRE compare paired SMPL-22 geometry over the complete test population.

Physical diagnostics

Foot slide, float, penetration, and jitter use Motius diagnostics. MotionLCM declares its common motion135 bridge explicitly.