REVIEW 5 cited by
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose TRAM, a two-stage method to reconstruct a human's global trajectory and motion from in-the-wild videos. TRAM robustifies SLAM to recover the camera motion in the presence of dynamic humans and uses the scene background to derive the motion scale. Using the recovered camera as a metric-scale reference frame, we introduce a video transformer model (VIMO) to regress the kinematic body motion of a human. By composing the two motions, we achieve accurate recovery of 3D humans in the world space, reducing global motion errors by a large margin from prior work. https://yufu-wang.github.io/tram4d/
Forward citations
Cited by 5 Pith papers
-
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
Sen-Cap estimates 3D human pose and global trajectory from flexible, uncalibrated combinations of LiDARs and cameras, setting state-of-the-art scores on Human-M3 and FreeMotion while tolerating point-cloud noise.
-
EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal
EgoHTR is a 55-sequence, 150k-frame egocentric 4D human-terrain dataset with a reconstruction pipeline, MoCap-validated benchmark, and perceptive locomotion policies deployed on a Unitree G1.
-
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
Motion-X++ provides 19.5M 3D whole-body pose annotations across 120.5K sequences with text, audio, video, and motion modalities.
-
Joint Optimization for 4D Human-Scene Reconstruction in the Wild
By jointly optimizing human motion, camera poses, and dense scene geometry with human-scene contact constraints, JOSH attains state-of-the-art global human motion and scene reconstruction from monocular web videos.
-
Reconstructing People, Places, and Cameras
HSfM jointly optimizes human meshes, dense scene pointmaps, and camera poses in a metric world frame, reducing world-frame human joint error from 3.5m to 1.0m on EgoHumans.
Discussion (0). Continue with ORCID to comment.