Pith. sign in

Paper Citation Record · LEDGER

Learning from Massive Human Videos for Universal Humanoid Pose Control

As of 22 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 12 inbound Pith citation observations for arXiv:2412.14172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14172 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:29:46.938716Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:11:12.977121Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T18:15:21.360726Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2f564c5-9444-417b-a4b5-0ed7c5570e54 · outbound

This paper cites Human- to-robot imitation in the wild.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human- to-robot imitation in the wild

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.262409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.536471Z digest=sha256:8909b477f691bf805c1f9cf7d223f0a69bdb2ba0bcaed822730dd22d7a8a6ef9

Observation 5a7f4b1a-5685-4bb5-b74b-1b22ade4404a · outbound

This paper cites Affordances from human videos as a versa- tile representation for robotics.

Learning from Massive Human Videos for Universal Humanoid Pose Control Affordances from human videos as a versa- tile representation for robotics

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.246202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.542027Z digest=sha256:c6bf7f581ca5952808274c12f971ea92b98c60afeb4a820439af5400a78eba0a

Observation 308e47e4-a51b-4d4d-ad9b-fa30e79de5d4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Learning from Massive Human Videos for Universal Humanoid Pose Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.547914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.547914Z digest=sha256:4f2cb086532bbc5947dd2fa32b2e7a4645a91570fe80924899995621d55f1217

Observation 1bb15ec1-a7db-4d7e-b6d8-f1ff694044a6 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.230438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.553333Z digest=sha256:404392803cbab849f56c6d714152778ca8808a87eaf4282e7e57cc5b135ecfc0

Observation 4d620fe6-aca3-49fb-8375-cddc6d5ba51a · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

Learning from Massive Human Videos for Universal Humanoid Pose Control Rt-1: Robotics transformer for real-world control at scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.214657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.558290Z digest=sha256:a5473747ce88e20d005fd8351b5d1959a3bcd67f8765a5ee1ece875b0783ce6b

Observation 60a506c3-a396-40f4-8831-0320133b3d99 · outbound

This paper cites Humman: Multi-modal 4d human dataset for ver- satile sensing and modeling.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humman: Multi-modal 4d human dataset for ver- satile sensing and modeling

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.198665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.564057Z digest=sha256:2f8b6867b8c8b0301c4699d2822e50af6d08090f1dde019a7882367ed4567ecd

Observation 2b2fa1f7-a3f5-4c9b-903d-acf8cad4804e · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Learning from Massive Human Videos for Universal Humanoid Pose Control A Short Note on the Kinetics-700 Human Action Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.570427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.570427Z digest=sha256:5ce93188e13ad04c6994cf7935c1ee1ec111e0b071f6b4db7667727667728d70

Observation 4b26f8c8-5f8c-477a-a893-9f635a6b3a40 · outbound

This paper cites Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.576024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.576024Z digest=sha256:11a1c8cdfe188b1bf4a52a714a63b98f9da4eb120643390f71503f7be4fd4385

Observation 57989dae-2787-40b3-9e70-b74f1bcb509c · outbound

This paper cites Expressive Whole-Body Control for Humanoid Robots.

Learning from Massive Human Videos for Universal Humanoid Pose Control Expressive Whole-Body Control for Humanoid Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.581588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.581588Z digest=sha256:0f1e8e03f051db8e8784d56b47450def5dd0084ec482ac14f76f99520b232f66

Observation 3bffc8b4-0ebc-48cc-ac6c-591ad609275b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Learning from Massive Human Videos for Universal Humanoid Pose Control VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.587377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.587377Z digest=sha256:cd72e95967a3cf5c19ca72ff5293b67b03068654662fe10953d54b65e0cc7372

Observation 278ff0fc-6389-48b3-a6f8-36114dbc7bc7 · outbound

This paper cites Haa500: Human-centric atomic action dataset with curated videos.

Learning from Massive Human Videos for Universal Humanoid Pose Control Haa500: Human-centric atomic action dataset with curated videos

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.182463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.592997Z digest=sha256:c64c43101fd80479eda9e6630882edb137d97f3d5662852e74af744993e5da89

Observation eb06418c-c5e0-478a-8ba4-05098bb59e37 · outbound

This paper cites Video language plan- ning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Video language plan- ning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.166929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.598049Z digest=sha256:b72573d5311d08721672e03e60f352fb96c27183b23c268480f8d4e80bf7b017

Observation b922c7b1-30cd-45c7-af6e-d7455ed72025 · outbound

This paper cites HumanPlus: Humanoid Shadowing and Imitation from Humans.

Learning from Massive Human Videos for Universal Humanoid Pose Control HumanPlus: Humanoid Shadowing and Imitation from Humans

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.602989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.602989Z digest=sha256:50d7970b92396b4ca3306398630ce151df5e26543ded49c34bb2cec7853f680c

Observation e01ab8e4-c80e-4fc0-bb86-92f286ee1c10 · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humans in 4d: Re- constructing and tracking humans with transformers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.150060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.608188Z digest=sha256:e2e1117ddacc16afc53f73caf62a29cc7033025a1fd489818a9395e462079fb2

Observation 37186e09-e50f-4a07-ba78-b5e2a85bf85e · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Learning from Massive Human Videos for Universal Humanoid Pose Control Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.119189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.613033Z digest=sha256:e64db3e2402415a992f4f91a6f40dcc830e962a63db3c25b4e16959833f9eafe

Observation 5587e8bb-166b-4092-a9f3-0626e88f4cee · outbound

This paper cites Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.617363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.617363Z digest=sha256:b895fcede7a8e4e1554293887682c7a08ebb99b2b46b489225280e1ed7dcc980

Observation fe179203-9085-403c-a846-227519eeb38e · outbound

This paper cites Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.622418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.622418Z digest=sha256:4f82ef617818e9684d766d2523f917180be6bcef73743d917692a64d13c12206

Observation cac6edd0-aae1-4390-a800-6454f9c51af6 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating diverse and natural 3d human motions from text

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.091669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.627544Z digest=sha256:4acb5e9bba70397e4ecfbaa8c6bdd1f90ff3e8b4568e2389b2bfeca28b63f45a

Observation 3c4be4ac-4728-433e-b851-35550912fb76 · outbound

This paper cites OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.633592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.633592Z digest=sha256:93ba26732e5b2de8ff8c31a6f0251fc94622ebfa1468332cc1130d01ba6055dc

Observation 4f5a56bb-33a0-4a9e-8214-bd25fdcb3a02 · outbound

This paper cites Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.639956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.639956Z digest=sha256:515b063fba9cde2a831a198b35658f0d3e018beef6f03febb2bcd6df7668538c

Observation 44e1aba8-3b7c-4a90-b374-45faeffde534 · outbound

This paper cites HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots.

Learning from Massive Human Videos for Universal Humanoid Pose Control HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.645736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.645736Z digest=sha256:350a342d7c65f53a4790e26532e8caefeeed8e5f6778040915dadb4d32439782

Observation 956b06a8-b3a9-4e95-b742-f724b9ffae7b · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motiongpt: Human motion as a foreign language

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.076822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.651226Z digest=sha256:43c8164ed577a050d6a9b6d08fbaf3f423f2de27bd76f7bea199d1220d445d04

Observation 32bab377-2f38-4602-9fb6-ca890e9b9468 · outbound

This paper cites Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions.

Learning from Massive Human Videos for Universal Humanoid Pose Control Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.656230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.656230Z digest=sha256:6ed3fc628955254c5d29ca350386c6e94cec4e07fb731875f632c6b8adf922d6

Observation 908142aa-3004-4d2a-9cfe-ee310cdb2244 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Learning from Massive Human Videos for Universal Humanoid Pose Control OpenVLA: An Open-Source Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.662385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.662385Z digest=sha256:6c9676dd0f299bc16781dc83dab1d12ea80f6f0e1efa22a16cab78a8b209a79b

Observation c0fa429e-0150-42d4-86ea-4eb0665ed30b · outbound

This paper cites Adam: A method for stochastic opti- mization.

Learning from Massive Human Videos for Universal Humanoid Pose Control Adam: A method for stochastic opti- mization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.061524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.668329Z digest=sha256:bdcd5f7d4a6fe5d4df10bbbe16c46cf09b6cc0c983805e10e7886f6c2f16e782

Observation c40744b4-17e1-412a-94e0-1c061b45e7d1 · outbound

This paper cites Segment any- thing.

Learning from Massive Human Videos for Universal Humanoid Pose Control Segment any- thing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.045121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.673251Z digest=sha256:30056f84be69703d87afc384633931230621ae88494b01841c4f71a3e55d4121

Observation 217d7b6c-05cf-4931-a882-03cf9945ddcb · outbound

This paper cites Vibe: Video inference for human body pose and shape estimation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Vibe: Video inference for human body pose and shape estimation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.030750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.678183Z digest=sha256:2e2975f975052f56e213e30497d6ecdc472fe5a43f02d24a0ab36996ca3d25e7

Observation ca51d2b7-a4ec-4c68-bf3b-15e1f43cf470 · outbound

This paper cites RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation.

Learning from Massive Human Videos for Universal Humanoid Pose Control RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.682966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.682966Z digest=sha256:c108c0dfdf46927ddb5a4245c38c7b3604e46ea9f6080babe79218d7aa18ba4b

Observation 8a85e45e-db21-4b27-8c02-3ada4aa16fff · outbound

This paper cites OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation.

Learning from Massive Human Videos for Universal Humanoid Pose Control OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.687823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.687823Z digest=sha256:222878acc0f708c0ecea6810794758ff2ba0306987d338e24ccebffb53ca8dad

Observation f274362c-2130-42e9-8c26-430eff8810e6 · outbound

This paper cites Robust and versatile bipedal jumping control through reinforcement learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Robust and versatile bipedal jumping control through reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:48.015089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.692355Z digest=sha256:4481b1eb75e83643d670a5e285c5a9d7b0d4ff4d14b5f36851de964fb7c9c44d

Observation 3603b37d-41b9-4728-a233-10d76baf0975 · outbound

This paper cites Intergen: Diffusion-based multi-human motion genera- tion under complex interactions.

Learning from Massive Human Videos for Universal Humanoid Pose Control Intergen: Diffusion-based multi-human motion genera- tion under complex interactions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.998739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.696590Z digest=sha256:4b5d8d32705f37d3f5e00949a039b8076ad6451e7a6dcbf3031f0ae5467f38c1

Observation f596ef4d-21e1-4a82-9deb-12573753a192 · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motion-x: A large- scale 3d expressive whole-body human motion dataset

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.982370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.702254Z digest=sha256:979f0e48f0f142c84cd0d5a100e62ddad8d04fc1e6032cfa279f32de0a59e3af

Observation 96108273-4c8e-4a15-8d60-c719cb406700 · outbound

This paper cites Smpl: a skinned multi- person linear model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Smpl: a skinned multi- person linear model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.966149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.706861Z digest=sha256:8cc777f664dad089dee12ab6b52bb66eca972f4fbb595b67ad3a887f693fa069

Observation 517272c3-5557-4c01-86ad-23a4afe65238 · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Learning from Massive Human Videos for Universal Humanoid Pose Control Perpetual humanoid control for real-time simulated avatars

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.950302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.712113Z digest=sha256:56aebce4c99e7df651618c2fbb4cfa25dc68e1abcf8cd9d9c7821be66d072241

Observation 798ca62d-351f-4ecf-9523-5480b283fbc6 · outbound

This paper cites Universal hu- manoid motion representations for physics-based control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Universal hu- manoid motion representations for physics-based control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.934415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.716847Z digest=sha256:50678ef7cc085837073083912df28ebae1eb499cab0ef0d95241ebc3ca547868

Observation ea053ca2-a14e-4fa4-be02-ea855ee10b59 · outbound

This paper cites Vip: Towards universal visual reward and representation via value-implicit pre-training.

Learning from Massive Human Videos for Universal Humanoid Pose Control Vip: Towards universal visual reward and representation via value-implicit pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.918913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.721637Z digest=sha256:d4c486218fa31e50dae88fde212d55cea639542a5b15988285e6ca7cdaa8f563

Observation a95dbc38-23fa-4a82-b71c-d2600a61a412 · outbound

This paper cites Troje, Ger- ard Pons-Moll, and Michael J.

Learning from Massive Human Videos for Universal Humanoid Pose Control Troje, Ger- ard Pons-Moll, and Michael J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.903796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.726818Z digest=sha256:388de678aa1969d7a81eb334baaa7fd1eaabcac102bcc7696cf52091215ae634

Observation 65f9f69c-0650-4628-8869-f0e6ec985c9c · outbound

This paper cites Struc- tured world models from human videos.

Learning from Massive Human Videos for Universal Humanoid Pose Control Struc- tured world models from human videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.888403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.731534Z digest=sha256:254779a1585035f1dffaec632a852c89c7cea4bfb2327931e769cb5d0e5309ef

Observation 19274c76-e8b9-4aa1-b6bf-a7628c4a8286 · outbound

This paper cites R3m: A universal visual repre- sentation for robot manipulation.

Learning from Massive Human Videos for Universal Humanoid Pose Control R3m: A universal visual repre- sentation for robot manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.871488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.736234Z digest=sha256:5e83d0f9a889040085d2f1afe8123a48b7d505322e1fa6b9e3bcee4c64a18bf5

Observation 99fe3a29-021e-4cb5-b20f-55cdef8871fb · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Learning from Massive Human Videos for Universal Humanoid Pose Control Open x-embodiment: Robotic learning datasets and rt-x models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.853314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.740954Z digest=sha256:17a9083a414d42a127b96116f5b04be9b73752b45bd5301fb6b5ba13f065d11e

Observation b364eb6a-cc80-40e1-b49c-5947c2333a09 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Learning from Massive Human Videos for Universal Humanoid Pose Control DINOv2: Learning Robust Visual Features without Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.745609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.745609Z digest=sha256:dc0f98667b29b726fe7cb7e8c4c370f951d463f566f4502c776c5b9dd5836e13

Observation fe1b216b-3fe6-43bd-827d-d850f6677914 · outbound

This paper cites Amp: Adversarial motion priors for styl- ized physics-based character control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Amp: Adversarial motion priors for styl- ized physics-based character control

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.835861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.750403Z digest=sha256:66bfa5429a7585f3199327e2148defb156a9b4069226090fd7da3560ef474c9b

Observation 10f5b3cd-20c4-41db-a2c9-55a494f34368 · outbound

This paper cites Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.819228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.755389Z digest=sha256:31d7a1f90e976d00fd97aac31c13be3252c7fc46c65b7d34bd5e4963b28b233a

Observation 78ac66b6-88a2-4d11-86fc-a6dd473ebcc9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learn- ing transferable visual models from natural language super- vision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.801583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.760711Z digest=sha256:d4667ae90e2744515534e66818a32411d3d33d94d42050c23d4e346ed820f8ef

Observation ee266928-3e85-49ad-8d69-b8bb4bc1b5b6 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.781821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.765413Z digest=sha256:c9589e246339a8ba55b8c0f694ce7a8e2a424ea971aa3233f0573b905854ba83

Observation 0daf2c91-ad0c-4166-88df-476db7ec7403 · outbound

This paper cites Robot learning with sen- sorimotor pre-training.

Learning from Massive Human Videos for Universal Humanoid Pose Control Robot learning with sen- sorimotor pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.763624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.770997Z digest=sha256:07a374183c88af43275bc33c490c1c72ef4511be0cf2da4cd8e18bffbaf01b1e

Observation 10a7882d-d20e-4a0f-97c9-dcdca3b827cf · outbound

This paper cites Learning Humanoid Locomotion over Challenging Terrain.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning Humanoid Locomotion over Challenging Terrain

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.775903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.775903Z digest=sha256:0fe2caadeffed56b2942fabc3aff736bfa50a6e7eb38f1c8ae4652b9414f6839

Observation 2c3abc2c-05fe-4465-b3de-c461f20efa04 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Real-world humanoid locomotion with reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.747592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.780721Z digest=sha256:a8cfe6ea2f6f4d9440fecfcefc30f356c4a8335e82e74a2551094d4db0950caf

Observation 5072d50b-85ac-4216-9185-58c8aa090ba8 · outbound

This paper cites Humanoid Locomotion as Next Token Prediction.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanoid Locomotion as Next Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.785407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.785407Z digest=sha256:7b0b51971b53e5206ac9e592184d9e6650d4f475c2c988a5d30364aed0c3ecfc

Observation 5d758e81-3be2-4853-881d-1c3cf994a5f3 · outbound

This paper cites Real-Time Flying Object Detection with YOLOv8.

Learning from Massive Human Videos for Universal Humanoid Pose Control Real-Time Flying Object Detection with YOLOv8

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.790270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.790270Z digest=sha256:a3038638935061c22cf4c2bef13108e6ec41ef9b243fbd9fb8a4420b0e838c1f

Observation e672f7f7-57b8-428c-8d85-e373661896e2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Massive Human Videos for Universal Humanoid Pose Control High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.730944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.795145Z digest=sha256:8e4b01826093b49b1c79c4cca142ef3f44ad13c80b190c7b40ed64f561f7705b

Observation 254612ce-8c2f-4b76-beb8-e991700c0aaf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Massive Human Videos for Universal Humanoid Pose Control Proximal Policy Optimization Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.800392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.800392Z digest=sha256:8e81149fce8169dfc8c628da54015e9296c01943e86f7711ac91bbf540cc0145

Observation 502b82d9-b4f0-43a8-96ae-960214be5daa · outbound

This paper cites Deep imita- tion learning for humanoid loco-manipulation through hu- man teleoperation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Deep imita- tion learning for humanoid loco-manipulation through hu- man teleoperation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.716151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.805497Z digest=sha256:56be9084dbca3ba5f1a0e6b1e2013bd30a8a1d1f0a018eb61987359962c8921e

Observation 33deea5b-0ebc-4203-9b5d-2474c9124a42 · outbound

This paper cites Human motion diffusion as a generative prior.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human motion diffusion as a generative prior

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.810547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.810547Z digest=sha256:1f09848ee2469ecbb84b176f347bea462cfeaed4fc14646c4b23970915f7e1de

Observation 8c3ae30a-4b76-4b83-ba13-ae8d1bb6a003 · outbound

This paper cites Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta.

Learning from Massive Human Videos for Universal Humanoid Pose Control Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.690173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.815541Z digest=sha256:3be1acc10a659476c60fd18b081ca1a764890cde7d21138a9cf3307c9434c1c6

Observation 7914f6a8-3442-4fa8-a1c5-39be0e96e3fb · outbound

This paper cites Grab: A dataset of whole-body human grasp- ing of objects.

Learning from Massive Human Videos for Universal Humanoid Pose Control Grab: A dataset of whole-body human grasp- ing of objects

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.674122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.820635Z digest=sha256:5b3f757c23fbd2beabc63373bf6a07daff809b7244114d49e328e73ef74ee493

Observation 94c5396d-f752-4489-bb9c-a6f390277232 · outbound

This paper cites Humanmimic: Learning natural locomo- tion and transitions for humanoid robot via wasserstein ad- versarial imitation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Humanmimic: Learning natural locomo- tion and transitions for humanoid robot via wasserstein ad- versarial imitation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.656426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.825451Z digest=sha256:6b41a95b4914bc2ccace7e5e4f0e150176f1cb2a426e71e949f64fea9fa855b2

Observation ca47395a-ea0b-42e3-97ae-66ec84079f37 · outbound

This paper cites Calm: Conditional adversar- ial latent models for directable virtual characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control Calm: Conditional adversar- ial latent models for directable virtual characters

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.639383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.830885Z digest=sha256:48a8394f6840db1803504e46134ae775bac36e7fff279cd1cf9fd60b32d26a9e

Observation 3e5067d9-fb3e-4365-9bde-ef2346c2fe07 · outbound

This paper cites Human motion diffu- sion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Human motion diffu- sion model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.623011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.836307Z digest=sha256:590bcce49910812a9333c95988a3db912ce57f900266885c05ded2334197413d

Observation bd2e05f3-8ec0-4b6d-beee-a0476f0be197 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.606132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.841442Z digest=sha256:5e3f0f09cc6e8ed67a23c76f64b2c3629f3df1c29b52fe008cfc9917b51b00e0

Observation 0a0d4209-fb94-4419-996c-43c326de9651 · outbound

This paper cites Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing.

Learning from Massive Human Videos for Universal Humanoid Pose Control Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance informa- tion processing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.589903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.846487Z digest=sha256:b1c46c2e02aefdfd6f91db9fe82ece76bd9e20de69e6a2c23b3c16de01ac45be

Observation c2937247-faa9-4cc8-a806-b2180b0c35a5 · outbound

This paper cites Neural discrete representation learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control Neural discrete representation learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.573317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.851796Z digest=sha256:3f45c520cb1c2ad4e3b52ca89f935b3f5d3aef20b31073b90e1951888d90c5e9

Observation fceb11b7-b548-4501-9ce9-34e27cc0f58b · outbound

This paper cites Attention is all you need.

Learning from Massive Human Videos for Universal Humanoid Pose Control Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.857444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.857444Z digest=sha256:e6944559d6fa895003f303b13f23f5f59cfaf325e88002b5145373d59470aaf5

Observation d09ae25b-e332-47ce-b25e-74e1bf9ee26e · outbound

This paper cites A scalable approach to control diverse behaviors for physically simulated characters.

Learning from Massive Human Videos for Universal Humanoid Pose Control A scalable approach to control diverse behaviors for physically simulated characters

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.546454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.862317Z digest=sha256:ba0b87ba0cfa9588c951928a908f4fb08e21b4f4ca52e042f8de30bcfeb0f0a4

Observation 8430144f-b435-446c-8bf5-dc82c031e98d · outbound

This paper cites Masked Visual Pre-training for Motor Control.

Learning from Massive Human Videos for Universal Humanoid Pose Control Masked Visual Pre-training for Motor Control

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.867522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.867522Z digest=sha256:6564e4fbd9f4c6c393fd34dcd6172aa0df4a915d7f8e12b03d9d87cdead56642

Observation 1f6e4142-c8f6-42ed-a677-7c129db4e9a2 · outbound

This paper cites Omnicontrol: Control any joint at any time for human motion generation.

Learning from Massive Human Videos for Universal Humanoid Pose Control Omnicontrol: Control any joint at any time for human motion generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.528594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.872612Z digest=sha256:b883292f718e298a99ff14713e6d871a098d5b4c1c5a2751526b14cce5da7f86

Observation d28dbbe8-b7a0-4635-9529-1842d532c4a0 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

Learning from Massive Human Videos for Universal Humanoid Pose Control Flow as the Cross-Domain Manipulation Interface

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.877246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.877246Z digest=sha256:b80d3843786931e86e01f9574043e849597a54710af0526a29cdb03aece500aa

Observation 46a9c989-981e-4610-9ec0-290bf0ef9ad0 · outbound

This paper cites Learning interactive real-world simulators.

Learning from Massive Human Videos for Universal Humanoid Pose Control Learning interactive real-world simulators

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.512792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.881947Z digest=sha256:32d4c2e39e1fe33f1874894f3e4ba2fe48d352c7c8dcd4a59c1cca581fe19005

Observation e98182d0-7005-4f6c-990b-52c5682e3e98 · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

Learning from Massive Human Videos for Universal Humanoid Pose Control General Flow as Foundation Affordance for Scalable Robot Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.886094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.886094Z digest=sha256:b2e606ca0853dc90e99d010eb286947cf603014b461f1f7e2346d5919b413843

Observation ebfb568e-d33e-4cc0-9c8c-e98c7a8b588d · outbound

This paper cites Physdiff: Physics-guided human motion diffusion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Physdiff: Physics-guided human motion diffusion model

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.496467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.891668Z digest=sha256:a986868c1001fa0f431e60bfc7eb10514ccd01355410680ae51eadd228a32499

Observation 1330e2fd-e66e-4280-b902-9ead981309e6 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:46.896458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:46.896458Z digest=sha256:70543acbf947008f9a9476439ee88b6a9d9b83b76c4459eb5eb432a11150a55e

Observation 21434a24-904f-4380-a395-71cde8870739 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating human motion from textual descriptions with discrete representations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.480757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.902019Z digest=sha256:67104050e28c240e213d2b99d1ce2fa2c0ee5c42292f948a2a25ee66c1513201

Observation 36b9c2f7-452b-4378-a5a0-5bd034a9005c · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Learning from Massive Human Videos for Universal Humanoid Pose Control Generating human motion from textual descriptions with discrete representations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.465678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.906707Z digest=sha256:4f2c445aa0e27306588923bd2e7826055aa8b8bc46cfbc3c96ae040b8a06f652

Observation dcfd5f5c-7898-49f9-824d-a563558e11a4 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.

Learning from Massive Human Videos for Universal Humanoid Pose Control Motiondif- fuse: Text-driven human motion generation with diffusion model

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.448262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.911733Z digest=sha256:8a40a2a43f7243a39ae0cf2a471bf1d095686a8cee15c02eae84a083e93f7b16

Observation e0919a9c-1a30-4b8f-ac8e-5e1529fec4c2 · outbound

This paper cites single person.

Learning from Massive Human Videos for Universal Humanoid Pose Control single person

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.431332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.916944Z digest=sha256:3dd89d931d05b0f1844f2f69ec4b0e218f9eba0dd6fd49c36a7fb8160218ef14

Observation 4dae4c7f-b7ae-4019-abd5-e711bcd73e31 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.415315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.923121Z digest=sha256:d57248d479d387bdd58cd6b3dd362853e46305b3c6f564517389dddc8d081db5

Observation 54ebceff-50f1-4195-8523-0f0630a45f92 · outbound

This paper cites a man/woman doing something [adverb].

Learning from Massive Human Videos for Universal Humanoid Pose Control a man/woman doing something [adverb]

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:29:47.399475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.928184Z digest=sha256:3bed7b10e8d8b50f5c4ab36f6ec2a1cbf9bdbd1b06dd384b19b7423efb85a4e6

Observation 7ecde2e4-bc2f-4941-9a02-b1071c7f43a9 · outbound

This paper cites an unresolved cited work.

Learning from Massive Human Videos for Universal Humanoid Pose Control Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:29:47.382743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.933936Z digest=sha256:f3885ba098c9e8e702d7f0c618b5f124ad4ba9bddbb468a70be0f42a1e6a6225

Observation 70b81db5-3d5d-4a96-8263-430869f7885f · outbound

This paper cites in the video.

Learning from Massive Human Videos for Universal Humanoid Pose Control in the video

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T12:29:47.366443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:29:46.938716Z digest=sha256:fabea48014526d54e63ddf96e3a7baa9419141479ab9c8582905e94161f78fc8

Pith citing papers

Observation e95e81a9-ca41-49c5-869e-436bdc0b47a1 · inbound

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning cites this paper.

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:12.977121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:11:12.977121Z digest=sha256:27017e2d7d43d362fc5d12c53ec7a3fe248d8de5e28ea560dbc69878a8d1d05e

Observation 76932038-a8d1-4943-ae15-49c66e385006 · inbound

LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning cites this paper.

LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:00:28.337040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:00:28.337040Z digest=sha256:75423ae127aac854a041b77cc900246c4da19fd34e589fb9c8811e3ce222f882

Observation 84d8370a-0097-4440-8df1-dbcd5f39b247 · inbound

KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills cites this paper.

KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:00.880531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:00.880531Z digest=sha256:9d23d8929c47ab3715eeb0a42dd63b2222e561950c66a30352e83cdfdcea43c0

Observation df6e441d-3b27-495c-ad3b-49a70f3d5b3d · inbound

GMT: General Motion Tracking for Humanoid Whole-Body Control cites this paper.

GMT: General Motion Tracking for Humanoid Whole-Body Control Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:37.670272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:37.670272Z digest=sha256:a35ab1ab86986f3090965dce1a9fcc919c7d903feba3a2c035d0809e7ada6fcf

Observation b180192d-dceb-43b6-9bd1-e13689839634 · inbound

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation cites this paper.

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:30.082678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:30.082678Z digest=sha256:154f214e5a9b6ee638039cf323fe2c85fdf0e4d5c57e6cd4e7c57d7612bef2d1

Observation 1152f504-9176-47a9-8d57-31cca7cfe022 · inbound

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation cites this paper.

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T08:37:20.378298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:37:20.378298Z digest=sha256:28646bce4ca0eb1d0818d410c1dbe542d712c6ddc22e949a39b405e284e45829

Observation 85e8596c-7c13-4273-9e09-2ee668e70951 · inbound

Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement cites this paper.

Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.606859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:42:58.700609Z digest=sha256:be55aa754318b24e88d1c8cbac368f7f392a3a3be6e4b7793d0f687b9be9dff1

Observation a58d7ac7-c965-4497-be50-23df6d32eb8f · inbound

An LLM-Driven Closed-Loop Autonomous Learning Framework for Robots Facing Uncovered Tasks in Open Environments cites this paper.

An LLM-Driven Closed-Loop Autonomous Learning Framework for Robots Facing Uncovered Tasks in Open Environments Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:14.659231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T11:27:34.992950Z digest=sha256:f6f18093ef9a98d2e5b8d7f228ddf1968b0ce3055e1f562d7984050008c11b8d

Observation 7f5625e8-70d6-4949-9b5e-9488ba635e50 · inbound

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking cites this paper.

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.700049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:01:00.409280Z digest=sha256:5d71f1fc83a3a5e7f863df99bc33ee3334cb8f8f23366cd6fe84f393acedf4f3

Observation 67260d33-1e88-4235-aee7-b73dc2155708 · inbound

LIMMT: Less is More for Motion Tracking cites this paper.

LIMMT: Less is More for Motion Tracking Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.550462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:13:43.198541Z digest=sha256:0dbcb4800138688be0870f220d2138d1e65b15abb80ed5cb87291537afc9ec7c

Observation 05cf4f94-8c67-41ae-9ab7-c4f45457429a · inbound

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations cites this paper.

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T18:15:21.362695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T18:10:36.258268Z digest=sha256:a6b0eda09f49f6c6fbbb44b2a441e77d9593217c7aa76b1f00dfacbade3d7334

Observation bbb99c54-894a-46b3-aeec-a6c55834a011 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Learning from Massive Human Videos for Universal Humanoid Pose Control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.294813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.294813Z digest=sha256:73ee414c657771135b49681a64ab5a669ce85fafd2b621b8c39feac68686807b