Pith. sign in

Paper Citation Record · LEDGER

MotionLLM: Understanding Human Behaviors from Human Motions and Videos

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.20340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20340 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:29.402677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.096987Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f6dfef2-cd9a-45d5-aff0-479e14399498 · inbound

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body cites this paper.

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:29.402677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:29.402677Z digest=sha256:c2fc508bfcab7fece7d6eb133e0fa66f00d070b2e175a9ba9ae81e688f96cc9d

Observation 5438a1b2-816d-41d1-a9a9-32d81d4ca416 · inbound

KinMo: Kinematic-aware Human Motion Understanding and Generation cites this paper.

KinMo: Kinematic-aware Human Motion Understanding and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:20.987065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:20.987065Z digest=sha256:151efe8507620208b0d3c30c83ffa100612bdba54d6fef1bafe23a4b162a57bd

Observation a233246d-6ccd-4f84-8363-eeed5e8b1729 · inbound

Human Motion Instruction Tuning cites this paper.

Human Motion Instruction Tuning MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:18.222752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:18.222752Z digest=sha256:7acebe8abe75e19e9fdeaca2a71c96a05b9e35a9c9178519b515121da9d24503

Observation db136ff7-cbf6-4773-95b2-aaaff2250ae2 · inbound

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension cites this paper.

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:17:48.174143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:17:48.174143Z digest=sha256:ddcaaf4a0f20a04dade4b3d15ac4b98fbd1beef13a4f284bcf8f46e888e99f82

Observation 07729790-99de-4d0b-82e3-90bd3ef760d0 · inbound

Fleximo: Towards Flexible Text-to-Human Motion Video Generation cites this paper.

Fleximo: Towards Flexible Text-to-Human Motion Video Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:40.506000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:13:40.506000Z digest=sha256:68dfbf1acb694673dfc2ff7661d608243a774c298f1a67bf72a32be555b38bb3

Observation 97492022-295e-4998-8e7d-6951bac0b6c7 · inbound

RMD: A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse cites this paper.

RMD: A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:19.604526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:19.604526Z digest=sha256:10c23a2baa49d481e6b2c309a651f780f28d9cbb86f0badbf0b3f93e59006329

Observation efcda853-e0c6-4029-b4b2-8c04e5191e84 · inbound

CigTime: Corrective Instruction Generation Through Inverse Motion Editing cites this paper.

CigTime: Corrective Instruction Generation Through Inverse Motion Editing MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:47:38.174812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:47:38.174812Z digest=sha256:1e076560b192341c69f8d0d06598e16dd0734ab1a223158919f0c96da49cab42

Observation c189a07d-1a24-4638-b501-b92ab1a7af94 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.084732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.084732Z digest=sha256:64361f9e2263eb396d53e2f4253ecb88aea9f4bf8fdaf5ee375ea6139b363289

Observation 4ca3ea59-49d4-4f91-802a-5c214e755715 · inbound

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset cites this paper.

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:22:52.371099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:22:52.371099Z digest=sha256:7e551a1f69b9decb1703722a6012e0d9811bf082cdd97ba739b0731fb45b589e

Observation ca6bcbf5-00b7-4aa8-8117-968ff3d1926c · inbound

MONA: Moving Object Detection from Videos Shot by Dynamic Camera cites this paper.

MONA: Moving Object Detection from Videos Shot by Dynamic Camera MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:27:06.184345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:27:06.184345Z digest=sha256:9e68b7ea96299dabfd43f859d3b3cd533aa9faad487dfb36a2f46a9fb2bd1f17

Observation de743046-e7ae-48a9-87c3-59bfa99d8306 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.176113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.176113Z digest=sha256:f7be385826730cc8ba735a7af43ea4b212a815b9c0d8e5e630c575ffa6750b9c

Observation d051129c-b82d-4049-80f4-6d83223506d9 · inbound

CASIM: Composite Aware Semantic Injection for Text to Motion Generation cites this paper.

CASIM: Composite Aware Semantic Injection for Text to Motion Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:32:01.163004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:32:01.163004Z digest=sha256:76d9927d7b206a6b0b7ff280c5ebe2a93469ca5f0ad77fcf87aa7f3c56d1b40c

Observation 778a610e-bda4-495d-a79b-d180f7739efe · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.244869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.244869Z digest=sha256:ded04f9276ba68d5f3198e59084eefe1d36661fcaaaa86f593e32b8abf54b06d

Observation 05b0f365-19fe-4908-9250-23e6dc6569f5 · inbound

IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model cites this paper.

IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.585941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.585941Z digest=sha256:362ca3959e7f25a6ee7fdf762961b22acf9f846aff69f9f373ab554fab3b47c1

Observation 8ce634ec-3c2a-48dc-938a-0d2ab47f9bc8 · inbound

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward cites this paper.

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:27.907738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:27.907738Z digest=sha256:66ddf9a7cfb4daf9fcb3075419f55e3b298f0cc470ead9b9cf6d1acd68ca3895

Observation 54b94232-4936-4c5b-83ae-3360eb5bea48 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.674994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.674994Z digest=sha256:f1ad0e4af36c41b8f2fef138b83bf87d5707537528b2f8be1f3a676be3c5430f

Observation 41e14b63-2c7c-442a-a5df-bdb37a9e620a · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.784420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.784420Z digest=sha256:05afe17751360fc66023f941a735a4acc9a08ca0c60713d3037e36ad7eb591b1

Observation f7cd73b6-a9aa-41d4-97c9-144cc1f960c5 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.839617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.839617Z digest=sha256:8c2d109bce4befba809813a8c7dd6fbd6bfe4b9b65a74355929858ed08e2826f

Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.349922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.349922Z digest=sha256:d746aed15a0b7a5d0c4d3c3f71a3fd371cb429c3e8aed172291c43ebafeb6bfb

Observation 5a7c04f5-6443-4215-8b94-e45c166ae1ed · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:09.468537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:09.468537Z digest=sha256:b1ecde2062d235eec90132ec23b5662dc4ef85bd2227a7c20524d1b10920a04b

Observation 65a20522-797f-474f-bae5-05e15825bd9c · inbound

Hierarchical Motion Captioning Utilizing External Text Data Source cites this paper.

Hierarchical Motion Captioning Utilizing External Text Data Source MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:35:04.892673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:35:04.892673Z digest=sha256:4105cfd086da454e42884e66c649302e4af5863b83489ad9c73301e7178b9fef

Observation 05ee7cd5-197e-498a-9d01-c2db73514fbb · inbound

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models cites this paper.

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:59:08.668703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T05:55:11.495430Z digest=sha256:62b178e90b66d59b05dd9a81bfbf4c8250256da60bb9381e5fe68fcc96c57568

Observation a3a45041-6704-43b5-98d6-654eeb4c829f · inbound

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation cites this paper.

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.699878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.699878Z digest=sha256:61eef36d0c8c8389c145f36a637b0000b602d5e3e969f5dc42bb126b04ff7670

Observation ca3d3740-6eab-443e-9fd9-2dfdf2a49c07 · inbound

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation cites this paper.

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T20:17:37.389126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:17:37.389126Z digest=sha256:07140cbfb14c0166de74a7cdc3d430f269b2d38a9e07152df96d18e66d660ae6

Observation 76407d53-ee9b-4346-81d1-5900af9b42db · inbound

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs cites this paper.

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:21:07.101462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T21:59:00.442135Z digest=sha256:acb1ebaa06a0554cef864eac3ab756f4631ca728e2a7b51ea3c2cedb563390de

Observation 5735b75f-2e32-4b48-8fbb-c2bace7baf84 · inbound

MotionHiFlow: Text-to-motion via hierarchical flow matching cites this paper.

MotionHiFlow: Text-to-motion via hierarchical flow matching MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.922429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T08:28:42.524111Z digest=sha256:5183f7159d98134b35d4214ef1e3e2f1b8ccca733d0bfbcfd2efb2d3ccab1657

Observation 4d25b8dc-0f5e-4756-b49e-222cac66069f · inbound

WirelessSenseLLM: Zero-Shot Human Activity Understanding by Bridging Wireless Signals and Human Language cites this paper.

WirelessSenseLLM: Zero-Shot Human Activity Understanding by Bridging Wireless Signals and Human Language MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:28:30.749765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:26:59.318141Z digest=sha256:8df15f1efa4aa72a0cda46fd70da8461b76f087122f8ce9e5adde4b776fa3b8c

Observation b8cca54f-72ac-49c7-bfe2-839abf946976 · inbound

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models cites this paper.

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.201459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T06:55:22.332372Z digest=sha256:fbffd75fe0705ae16059cf112d86acab31d72a0dcbefefdd4b7d42b73706cee6

Observation 871a9618-42eb-4be6-9d14-4a500309b68d · inbound

Fine-grained Human Motion Understanding with Language Models cites this paper.

Fine-grained Human Motion Understanding with Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.100697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T17:55:49.866744Z digest=sha256:b62d739f9791a429a0fc86c83c203e338c2fdb8d76f070dba05725c96d7ba3de