Pith. sign in

Paper Citation Record · LEDGER

MotionLLM: Understanding Human Behaviors from Human Motions and Videos

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.20340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20340 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:29.402677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.096987Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f6dfef2-cd9a-45d5-aff0-479e14399498 · inbound

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body cites this paper.

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:29.402677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:29.402677Z digest=sha256:2acacf63094e075db76f57c35dd00909468cda201e732b0782e2295177ec8ae0

Observation 5438a1b2-816d-41d1-a9a9-32d81d4ca416 · inbound

KinMo: Kinematic-aware Human Motion Understanding and Generation cites this paper.

KinMo: Kinematic-aware Human Motion Understanding and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:20.987065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:20.987065Z digest=sha256:cf4693e213b98e7cc55bdfd7d4f4adc767b191b1d2cc403dd3f8d4505769bf54

Observation a233246d-6ccd-4f84-8363-eeed5e8b1729 · inbound

Human Motion Instruction Tuning cites this paper.

Human Motion Instruction Tuning MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:18.222752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:18.222752Z digest=sha256:581dbca2b3502aaac6ba185e8c01a3f93586ef414c7a717a2634d394d6be42d6

Observation db136ff7-cbf6-4773-95b2-aaaff2250ae2 · inbound

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension cites this paper.

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:17:48.174143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:17:48.174143Z digest=sha256:e5d7635400e06d1180298f560e9e0a5c702da8203aaaf7151cf9828e754962a9

Observation 07729790-99de-4d0b-82e3-90bd3ef760d0 · inbound

Fleximo: Towards Flexible Text-to-Human Motion Video Generation cites this paper.

Fleximo: Towards Flexible Text-to-Human Motion Video Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:40.506000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:13:40.506000Z digest=sha256:732ea5f357632628e3b2b4a7fdb3346f5512a61167bbb1c9d783f88eba5b493d

Observation 97492022-295e-4998-8e7d-6951bac0b6c7 · inbound

RMD: A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse cites this paper.

RMD: A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:19.604526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:19.604526Z digest=sha256:26c8b058bdd7994c365d9ba8517622352e550c6b2b65164373e83926839236cf

Observation efcda853-e0c6-4029-b4b2-8c04e5191e84 · inbound

CigTime: Corrective Instruction Generation Through Inverse Motion Editing cites this paper.

CigTime: Corrective Instruction Generation Through Inverse Motion Editing MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:47:38.174812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:47:38.174812Z digest=sha256:99cd2b6472b283142215bcf43ef0642548072c530c1f45fc5b7d36d61cc1bbb6

Observation c189a07d-1a24-4638-b501-b92ab1a7af94 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.084732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.084732Z digest=sha256:27f31f8441ad55419211b9b3e4902cc14527039c7b81d5fcb0e966f522f597e2

Observation 4ca3ea59-49d4-4f91-802a-5c214e755715 · inbound

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset cites this paper.

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:22:52.371099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:22:52.371099Z digest=sha256:91b87cf65bfadb5c038a98a7e8cc50331c1f8747b65f40b8086c69f69b4c6b96

Observation ca6bcbf5-00b7-4aa8-8117-968ff3d1926c · inbound

MONA: Moving Object Detection from Videos Shot by Dynamic Camera cites this paper.

MONA: Moving Object Detection from Videos Shot by Dynamic Camera MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:27:06.184345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:27:06.184345Z digest=sha256:da949e7dccb247a233ddb27d29d5159101a895f57df8fb856f4dc5dfb59c2471

Observation de743046-e7ae-48a9-87c3-59bfa99d8306 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.176113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.176113Z digest=sha256:c03b391e02e064a6217013d25b5a67b37e38e82531fbc70f648a816d17a86b0d

Observation d051129c-b82d-4049-80f4-6d83223506d9 · inbound

CASIM: Composite Aware Semantic Injection for Text to Motion Generation cites this paper.

CASIM: Composite Aware Semantic Injection for Text to Motion Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:32:01.163004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:32:01.163004Z digest=sha256:64e8abe11dbf7257553b019a9255c3eb5ca4e3d2de1bed306ac3176d8c9aee7b

Observation 778a610e-bda4-495d-a79b-d180f7739efe · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.244869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.244869Z digest=sha256:45a340dbf1a67a505097d2b0c3df3ca236d637f564e513a4cafebcc96b9d156c

Observation 05b0f365-19fe-4908-9250-23e6dc6569f5 · inbound

IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model cites this paper.

IKMo: Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.585941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.585941Z digest=sha256:504f0bcc0d58d20ffd12446c360b4160bc958f82957f7c4ec699371797825673

Observation 8ce634ec-3c2a-48dc-938a-0d2ab47f9bc8 · inbound

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward cites this paper.

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:27.907738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:27.907738Z digest=sha256:61305b34c0657c030c5443629b57612a508d4831fd8375203e42e061ff6d8a2f

Observation 54b94232-4936-4c5b-83ae-3360eb5bea48 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.674994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.674994Z digest=sha256:2784d24a4f7a4b3568c75f093b166de0aa64be12f943c5be16fb8e571f0fd146

Observation 41e14b63-2c7c-442a-a5df-bdb37a9e620a · inbound

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model cites this paper.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.784420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.784420Z digest=sha256:1473e11f4f6abd7d211a02e79f501b77184d3514236c95aca116b879211db0eb

Observation f7cd73b6-a9aa-41d4-97c9-144cc1f960c5 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.839617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.839617Z digest=sha256:11a0e27c6314efdf19c61e32e89c96e98b6619772c0bf971879a4ae03c771129

Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.349922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.349922Z digest=sha256:88df051162c63b661e95f7c7a29bbdd4f5abeadf50708ddf25af13c0dac15ad5

Observation 5a7c04f5-6443-4215-8b94-e45c166ae1ed · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:09.468537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:09.468537Z digest=sha256:8e5242965f160c7c4fab3e69baa14a179caada3f4545ae573dd841a6a5a06b9e

Observation 65a20522-797f-474f-bae5-05e15825bd9c · inbound

Hierarchical Motion Captioning Utilizing External Text Data Source cites this paper.

Hierarchical Motion Captioning Utilizing External Text Data Source MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:35:04.892673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:35:04.892673Z digest=sha256:533d69182297285cbdc45a830f8a94eb689cf94829c4480e17c83d7f674abd5b

Observation 05ee7cd5-197e-498a-9d01-c2db73514fbb · inbound

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models cites this paper.

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:59:08.668703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T05:55:11.495430Z digest=sha256:99483c28233100305dac441f57698c9f4adfaf984a3cfb874675903453694c5b

Observation a3a45041-6704-43b5-98d6-654eeb4c829f · inbound

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation cites this paper.

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.699878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.699878Z digest=sha256:6a19a84974e33ef79562cbbec9f9d3390b2513b4273b6074f6b01868fed40484

Observation ca3d3740-6eab-443e-9fd9-2dfdf2a49c07 · inbound

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation cites this paper.

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T20:17:37.389126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:17:37.389126Z digest=sha256:e512ac5bded149e862951316029cdb93b8e0f8672375f9140137d121248a365b

Observation 76407d53-ee9b-4346-81d1-5900af9b42db · inbound

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs cites this paper.

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:21:07.101462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T21:59:00.442135Z digest=sha256:8f71df6e2e43344743112f90f3548089deb231094e04aba4c6096349b66fe401

Observation 5735b75f-2e32-4b48-8fbb-c2bace7baf84 · inbound

MotionHiFlow: Text-to-motion via hierarchical flow matching cites this paper.

MotionHiFlow: Text-to-motion via hierarchical flow matching MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.922429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T08:28:42.524111Z digest=sha256:79ca78568f44aab44d25eb111ddc4109c5b8ca9c9144625e6727bdb75ed7a482

Observation 4d25b8dc-0f5e-4756-b49e-222cac66069f · inbound

WirelessSenseLLM: Zero-Shot Human Activity Understanding by Bridging Wireless Signals and Human Language cites this paper.

WirelessSenseLLM: Zero-Shot Human Activity Understanding by Bridging Wireless Signals and Human Language MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:28:30.749765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:26:59.318141Z digest=sha256:e55b3f1ea29337bfc6e4e91b7d347edecbd574b8db89628a3644f6f11faec953

Observation b8cca54f-72ac-49c7-bfe2-839abf946976 · inbound

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models cites this paper.

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.201459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T06:55:22.332372Z digest=sha256:1d5baff76878af51c52d59e2062368ffc5262470e9d37014fdea73e37ad304f7

Observation 871a9618-42eb-4be6-9d14-4a500309b68d · inbound

Fine-grained Human Motion Understanding with Language Models cites this paper.

Fine-grained Human Motion Understanding with Language Models MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.100697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T17:55:49.866744Z digest=sha256:6027260c46acdf4aad561b8c7e3e02352c2049f088adf08ec9a964a5e41ffdd6