Pith. sign in

Paper Citation Record · LEDGER

HuMoCon: Concept Discovery for Human Motion Understanding

As of 17 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2505.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20920 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:48:17.043562Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a04c1be-3b50-48f9-a3aa-6d6e70c25472 · outbound

This paper cites Teach: Temporal action composition for 3d hu- mans.

HuMoCon: Concept Discovery for Human Motion Understanding Teach: Temporal action composition for 3d hu- mans

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:09.924564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:09.924564Z digest=sha256:b8df941d9ef0e572f22bd1f7fe7cc4da92df819906a298feb348036cc1896e07

Observation 20f56a7f-9e9b-4041-9c62-13c5e3c02fe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.034634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.034634Z digest=sha256:fd3f629da69a7dcbab66a4976a43095a9d6e5d5ce930da06cae2a35f399b7777

Observation f1855ea7-529f-4ad9-8e66-2ca7a18ed8db · outbound

This paper cites Ac- tion quality assessment with temporal parsing transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion quality assessment with temporal parsing transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.114885Z digest=sha256:bf864d78c5c2200255726277abaa77aeb71608c3ba644ac8f44223b055d37354

Observation b59fd234-9607-451d-8ce8-c0a755cf1132 · outbound

This paper cites Implicit neural representations for variable length human motion generation.

HuMoCon: Concept Discovery for Human Motion Understanding Implicit neural representations for variable length human motion generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.177615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.177615Z digest=sha256:7fff1114c9626395a23ba19841ee4fdd5573b40d56666496c44cea1ce9e8edad

Observation 778a610e-bda4-495d-a79b-d180f7739efe · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

HuMoCon: Concept Discovery for Human Motion Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.244869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.244869Z digest=sha256:b8712ab85d2df44534d386038db60b66f2adab2221ade8e6f66a3af6696842db

Observation ca57b96e-3861-4524-8fc5-a5addd528d5c · outbound

This paper cites Pose Trainer: Correcting Exercise Posture using Pose Estimation.

HuMoCon: Concept Discovery for Human Motion Understanding Pose Trainer: Correcting Exercise Posture using Pose Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.314111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.314111Z digest=sha256:0842491c649bb2c6dda2f098edfde6023cabfa4dad8c8510026a2e10a57156e4

Observation 8ce6c9d6-75e6-4b1e-a17a-a4a5be0c0cc4 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

HuMoCon: Concept Discovery for Human Motion Understanding VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.413560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.413560Z digest=sha256:011b9d9a217a9a7aa7220bd68dae78860166829fcdc2fd94e686c3f9b6ed8274

Observation 34b2f955-4cd2-4802-847c-9557cc812b52 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.501301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.501301Z digest=sha256:7f7a8e8d7d3d7b460fcd381edd2b4cd9e71013b60e591dd5c50b8ec0d8743b8d

Observation 43231859-4654-4fd6-bf88-baf8537fb9ec · outbound

This paper cites Posefix: correcting 3d hu- man poses with natural language.

HuMoCon: Concept Discovery for Human Motion Understanding Posefix: correcting 3d hu- man poses with natural language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.565338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.565338Z digest=sha256:5df2000475635d826102012a04015661ed751b8ae53f12d1073a620d72e75132

Observation 2db973a8-70d5-4453-a49a-047820280c13 · outbound

This paper cites Behavior recognition via sparse spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Behavior recognition via sparse spatio-temporal features

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.633738Z digest=sha256:2b570710861d7056f38b76cbcbd5638c67eb2b86d4f59a9005d789c1e27fab4b

Observation 4e4dc90b-7635-42b4-b240-1c6fc46fbd75 · outbound

This paper cites Hierarchical recur- rent neural network for skeleton based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Hierarchical recur- rent neural network for skeleton based action recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.760959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.760959Z digest=sha256:7e8ee42ffacc2918bbca429a9091159fbb4f75dfd5dce52891178a103d22c72d

Observation 1489a40f-1f6a-4fa6-8b8b-bb7f10d95df4 · outbound

This paper cites Clap learning audio concepts from nat- ural language supervision.

HuMoCon: Concept Discovery for Human Motion Understanding Clap learning audio concepts from nat- ural language supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.850580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.850580Z digest=sha256:7d295d0a994191165a7f18c49ce2d8a8833a8cfa884b0c483946aeb4530849b4

Observation c782aa6f-b0b6-472c-be5e-b22c87b436cd · outbound

This paper cites Motion question answering via modular motion programs.

HuMoCon: Concept Discovery for Human Motion Understanding Motion question answering via modular motion programs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.909322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.909322Z digest=sha256:194097e99c016a0bd69d2e3d9696f34e2bf3ff7bae28da9c83313898e23ca0dc

Observation 417681e7-4df5-4451-9114-711ab491d4b7 · outbound

This paper cites Towards accurate active camera localization.

HuMoCon: Concept Discovery for Human Motion Understanding Towards accurate active camera localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.992617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.992617Z digest=sha256:14e9e573e623dea60ee49abd27c4ed371229265765bde2cbec098c641cf727a6

Observation a96723f5-5384-4022-9933-3d07570c428b · outbound

This paper cites Cigtime: Corrective instruction generation through inverse motion editing.

HuMoCon: Concept Discovery for Human Motion Understanding Cigtime: Corrective instruction generation through inverse motion editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.057582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.057582Z digest=sha256:e7977ec78e5dc02e8b228ed134a5007dcf3713b335656c724bf46cbdfbeee49a

Observation ff860867-19ff-4e5c-bc38-fb68429e0236 · outbound

This paper cites Slowfast networks for video recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Slowfast networks for video recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.903825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.142433Z digest=sha256:3f875705bebd4fc17761cc1288027c46e5a21697b9187caa95389214c80061c8

Observation 3eac355a-cfcc-4cf5-911f-9df17de7310a · outbound

This paper cites Chatpose: Chatting about 3d human pose.

HuMoCon: Concept Discovery for Human Motion Understanding Chatpose: Chatting about 3d human pose

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.889083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.225495Z digest=sha256:f112dea91b1214c940c13601b9c7e2cf4b11a6d4e764fabda1130a1cde5b20fe

Observation ea2b70de-3e15-47dc-bbb1-6c4ae6750b3a · outbound

This paper cites Aifit: Automatic 3d human-interpretable feedback models for fitness training.

HuMoCon: Concept Discovery for Human Motion Understanding Aifit: Automatic 3d human-interpretable feedback models for fitness training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.871720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.320638Z digest=sha256:00de7a30c8f0bfe7d690da0edf94a49598474f77c9d251f46ec8793c282b3977

Observation 20667788-0537-40e8-a2b5-447e05ac8a1c · outbound

This paper cites Imagebind: One embedding space to bind them all.

HuMoCon: Concept Discovery for Human Motion Understanding Imagebind: One embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.742066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.391368Z digest=sha256:6a1c54beadb96e461b45439f0dc391d26ec25a46224e9df1b0e4753a64cba1e1

Observation 51dfab67-7197-434c-aa1f-0087d77241be · outbound

This paper cites Ac- tion2motion: Conditioned generation of 3d human motions.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion2motion: Conditioned generation of 3d human motions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.407672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.475504Z digest=sha256:77ad673b1f773974cdfd361d21ff8d552d58cef20ae3fb9156637ce006f0a0d7

Observation 4997ba37-e887-4c0a-b5a8-33be0db7c361 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.123330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.547308Z digest=sha256:9738eb974bf62dbae97e5e4fa4a491df21d1c93ffd8b0cde1a9a49f47ee72ebc

Observation e775fcd9-ebfd-439c-8d60-a8bcd4cb95e9 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.893693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.636934Z digest=sha256:d5dde885055decf8a0bf304ae55a9849ed650b71483dd2b70cbd989e0286a011

Observation e15425d3-6e93-4e05-81ba-f26ababa7d1c · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts.

HuMoCon: Concept Discovery for Human Motion Understanding Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.730244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.730244Z digest=sha256:c10e05f70927e87f7c8572b1fa893d573308e0b0d2693cfa9038f12fbcf474df

Observation b78b7213-7196-48e1-8d36-82aa6765d731 · outbound

This paper cites Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.653995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.815961Z digest=sha256:137b82dbdc89adb349146f700b66291ee6d393dcf2a0ca63b7ac667aa6e806ef

Observation 29a51d22-1927-4a53-9eef-ab08a9b70ff6 · outbound

This paper cites Autoad ii: The sequel-who, when, and what in movie audio description.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad ii: The sequel-who, when, and what in movie audio description

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.434143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:11.935438Z digest=sha256:a8049f9b09df6777e16c455c23b16b05b8268de3a5f3d399825557663f78f6ee

Observation ba9114fe-6119-4eb3-9b85-ac1b458f5310 · outbound

This paper cites Autoad iii: The prequel-back to the pixels.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad iii: The prequel-back to the pixels

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.144137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:12.033488Z digest=sha256:bccc600e034e2cc284cc439404c748b41b18e0de4d2190c77f2957fb885a0f85

Observation b07742b7-366d-4efd-96cc-966dcf1311c4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

HuMoCon: Concept Discovery for Human Motion Understanding Masked autoencoders are scalable vision learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.973255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:12.126802Z digest=sha256:92bd43f599104fd4db34946744ad225d95b9c5259b11613835f854afc3180be3

Observation b482f4d7-6ad4-4402-ae8c-d7f83a1888bb · outbound

This paper cites Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.817282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:12.239768Z digest=sha256:c4a2e78024eb438b0eec7b3f63288be738bd00b4c237349c5f6e0546ed14ce31

Observation 03b1b41f-5f04-46c5-9cc2-4d0979884ee3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.389567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.389567Z digest=sha256:9091118977a73ddecc5bd8c17dd030cba922aeca59035bc83ceef47b4be3a6a9

Observation 03f14e86-f0ea-4dca-9e3a-63bced0f1254 · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.495813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.495813Z digest=sha256:fca1462f9ae755c16456e60dae3479189796fb678f7ab953816f1a3e6610fcfe

Observation e528ae26-35a2-4910-b0f0-dbd5a8e40684 · outbound

This paper cites Hand-object contact consistency reasoning for human grasps generation.

HuMoCon: Concept Discovery for Human Motion Understanding Hand-object contact consistency reasoning for human grasps generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.639694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.639694Z digest=sha256:0bdb43a8ae90146d2fb5b9689f9ada2cee379042e40d995ad9f213d876519419

Observation 64e433a8-cb03-4d13-a276-bf17fe350259 · outbound

This paper cites SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations.

HuMoCon: Concept Discovery for Human Motion Understanding SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:12.726047Z digest=sha256:c8590fabb647fd245e7f939ddde3d7443894be1f82d864627f2b016cc36024b8

Observation 9104fc61-8821-44f7-a4be-9cdafa246fb6 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

HuMoCon: Concept Discovery for Human Motion Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.856895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.856895Z digest=sha256:c6876933c1aa0bf146fccc085668ae682c98d6c509b6cc7ea64784421b8e0bb6

Observation 208464c2-a052-4d9d-92b3-afff0104b3fe · outbound

This paper cites A new representation of skele- ton sequences for 3d action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding A new representation of skele- ton sequences for 3d action recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:12.907945Z digest=sha256:18f9e3964c1856b6b16bd55323d7e2d826ab9e0440d52090b3325ef6e2fa5495

Observation 87faa928-de0b-4a98-b43c-6fe7b6f03427 · outbound

This paper cites Flame: Free- form language-based motion synthesis & editing.

HuMoCon: Concept Discovery for Human Motion Understanding Flame: Free- form language-based motion synthesis & editing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.961097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.961097Z digest=sha256:e9c3650f85c951be0c3a7a1475c935a61328ad452d939707f57b761472905fc6

Observation ea5edf5d-c586-4251-ab3f-d4051cdb4253 · outbound

This paper cites Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.598734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.005451Z digest=sha256:72a1b8e0de7a28be7fa5f4d01471753e9a72414b4d7d737a7e0d7beefcc4dedf

Observation 1721af8e-1829-47c5-9651-76f8e441ebc4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HuMoCon: Concept Discovery for Human Motion Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.085093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.085093Z digest=sha256:f681189444aa89c940e564aeef1a6e3e2626c05f608ced16c20df0e208c8c712

Observation 57916962-adad-4b7f-9fd7-09b260c88146 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.144440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.144440Z digest=sha256:a46330391711562f39dbc06f1f7b529f778d73f5d93e60d0204fcee353f304c0

Observation cd0b5d30-6ae9-4a28-bdd1-cbb6461c1914 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

HuMoCon: Concept Discovery for Human Motion Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.472147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.202566Z digest=sha256:92c57dbedfddf9bf4181dd98c5281f47f1bcaff29dea840b9bdf70250cf69c87

Observation 16f9336b-0b54-4c04-8488-460a6540235f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.245692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.245692Z digest=sha256:424432f0194bd7f6ba5c7ec67424c8098815062eb3f11c74677268110dfa776f

Observation 0b529021-9209-4e40-942b-db8c8cd4b16c · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.321591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.341764Z digest=sha256:e3b330fabe20f53fc4260ece54388ccfc1e4a7e30570a1fea4e759cefc4d4e53

Observation e6a15de1-018e-4550-b7a3-1cafb400d427 · outbound

This paper cites Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.182622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.427954Z digest=sha256:0fa3ec583452f1adb0c3addffb8a9408d2cf47be10ecb73423be2fd7979dc725

Observation e4c8e21f-edb1-4d66-b8df-79f53e57b561 · outbound

This paper cites Towards unified sur- gical skill assessment.

HuMoCon: Concept Discovery for Human Motion Understanding Towards unified sur- gical skill assessment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.994813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.478624Z digest=sha256:d32f3e469f5292ea57b62d83cb25deef628d084d5449323fe1f799ef5f07a072

Observation c5dd6ce4-9680-4d59-b44d-fd72a6387429 · outbound

This paper cites Spatio-temporal lstm with trust gates for 3d human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatio-temporal lstm with trust gates for 3d human action recognition

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.822773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.547011Z digest=sha256:0140b847ac9bf42c3ee16a4e9e142f3acc0d04582cce6551d261361ff9e8dd20

Observation 187ab89a-4c1d-4011-bfd7-2c4d3ab69860 · outbound

This paper cites Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.683090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.607043Z digest=sha256:76d2ee93dda86a205aa62ee66dbd9cdb0c3b518251a9debc449403b1389ab1e2

Observation d90d895f-741d-4b2e-8605-da08a2e85ae3 · outbound

This paper cites InfoCon: Concept Discovery with Generative and Discriminative Informativeness.

HuMoCon: Concept Discovery for Human Motion Understanding InfoCon: Concept Discovery with Generative and Discriminative Informativeness

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.297953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.659969Z digest=sha256:fc760760266c8e0b9280f74bf79ff089b70518b4fd2d9d01905128ef484dd474

Observation 00d2d0fb-ce91-421d-9212-bbf7b32180b8 · outbound

This paper cites Posegpt: Quantization-based 3d human mo- tion generation and forecasting.

HuMoCon: Concept Discovery for Human Motion Understanding Posegpt: Quantization-based 3d human mo- tion generation and forecasting

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.435492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.714895Z digest=sha256:772e9e1ba767b94509f8b574eba5e0b4d40e894fef04c42809267ed00b298b8f

Observation 8f552295-2322-4b0f-8068-6be3fc8388ce · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.253372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.805257Z digest=sha256:b34441d6c2deef3e2c2bdf0ca398bac1ccb94b27fd11d7437b9a0edbe9fe9ed1

Observation 28359c28-23f9-4524-89dc-45a73fa2755f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.881665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.881665Z digest=sha256:cfd4e09acaaecfb43b3b7f23bf8574998f5360b5e87f88a51e91e73c78d8bd39

Observation d5e98ac8-6eb9-4368-ae87-f5f683df8823 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.947245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.947245Z digest=sha256:11b25341385ecc3147ba601d786ab3cc3f09f8fd070a401b47e2c654eea3b01b

Observation 9b0fb8e3-4431-46bb-a5b7-380fcf050f4d · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

HuMoCon: Concept Discovery for Human Motion Understanding What and how well you performed? a multitask learning approach to action quality assessment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.136730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:13.998154Z digest=sha256:bab15992aa449d3c6713eedfaa14d104953f9503bf67671e667956e4c3a4f9dc

Observation c3457ee8-ab87-472d-853d-dca2a07d5498 · outbound

This paper cites Action- conditioned 3d human motion synthesis with transformer vae.

HuMoCon: Concept Discovery for Human Motion Understanding Action- conditioned 3d human motion synthesis with transformer vae

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.873624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.050854Z digest=sha256:82b50005ba092f8c271cd89cb468fbdab8c0a7a7dc98f45e126a42bac3e242b4

Observation da2de4c7-d835-4237-985d-8b695a69468b · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

HuMoCon: Concept Discovery for Human Motion Understanding Temos: Generating diverse human motions from textual descriptions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.094528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.094528Z digest=sha256:32bed57ea30b7d80ef069e9af3917d71128a3fbb975bf995145e9336bbbedbae

Observation 7a7255e9-6b35-4a8c-b1d6-d10d5156d21b · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

HuMoCon: Concept Discovery for Human Motion Understanding Babel: Bodies, action and behavior with english la- bels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.607481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.138546Z digest=sha256:576b7ca9ed1ab41bf5e6194f605dd012635c60a94f2c9c374bbcaa69b23b5955

Observation 65f19c98-0b95-4647-bcf4-e2983ec9d777 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HuMoCon: Concept Discovery for Human Motion Understanding Learning transferable visual models from natural language supervi- sion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.425341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.215738Z digest=sha256:9c8dfa61a2a41513d9e17da7f33575d8d405b1977d5518f5ef1e52e2aa279bf0

Observation 831fc922-753d-41b5-aa46-34fb4338c55c · outbound

This paper cites On the benefits of 3d pose and tracking for human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding On the benefits of 3d pose and tracking for human action recognition

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.196705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.279679Z digest=sha256:1626916141e652eb2b4e8bc3e4e16e46a691b5e6b9f60d93018041e9cfec0dda

Observation 7fe7f90a-a16d-4072-87f5-c4acfd15452d · outbound

This paper cites Zero-shot audio captioning with audio-language model guidance and audio context keywords.

HuMoCon: Concept Discovery for Human Motion Understanding Zero-shot audio captioning with audio-language model guidance and audio context keywords

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.346070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.346070Z digest=sha256:55d428dd446933b99f46833062fe9806ea755b37eb50436b98e73e11bb7178b1

Observation 8c342166-b30e-4688-bdf9-2c447fdf3826 · outbound

This paper cites Two- stream adaptive graph convolutional networks for skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Two- stream adaptive graph convolutional networks for skeleton- based action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.862099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.405018Z digest=sha256:0f3107572faea1bb5f17bd2e4e7b58a41fb9d540bd3757d281e0abafad861e37

Observation 2b58ca6e-3251-406b-bc50-9d004b6fdca0 · outbound

This paper cites Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019.

HuMoCon: Concept Discovery for Human Motion Understanding Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.647262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.464679Z digest=sha256:851856853e9ead9d54e0bd9f000b7f47248afe87f2dc27f2c5dee7aaaf383b1c

Observation 3af641c8-f034-457d-b777-cafcd8bcf05a · outbound

This paper cites Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020.

HuMoCon: Concept Discovery for Human Motion Understanding Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.523931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.523931Z digest=sha256:31a4e07f87f0c3d247cd3504e959bcbd21cff6661e07f2210371226c7655dc9e

Observation 6663db44-167f-4f5f-923c-045a26aecb86 · outbound

This paper cites Deepphase: Periodic autoencoders for learning motion phase manifolds.

HuMoCon: Concept Discovery for Human Motion Understanding Deepphase: Periodic autoencoders for learning motion phase manifolds

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.436539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.601668Z digest=sha256:8196a756af37b560152ce8fa2d23a319aaba89f7d7e44b0d8e23629e5fb21f1a

Observation d06102dc-abc2-48cb-9248-6985586b7706 · outbound

This paper cites Convolutional learning of spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Convolutional learning of spatio-temporal features

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.236498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.704401Z digest=sha256:3c86aac1ee16626bb47c95cae7e6913298326cedaa2597840f60ac3f4ad6c5f1

Observation d1fe6ca9-1bbc-405f-873d-1e16e0da08c3 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

HuMoCon: Concept Discovery for Human Motion Understanding Motionclip: Exposing human motion generation to clip space

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.032251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.788254Z digest=sha256:b7d559e0b9e92d2d43e7798197ad3406cec1fc90ffc4d2e69577142251b97d4c

Observation 5e1ffcd6-416e-4ea7-bf77-1091569b3774 · outbound

This paper cites Human Motion Diffusion Model.

HuMoCon: Concept Discovery for Human Motion Understanding Human Motion Diffusion Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.844576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.844576Z digest=sha256:95c2d3d383549b5aadcd4e3bb00d45805b8e349becd20cbf76b603d32ebeb987

Observation 3c5815f4-3f2d-4da3-bc78-a8f56306f38c · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

HuMoCon: Concept Discovery for Human Motion Understanding Learning spatiotemporal features with 3d convolutional networks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.909199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.909199Z digest=sha256:f9c547e32c1f100158474d34aadd0f4f9133e159bfb1fab67fa4d51640355c2a

Observation 56c8bae8-3a4c-453b-a1be-7e88c2a6ed36 · outbound

This paper cites Human action recognition by representing 3d skeletons as points in a lie group.

HuMoCon: Concept Discovery for Human Motion Understanding Human action recognition by representing 3d skeletons as points in a lie group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.873807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:14.987206Z digest=sha256:f259c49fb34ab4a2ec1dab01157c284bc77a4e835c2f7cff1d419c02f63457a7

Observation 06f1a0dd-c61b-4c52-9b0c-a9bd360d9b94 · outbound

This paper cites Action recognition with improved trajectories.

HuMoCon: Concept Discovery for Human Motion Understanding Action recognition with improved trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.555895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.061942Z digest=sha256:69018d0a1ae9eee90100194034cf4504a91146225af26af8de2e17c732a66b3e

Observation a3db810f-40c1-4d18-ab05-227898ea01bf · outbound

This paper cites Learn- ing human dynamics in autonomous driving scenarios.

HuMoCon: Concept Discovery for Human Motion Understanding Learn- ing human dynamics in autonomous driving scenarios

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.420267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.133096Z digest=sha256:dae45b3f743d103a80ea18a9234d647b8ad6d144e48066da8080771574d91700

Observation 59f471f3-70f9-4451-b87b-921b62699906 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

HuMoCon: Concept Discovery for Human Motion Understanding Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.259990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.244074Z digest=sha256:9937271a0d2f76cd41586ad5d8ac671def806aac885cd02004721559129c409b

Observation 8401b309-0b16-4b9d-ac1b-30f5586f3016 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.328435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.328435Z digest=sha256:719afc5f9559e8ea679359a83e95206eb6737bf12ed5c16567cef1c5f649e735

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:42434a0525183702e0edf10769ce139d3ddeed01788e89aea119235474744cf0

Observation b6deda1a-9d71-425a-98b1-b0e2b25f293a · outbound

This paper cites Unified Human-Scene Interaction via Prompted Chain-of-Contacts.

HuMoCon: Concept Discovery for Human Motion Understanding Unified Human-Scene Interaction via Prompted Chain-of-Contacts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.499056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.499056Z digest=sha256:5a906e4f0041fa3ddeb6d5e4f8b2e219c15b24a3b18b74725215d2a5cbed5707

Observation be953bc4-6c03-4eb8-8487-f0ea8c17fd8c · outbound

This paper cites AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description.

HuMoCon: Concept Discovery for Human Motion Understanding AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.573298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.573298Z digest=sha256:abc9eb7204cf8705027d218093ec915dd64e780082074fad7dd4850e57ba6c71

Observation d007fada-1dd1-47bc-89fb-6af1778e96e1 · outbound

This paper cites Dexterous grasp transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Dexterous grasp transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.011797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.647947Z digest=sha256:c10319bc68a44d580f4551116e43a62e400fa7f2e1fc0a8dbe30736df844e21f

Observation f76756c4-f32c-49d6-9d50-f76b9c548aa5 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

HuMoCon: Concept Discovery for Human Motion Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.720080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.720080Z digest=sha256:7ce33ee5f930cfa5d0c71ef5dc63619378d8d2bf675c78020248d2ad587f147f

Observation 10c3c04d-cff4-4210-8c7f-e449cfe56b35 · outbound

This paper cites Spatial tempo- ral graph convolutional networks for skeleton-based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatial tempo- ral graph convolutional networks for skeleton-based action recognition

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.777468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.777468Z digest=sha256:70fab2bf82916442b8ce16866ba413dd8a8740ff90d4d40f537ed518a4571a7d

Observation 37fd70ca-82b3-4a2e-8099-524b47f3f24d · outbound

This paper cites Learning to use chopsticks in diverse gripping styles.

HuMoCon: Concept Discovery for Human Motion Understanding Learning to use chopsticks in diverse gripping styles

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.789211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.830957Z digest=sha256:e04cc899049054e4f4ac3d7dc5183623949a0b10e58d97243195987bd2b0e8ac

Observation d7b03f79-b05b-4883-8c10-c6785a696709 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

HuMoCon: Concept Discovery for Human Motion Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.523041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.910767Z digest=sha256:28f271c1748a4e6de3e65be986e584badbbba3a8d5a1a49908600967ba9a7902

Observation d21bf28c-63a4-4cd6-ac0b-94068ef65ac2 · outbound

This paper cites Sigmoid loss for language image pre-training.

HuMoCon: Concept Discovery for Human Motion Understanding Sigmoid loss for language image pre-training

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.278946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:15.969101Z digest=sha256:891d06194e1c269db96e0081124fad8b2e0a5ffa582dc1d7d78428c3a225e175

Observation 5c7172aa-8f9f-46c2-bc77-e4df44dd84ca · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.029474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.029474Z digest=sha256:6c02b3168180ca7775be344ea1ebf1a17fe4cd85486db7e3d170f0bffbfea0b7

Observation 47b9ecd6-67b5-43ac-b2b8-7def70600694 · outbound

This paper cites T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations.

HuMoCon: Concept Discovery for Human Motion Understanding T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.105127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.105127Z digest=sha256:51bd19f8ea48a01fa6f7607677641e8d230190d91d5eb88a418cd6105c3a6781

Observation 9d0d6c33-8e81-444f-b42b-1d09f4fc4fd7 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.045164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.159016Z digest=sha256:74fa53ea870b114990dd7aaf87b98373c5eb693bc1e129035daa6fc2514a9ec7

Observation 54025af5-d1e1-441b-94d3-973e27b5fe10 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

HuMoCon: Concept Discovery for Human Motion Understanding Pointclip: Point cloud understanding by clip

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.857074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.224869Z digest=sha256:f5cc50be27efa761d826d8f994e6c0cece73b5e937b5f98a2be4d9d450faf611

Observation 3d7d98d9-8c2b-4cdc-943d-70a929f975b8 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

HuMoCon: Concept Discovery for Human Motion Understanding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.292122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.292122Z digest=sha256:2cd637d5d0504982cf20a5310750487a5ae9c1c1cb329655f33b53f3bcca5ab9

Observation 7d80e39c-03c8-4383-9fbd-44a6bb6f78f5 · outbound

This paper cites Streaming dense video captioning.

HuMoCon: Concept Discovery for Human Motion Understanding Streaming dense video captioning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.710966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.441949Z digest=sha256:338907eb63886d0674747925f897e664b7530f9a4abae136d76212848b95d014

Observation a48bdeba-b3a4-4799-b1b4-30ab37674bb7 · outbound

This paper cites Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.526788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.539175Z digest=sha256:cf99e0149c7bf9faea313280ecbdad05302bcfa708e1a85a1a58ff5b3db63af4

Observation e2d842ec-0c37-4a27-a756-081035a8de30 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

HuMoCon: Concept Discovery for Human Motion Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.643392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.643392Z digest=sha256:c4414e65da5ddbf16c3b55f8ff48f0a0b188809972067609325b4612e99805bc

Observation 15700c2e-c4d9-429e-96b6-bb2ca64ecb9e · outbound

This paper cites Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5].

HuMoCon: Concept Discovery for Human Motion Understanding Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5]

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.383910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.773797Z digest=sha256:043b568332066fb8d3850ad85e86f43f5c672c955d2a45d24380afe26e816b68

Observation d23a2469-1fb5-4239-8f6e-29726d222219 · outbound

This paper cites Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig.

HuMoCon: Concept Discovery for Human Motion Understanding Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.128667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.873059Z digest=sha256:763473ba0acfd323f35c6e6e05c5bd25cd10e2c1b7ce1ae91bb14211c05d3334

Observation 20a2d8f4-4502-4b51-8466-9ff35e6c1ca7 · outbound

This paper cites Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80].

HuMoCon: Concept Discovery for Human Motion Understanding Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80]

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:17.934458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:16.966088Z digest=sha256:cba7660cbe0802cf716b4b458ca2e491c88baf5a9fc416a4c115309e6ef29f55

Observation e97ab098-55cd-4ba2-9b0f-dcf258105aed · outbound

This paper cites Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46].

HuMoCon: Concept Discovery for Human Motion Understanding Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46]

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:48:17.751549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:48:17.043562Z digest=sha256:f1a84e75403db0752e40ff27ca0aa4fecc266095fbbc01eed44899f91e704de3

Pith citing papers

No inbound Pith citation observations are available.