Pith. sign in

Paper Citation Record · LEDGER

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding

As of 11 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2507.12463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12463 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:49:54.185917Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71839772-d37c-4bb6-8090-7cbf3de9a016 · outbound

This paper cites Auxiliary tasks benefit 3d skeleton-based human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Auxiliary tasks benefit 3d skeleton-based human motion prediction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.455478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.455478Z digest=sha256:322edb570b3bad67aa00382da26c0fc4195a0da2e4f624b14e96076a49d1c6ca

Observation a96daa7f-fbc2-4f4a-880c-6c94bc60c808 · outbound

This paper cites Tamformer: Multi-modal transformer with learned attention mask for early intent prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tamformer: Multi-modal transformer with learned attention mask for early intent prediction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.600197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.600197Z digest=sha256:838ec3830fb4e4ea19131d73e380c0795bcc2872c8e5e7ba4e1aec7e7f0db02c

Observation c47b113e-39f4-4286-8073-9c025b0ab892 · outbound

This paper cites GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:49:55.192546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:36.704505Z digest=sha256:e756275ed400f4d952c7ed83e9ded15b946c929413e656a950bea7c9d7d52f80

Observation 27815572-f287-4085-a7ad-e23add9bdce2 · outbound

This paper cites Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.860276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.860276Z digest=sha256:e4da0da3d67f28b56b9434442baa5af5063701d4b9fbf5b65f91858191ce9323

Observation e65c5890-52c5-41db-8aec-26f07361ddfc · outbound

This paper cites Incorporating physics principles for precise human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Incorporating physics principles for precise human motion prediction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.992085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.992085Z digest=sha256:0b5fb10d435aef3e63ed88b18b09135581cdb74f6cf14abe480ba49c442f79a4

Observation f19f81f0-32c3-4616-976b-4e360e651e53 · outbound

This paper cites Behavioral intention prediction in driving scenes: A survey.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Behavioral intention prediction in driving scenes: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.104896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.104896Z digest=sha256:190cf65bedc41339f750f478580bdaf13a6046cd810fc220bde1bb34eb4b30cb

Observation 668c5fa7-85d0-4f6d-a0aa-a5a34ede1a94 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Scalability in perception for autonomous driving: Waymo open dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.233375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.233375Z digest=sha256:504b60e3df68b297e73757bf8495e1ee8ef784e1c4b86f1a733662f3db0e86da

Observation 8f10b2ba-4ba7-4d84-bdfa-d7cb13c856f6 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding nuscenes: A multimodal dataset for autonomous driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.369537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.369537Z digest=sha256:6d708b9f7d2679b9a50e4de4ce3f953292df8031a087e6cdea09a142b066b288

Observation 6b0d5c9a-88d9-44fe-b3ba-60febd80c9e9 · outbound

This paper cites Vision meets robotics: The kitti dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Vision meets robotics: The kitti dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.511468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.511468Z digest=sha256:b2185ebeb7f745d313d88f6b226d45f5359c4bd4e69693e3040e4767daf3b796

Observation 6a6153d1-418d-4584-bf64-aab75b599cfb · outbound

This paper cites Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.661119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.661119Z digest=sha256:3c1d3e9ad576b3c43a037318155343422b34534a9bd953f627ea35345d28dddb

Observation b50f0152-fb57-4c86-824d-22574cc8200a · outbound

This paper cites Euro-pvi: Pedestrian vehicle interactions in dense urban centers.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Euro-pvi: Pedestrian vehicle interactions in dense urban centers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.815286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.815286Z digest=sha256:d08504ff2d6293f9c2a3ba0357f18039edfa009ad8a1702b3d9c5d4a7fbb2dd2

Observation 75c4dd4a-6aa8-4f10-b74c-b9cbea0c36f2 · outbound

This paper cites Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.957541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.957541Z digest=sha256:ce7ed9cdfce56562d40ccdacffc9db8644ba621ed15adcb1b603469763a188d8

Observation 2c03a62f-90fb-4846-abd9-0525de31d5b6 · outbound

This paper cites Recovering accurate 3d human pose in the wild using imus and a moving camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recovering accurate 3d human pose in the wild using imus and a moving camera

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.083776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.083776Z digest=sha256:a971e2e28db6c7a77f36fd8fd20fd5cb0044fc2d1dd32038ec745da89262140b

Observation de699b7c-0835-46f1-9b34-031d2a002a56 · outbound

This paper cites Pedestrian motion reconstruction: A large-scale benchmark via mixed reality rendering with multiple perspectives and modalities.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian motion reconstruction: A large-scale benchmark via mixed reality rendering with multiple perspectives and modalities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.220473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.220473Z digest=sha256:ea4d11ed07ec8478d13029d8f91a2c5179f9ad8dc84912a3fd1ae5e4ba0d76d8

Observation 35a10f2d-3953-4c15-83fd-a9fd04b6a633 · outbound

This paper cites Text to blind motion.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Text to blind motion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.336758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.336758Z digest=sha256:e799182c4421b370ca3c73e4db223a0aff00c2eafcdc1dbfc0f09adca64423fe

Observation 522147a0-1c2e-4da7-b5dc-f7fc850ea232 · outbound

This paper cites Learning to generate diverse pedestrian move- ments from web videos with noisy labels.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Learning to generate diverse pedestrian move- ments from web videos with noisy labels

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.508404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.508404Z digest=sha256:10a4c34699feb14cb69a6e01b300c7e83e26ac78712a286133b5149d41358ee8

Observation 5f264ace-4e2b-475a-8332-4e25d805e24b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.658154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.658154Z digest=sha256:7dd6560ceab3af78f279c77c562691fc64c6f83768aeed7d90e53f9054025619

Observation c37d2d69-251c-4188-a3a1-f199219b0a40 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Vila: On pre-training for visual language models, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.780959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.780959Z digest=sha256:9db03f5148df95e8ef462436fe30f7ade9b33c37583389016ef4ddbe986af457

Observation 7ba56018-4ad9-4ac4-91cf-f28fe1dadac0 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Longvila: Scaling long-context visual language models for long videos, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.897379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.897379Z digest=sha256:8613f5f47353e43c1a9edb17e2b2055a4a13c587fbff4de44edb230f9350df30

Observation 693cdbeb-5477-427d-9c0c-ccac7a52ccb5 · outbound

This paper cites Nvila: Efficient frontier visual language models, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Nvila: Efficient frontier visual language models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.033356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.033356Z digest=sha256:f580d829e8d545d9b074818357ae9eb8fa7977b180dd9eda5c005da9b39c50f5

Observation 297c659b-7536-4a3d-ad99-c81590c0b5f8 · outbound

This paper cites Visual instruction tuning, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.144588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.144588Z digest=sha256:8902391508d40041224b97589625ff6f27ce9364b2ff61ffb5785c96059f0123

Observation a63ec538-4950-44c5-b9a0-10b1c4594b7f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.289917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.289917Z digest=sha256:a8eef118e2193871527544b2307258a22e1403391f933ef5adba2a2a0cb3ff63

Observation 1453a2f1-8bde-4cfd-9337-264e70df223f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Improved baselines with visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.431858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.431858Z digest=sha256:c18cbc536faa41b64d24da81ae95a4d147a279ea80afd57fcdb2cd8921c787ea

Observation 2eb5d404-b758-49a9-b3a4-66d32a56798c · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Video instruction tuning with synthetic data, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.607726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.607726Z digest=sha256:e104b865bd48c5a54c0005421251750360b0ba3a6c525b544df37d54b27e3da8

Observation aea0b38f-be94-4ce3-8094-dbf9f85e5196 · outbound

This paper cites Qwen Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.726727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.726727Z digest=sha256:d21a475fbe39fcce347e5fc35eaa0239fe234eb350ac8bac55406e7386be753c

Observation a666b112-4d18-43ab-b4ba-28345d4a0e29 · outbound

This paper cites 2d human pose estimation calibration and keypoint visibility classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding 2d human pose estimation calibration and keypoint visibility classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.878547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.878547Z digest=sha256:7d2afcc4c00bb7ad66e3ab21d37b31d9d4692d42ae307a67b2fc4c7ec412d8e2

Observation c57d3b76-d08b-4d33-b7f5-c64d243f7602 · outbound

This paper cites Recurrent human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recurrent human pose estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:40.039779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:40.039779Z digest=sha256:cd06b3d4b53695dcb75dc8c109f642fb38906c2a33d563925889f45473bb10ba

Observation fa480358-192c-4714-9253-d3f1016384a7 · outbound

This paper cites Rethinking the heatmap regression for bottom-up human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rethinking the heatmap regression for bottom-up human pose estimation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:04.784210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:40.230695Z digest=sha256:2e77099d58b6bac6f02593ba10258ad94a388e2fceb9a571afb8373354592161

Observation 892c2586-5580-411a-8265-3ab02a313167 · outbound

This paper cites Tokenpose: Learning keypoint tokens for human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tokenpose: Learning keypoint tokens for human pose estimation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:04.659736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:40.355038Z digest=sha256:55ee3f312e944052112d38a198810316f5b5333b59676a045105c249fac9bc58

Observation 8a758109-a470-4bd3-a647-71bfc7302460 · outbound

This paper cites Whole- body human pose estimation in the wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Whole- body human pose estimation in the wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:41.266257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:41.266257Z digest=sha256:28763042d922e6a898fe92ce03af0069c4f5c513913919b7848e901d4fb242af

Observation 94e049d9-f517-4a89-a27f-b739d48817cb · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:04.442700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:42.533122Z digest=sha256:709496b3c3750e6689db9440c794609a5499752ea5a11f24a097eb9615360786

Observation 9b883b61-44b7-437b-bcc5-0cd0cc3b1c58 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:04.210756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:42.642434Z digest=sha256:6292f72d906e39f2472e52683ee99e7131a66d449efe4e6bd2cc8bb92629f094

Observation 59899439-f75b-4e85-a161-55b8cac73f39 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:03.926076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:43.103542Z digest=sha256:8816385960a305cedb8942ba758ac7168222b62655f76aebdb2ee6da8c93b0d4

Observation 1c010a3d-66b7-4b03-9a71-7c68846f1581 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:45.787654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:45.787654Z digest=sha256:41578ad4180226a3b5ee84ab9484fb0aeec89f7fa1361f3a0807cd3e5320cc1f

Observation 0a50da58-49dc-4d31-8ab4-40816ba7ab2e · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:45.980246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:45.980246Z digest=sha256:9b6cf33f52007d71c2d94d1e451b92a51e050d2cf3d5882dcab82e63a4735538

Observation 27b6e82b-430e-42c8-9144-895cccf7a313 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Amass: Archive of motion capture as surface shapes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:46.171972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:46.171972Z digest=sha256:bfd300e80fa21cf2a2470e16470b53cd32d220232074f2bd5363737dbc10eb63

Observation f4e72f73-3dac-49b6-a2cc-7bbff6b27495 · outbound

This paper cites Motion- x: A large-scale 3d expressive whole-body human motion dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motion- x: A large-scale 3d expressive whole-body human motion dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.653273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:46.305846Z digest=sha256:ef1f489d2eeb20f2cce428c632b49a243cf3d8444674a60c30e6bec7cc19d7af

Observation d30fa8f4-654e-4a01-9c9a-f20182624064 · outbound

This paper cites Dynamic multi-person mesh recovery from uncalibrated multi-view cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dynamic multi-person mesh recovery from uncalibrated multi-view cameras

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.383774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:46.407522Z digest=sha256:e41f678fc02c4a6cdf6d8f11a43ceb96e73f3d68166ced5f11cc505ddc6fe61a

Observation 867a74ba-8797-4300-a5cb-9d6c5274a6f9 · outbound

This paper cites Motion capture from internet videos.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motion capture from internet videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.152948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:46.555723Z digest=sha256:fd2c8cb35f3bf4f97db488ed9b5a7b0a14ba5d87c0ec9e8eef64ca2ede072035

Observation 8d9a3d61-079c-4639-a8bb-838233d8e2f8 · outbound

This paper cites Scene-aware 3d multi-human motion capture from a single camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Scene-aware 3d multi-human motion capture from a single camera

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:46.756807Z digest=sha256:18f08c4d30a1ea083e8d914f681b5045da2db06f779c4023c921075d8743259c

Observation 2291b9e1-8e3d-4ee4-afc5-36e6e2ef7029 · outbound

This paper cites D &d: Learning human dynamics from dynamic camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding D &d: Learning human dynamics from dynamic camera

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.746385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:46.912517Z digest=sha256:486cb6aaad4d00b8d83777a6f07c7dcca8a6b7f80ec83fbda55ddadca86f7deb

Observation c096dc11-32ef-427e-8d4b-02f99abb36f5 · outbound

This paper cites Decoupling human and camera motion from videos in the wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Decoupling human and camera motion from videos in the wild

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.536495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:47.066608Z digest=sha256:1d8408731a4a92c1e0d31471e45ad5595c509e58d42ad82cbf26803ef771775a

Observation 6aefdd6d-1981-4a44-a1f1-0dddfaf1523e · outbound

This paper cites Glamr: Global occlusion-aware human mesh recovery with dynamic cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Glamr: Global occlusion-aware human mesh recovery with dynamic cameras

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.319830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:47.201746Z digest=sha256:3370f2f4979a15cf8221e8bd099f2182c185798ff637cb98e1983807f9939cef

Observation 9820c467-ba5a-4440-aa1b-2bf2e57f7dc5 · outbound

This paper cites Implicit neural representations for variable length human motion generation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Implicit neural representations for variable length human motion generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.098770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:47.271744Z digest=sha256:e0994e6b90bfba62e523373f95662979fbc3c0d7d05e5075f0efa7b21598bf00

Observation 376615c3-cc43-41f0-9c2c-c1b895599a2c · outbound

This paper cites Human Motion Diffusion Model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Human Motion Diffusion Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.444311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.444311Z digest=sha256:218be8ebb63185b6a857139813c9454524320767b45193aaebd2abca2803dd2f

Observation 03f7e6d3-2a41-4896-8caa-e2297dc1c84a · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motiondiffuse: Text-driven human motion generation with diffusion model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.619490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.619490Z digest=sha256:6df193a37610c0870ebd5d65fb3fd7b4c9da0f93dc3465bbcda06a848de8511a

Observation 9ea47c99-d57c-4a00-904f-4631a1ac504c · outbound

This paper cites Home action genome: Cooperative compositional action understanding.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Home action genome: Cooperative compositional action understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.780696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.780696Z digest=sha256:24387c7aa99e283b05db0769f65bdbc3380ed5b67a3dfd385cd2aedc0ba44600

Observation b4758825-c618-4953-bb34-bb8a00c56d74 · outbound

This paper cites Babel: Bodies, action and behavior with english labels.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Babel: Bodies, action and behavior with english labels

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.923957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:47.926998Z digest=sha256:6b1bc2b1fc687feb0ffb7a3e54864a9317c91d30c19f36a2fe40391bbb1da6e3

Observation be6c8ef9-7a1f-4292-a862-7d67243e69d2 · outbound

This paper cites Mining actionlet ensemble for action recognition with depth cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Mining actionlet ensemble for action recognition with depth cameras

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.710745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:48.054445Z digest=sha256:2e52fc810af4cdaebc07d5c58b597e06c4c03b9b9c9609fc761c5513ed7f3a63

Observation 92a5830b-19b3-47c5-a2b6-c13a1da481bb · outbound

This paper cites Action recognition based on a bag of 3d points.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Action recognition based on a bag of 3d points

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.495068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:48.182938Z digest=sha256:7024ea3223d93698182a520cd821ebc263c81675818c267b4d2d0d12c7578650

Observation 6a2104d7-8a74-41c3-be0e-22d7490af9a9 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:48.439467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:48.439467Z digest=sha256:2a714828f70db28e9c81457b7f5b51460710454df6649a08497a80b03833c0a2

Observation d4bca55f-f114-4a84-8c6b-b03b8097020a · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Modeling temporal structure of decomposable motion segments for activity classification

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.314815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:48.662758Z digest=sha256:5de142b478191626e56d759881892779747c76f3fb3fb424ff2463c099bc3f5c

Observation c58ba53b-e930-46e4-8461-78f9e80d71dd · outbound

This paper cites The Kinetics Human Action Video Dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding The Kinetics Human Action Video Dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:48.737903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:48.737903Z digest=sha256:30fe6854a87be66713eef9603b7361a7a11b3b4823b1a85430271b812494e824

Observation 92b440a4-e734-4821-a45c-8212ea0819ed · outbound

This paper cites Actions in context.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Actions in context

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.161871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:48.924740Z digest=sha256:921051a00ae12233c8e37a5a66ed62f7091ac86f28348c24ab8306c00f34f595

Observation 164e9ab0-c904-4bf8-8d68-3c8a9b41626b · outbound

This paper cites Auxformer: Robust approach to audiovisual emotion recognition.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Auxformer: Robust approach to audiovisual emotion recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.946957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.076536Z digest=sha256:4def42a60c6e1fd21f202522157ac0e18ec8062578a21fcab5559a8394c3fe42

Observation 7ead9bb5-7390-4fe1-8aaa-50dc25b53977 · outbound

This paper cites Context-based inter- pretable spatio-temporal graph convolutional network for human motion forecasting.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Context-based inter- pretable spatio-temporal graph convolutional network for human motion forecasting

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.804751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.207068Z digest=sha256:2f861cedcffbd01222b6df4a71ba10ef7500eae3bb1e3ae52769220c517cc30f

Observation c0391ee2-d79c-45bd-99b2-d74667c2bae5 · outbound

This paper cites Gcnext: Towards the unity of graph convo- lutions for human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gcnext: Towards the unity of graph convo- lutions for human motion prediction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.618553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.340159Z digest=sha256:953152786b29789e51475770497040e4a0acb116fa96e041e63961ce9fd41cbd

Observation 2f2d564a-098b-42ee-927d-698371054ade · outbound

This paper cites Back to mlp: A simple baseline for human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Back to mlp: A simple baseline for human motion prediction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.443157Z digest=sha256:2e4cf10eb3a591eac03dcfe45dc26c73ae061f869843c4a7c87ede4324fbe63c

Observation ed0b00e1-56de-4771-950a-999aac11ce18 · outbound

This paper cites History repeats itself: Human motion prediction via motion attention.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding History repeats itself: Human motion prediction via motion attention

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.194250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.578637Z digest=sha256:16a1ecad9b3fe26398134cce45e6e566ef09c35345a751c8bfd8cc6760b2aaca

Observation 85d4f0b6-6f2d-4329-bc6b-378e3b4f8379 · outbound

This paper cites Trep: Transformer-based evidential prediction for pedestrian intention with uncertainty.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Trep: Transformer-based evidential prediction for pedestrian intention with uncertainty

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.936136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.698402Z digest=sha256:2f240ccf397c7b7bbd6a0870a154c38624807be1d0671093b93a683d2002b1bd

Observation 2e0955c3-2a8a-4fe0-bb23-cbb3bf827409 · outbound

This paper cites Visual–motion–interaction-guided pedestrian intention prediction framework.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Visual–motion–interaction-guided pedestrian intention prediction framework

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.759578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.805932Z digest=sha256:b91d3549294c735dfa283637601c1e30caed956563b3c02e35158dd4e800c1ae

Observation f0fe421d-6901-4329-b2f8-aab7dc2edacd · outbound

This paper cites Benchmark for evaluating pedestrian action prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Benchmark for evaluating pedestrian action prediction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.547880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:49.841996Z digest=sha256:1de260c7c550aa20034ec1b2d951feca2a7b89ea1a1aea5d2556ad1dac0d301b

Observation 8ceeceaf-1327-4ba8-bea8-02410bca2679 · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Agreeing to cross: How drivers and pedestrians communicate

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.335344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.076568Z digest=sha256:2e859536f2bc89c384fa4d66f778701a6be49c607d3bbf49ef53414e61277d52

Observation 179f464a-4f06-48cc-8735-37ee93e2c362 · outbound

This paper cites Pedestrian intention prediction based on dynamic fuzzy automata for vehicle driving at nighttime.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian intention prediction based on dynamic fuzzy automata for vehicle driving at nighttime

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.134669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.298586Z digest=sha256:1ddf2cec50d922941bd06547f39f4055604782b76587db1310c64506dd2e3cde

Observation 7fa810cf-0bc2-4a02-b7fd-52823e3d55f3 · outbound

This paper cites Pedestrian intention and pose prediction through dynamical models and behaviour classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian intention and pose prediction through dynamical models and behaviour classification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.977044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.432971Z digest=sha256:4bad710beccd8e4a5b08a58a9b00a96e30902d148d4a809b6622048a4fdff902

Observation bea57fc8-781b-4767-99e4-422fe8c50f5d · outbound

This paper cites Pedestrian path prediction with recursive bayesian filters: A comparative study.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian path prediction with recursive bayesian filters: A comparative study

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.799103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.558560Z digest=sha256:4990d081cf3cc7c12a4de11247074b06644562fc9e30d1e8abb515c0bdd2fff2

Observation 51bebc89-e5e6-4884-b173-a92fa96ceabd · outbound

This paper cites Dolphins: Multimodal language model for driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dolphins: Multimodal language model for driving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:50.712847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:50.712847Z digest=sha256:22d5ef9fb2a08f47e69c793434009ddc4aa6b657829d4ab982c3d8dfbb24c6d7

Observation 2d51aa2e-fb38-43af-8dac-c2b5fcf3ba25 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Drivelm: Driving with graph visual question answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.596482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.803207Z digest=sha256:7c28c5ed9c36976b02670fa9776f0e285ee18999a3cd74b1944e375aefd3b911

Observation 10676d0d-7943-42c1-9a00-4e7cbe2b0cd8 · outbound

This paper cites Tem-adapter: Adapting image-text pretraining for video question answer.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tem-adapter: Adapting image-text pretraining for video question answer

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.389195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:50.879641Z digest=sha256:249b82958e61201334d992b840a984f5df6cd265741c4efef610f43fbf80ebe9

Observation e6c44bc0-0547-4c9c-bdfb-b981eedf1829 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Lmdrive: Closed-loop end-to-end driving with large language models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:50.978757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:50.978757Z digest=sha256:3bbdcecd898447d8aa77a8b1933a6878e095c8e6fdea59505c708f11e2aee13b

Observation 0d2be7b6-6737-4ed4-ae54-7934b12a7f6f · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.079501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.079501Z digest=sha256:bdf96e0b8a91066d7157c40fda0e970b1dc29246d65728a82499356394515960

Observation d39b9996-86f6-481e-8701-3389039ebcf5 · outbound

This paper cites Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.158814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.158814Z digest=sha256:0c03efa8db7ecd4895a22541ad0d627018737aceee27d837c4ad738b4d79493a

Observation 68838d8b-e6f7-43f1-872d-d54892d9d338 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.167373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:51.254950Z digest=sha256:f35b164d653041faec038537c45ac100e42aaca98c1aab05642496cb2b447211

Observation da9bb143-f317-4cbe-8537-f40b3002c5e2 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Lingoqa: Visual question answering for autonomous driving

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.984400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:51.438008Z digest=sha256:ad020cf5c9ac063778146112c0a2d096b5677a350608e198c3f8e6b5a15c26a4

Observation fd6b48ab-d441-4973-8c40-d5b1556c37ac · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Covla: Comprehensive vision-language-action dataset for autonomous driving

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.633873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.633873Z digest=sha256:f9b99a30d57987e78fa44c0b7240696ff88305a8f088037c45d68a617cf7002b

Observation 3e784a95-c9bb-45c9-9b7d-d40f50b91555 · outbound

This paper cites Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.751162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:51.737098Z digest=sha256:c1537db4bb1c89764859069edaacd7f267543460896ac768ff2ab1ef2c3bc2b1

Observation 607d30e9-3a00-4aaf-9820-30d0b87c37e2 · outbound

This paper cites Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.500537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:51.910116Z digest=sha256:e6a0de669698b2639be63a0b560a1aa8ae2a11dfd91a5aae4c8c20cddb6aa5fc

Observation e5f7465c-63d7-4549-83db-691b68c40cc8 · outbound

This paper cites NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.051367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.051367Z digest=sha256:4833799d1cb113501db2c7b7ad3df28ab8ab22b6fdce29a15ccf44e2c7a7291c

Observation a79d6fa2-8fd4-4d0e-bd84-197aed7768f3 · outbound

This paper cites 1 year, 1000 km: The oxford robotcar dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding 1 year, 1000 km: The oxford robotcar dataset

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.288537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.224282Z digest=sha256:f0730e04c9a59ee8a42e1ea3d157f35f2c2b6339b390743cf3723840c2f1dd50

Observation 9bb8aba5-a4b5-4106-9d54-f1f3eba09c6d · outbound

This paper cites Dada: Driver attention prediction in driving accident scenarios.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dada: Driver attention prediction in driving accident scenarios

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.106647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.296701Z digest=sha256:47b6b787fa598d6513b66ebffe0edf3c67e2b018b22db81da47b9b29d0116794

Observation ef3aa4fc-2030-4b75-91b9-11c50c2c4fe4 · outbound

This paper cites Snow removal in video: A new dataset and a novel method.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Snow removal in video: A new dataset and a novel method

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.933123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.391187Z digest=sha256:1a8846b02e4806f75d4eae60e03c0ae5121cbd7e44da38d98702e10039966107

Observation 633e92b7-df81-4138-bd02-74811c27fd65 · outbound

This paper cites Semantic foggy scene understanding with synthetic data.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Semantic foggy scene understanding with synthetic data

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.711029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.460177Z digest=sha256:81c1cd47266a98f571b22b5e0d8bd4d3573463a09eee5c5b3deb469343bb3e70

Observation 14e4f89f-40d7-4f5b-9d28-723b02c3b296 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.501706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.544376Z digest=sha256:5d0f448da209e3c60984da56f81d71bc5e03ebc1d2ae816d356cb8a25da62b67

Observation 5dfdc58d-34c0-4e05-afb3-b59c6fd130fe · outbound

This paper cites Recov- ering accurate 3d human pose in the wild using imus and a moving camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recov- ering accurate 3d human pose in the wild using imus and a moving camera

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.335057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.648254Z digest=sha256:5ab4554e1430bfe994d164984d69086d31fbf1f0201267c377afaa9adeb8d2f5

Observation d3a61028-8328-42fc-adfe-ee4e7bb0957b · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Drama: Joint risk localization and captioning in driving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.765409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.765409Z digest=sha256:6302f554a3ba70cea526f54f2dde83b8f6e85c28eb18b3484cecb409363b84ea

Observation a7cc4c9c-ef71-4152-930d-8a4d0a52d6f3 · outbound

This paper cites Posescript: 3d human poses from natural language.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Posescript: 3d human poses from natural language

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.146203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:49:52.848079Z digest=sha256:2235b142db9b29f82a4eaf9cdddb06fc44a7286c2b3177a8c3bb31e1b41fceb2

Observation d09a7f04-0b4f-465b-b0ce-31e679d39e3b · outbound

This paper cites Motiongpt: Human motion as a foreign language.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motiongpt: Human motion as a foreign language

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.954611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.954611Z digest=sha256:d9e5ad3bfdcb6cfe72b00c1c04b22a8dfe532dc14d07c6ec49e434bfda254dba

Observation d6efac6a-c297-4a31-abc6-8c9401a22094 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.064565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.064565Z digest=sha256:ebfaa4168dc5e92f3ecc0966bfb5dc97355589ef79c3ea84acfecc2486410d62

Observation b8fa220a-9394-478e-ae1a-55896c6ea69d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.159318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.159318Z digest=sha256:0d3369e5e5413575fb2445d0b80caa90203006ff723182a0064537a6977d5128

Observation 716f75ef-6c74-4cd4-9f30-f5613f70e6bd · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.272808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.272808Z digest=sha256:ee537bed8150cadf15ae45d697fb216fe2ac98d626137873faf4cd7f456c93c0

Observation 7200b3b0-b926-4848-bd33-5343c9a6842c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.371557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.371557Z digest=sha256:80bffef78c1207c44abd3a8e0493ac3e0a95524fe801a148a7ea2e53b48a93cf

Observation a5b2b250-4156-402e-98a8-275f3f32c366 · outbound

This paper cites Qwen2.5-VL Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen2.5-VL Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.451178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.451178Z digest=sha256:9862dc8ad01338420edb03fff938f02ac741f766e118d5f2088051df76289911

Observation 26395a82-440b-4d02-9634-af4b864fd32f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.541236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.541236Z digest=sha256:1a6ae1e229fc6e2b8ffa55b78258c432cfbeda45558cc3815229a0f494f0c5d4

Observation 68ea932b-35ff-451a-91b8-44938f6df4d9 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.632371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.632371Z digest=sha256:f128b08f37a2204f58e498a90e26bfab99d5f074719c1eebdcdc543061e35782

Observation f3e330ce-1d74-48eb-a971-be76d1ae37d0 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions., 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Building and better understanding vision-language models: insights and future directions., 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.709943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.709943Z digest=sha256:1a6fae244e6e8939a72e5440f3d1853a91b89c37e1b65a185c2a9024de361c99

Observation de930206-53e1-4297-90c8-e1418f1cbca7 · outbound

This paper cites Pixtral 12B.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pixtral 12B

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.785315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.785315Z digest=sha256:c4c2a111e0180871d56706dd564ff6773ba3c0bd6334a6f3d52f7bf6a521e880

Observation 31e8e320-eaac-4ab4-8bc5-b55eee0c7e62 · outbound

This paper cites Gemma 3 Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gemma 3 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.884739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.884739Z digest=sha256:e22c1740179d3b7c9f417ef2e97fbb9211c455ec419e4b6ae6118df76915ced1

Observation f4b662be-748a-4797-b619-ddf3794da39d · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.990949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.990949Z digest=sha256:bfd25322ee3d8e3b184d9145e893b739fcd34409c19a12893760e0fa09c40d9e

Observation 84beb5f3-69ff-4066-9f02-b08cae3daa5f · outbound

This paper cites Kimi-VL Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Kimi-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:54.072518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:54.072518Z digest=sha256:a05532a057212b73061b2937abd84d6858fb4f718da7e54526e66cc820abd9e6

Observation f1010b81-99b5-464e-8195-5401d9bd769b · outbound

This paper cites Joint Attention in Autonomous Driving (JAAD).

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Joint Attention in Autonomous Driving (JAAD)

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:54.185917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:54.185917Z digest=sha256:2fcd87327ad0f9cc9e346fde6ca0b9ef1679e371d1e59d8bf29b99eb15543e09

Pith citing papers

No inbound Pith citation observations are available.