Pith. sign in

Paper Citation Record · LEDGER

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding

As of 9 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2507.12463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12463 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:49:54.185917Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71839772-d37c-4bb6-8090-7cbf3de9a016 · outbound

This paper cites Auxiliary tasks benefit 3d skeleton-based human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Auxiliary tasks benefit 3d skeleton-based human motion prediction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.455478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.455478Z digest=sha256:322edb570b3bad67aa00382da26c0fc4195a0da2e4f624b14e96076a49d1c6ca

Observation a96daa7f-fbc2-4f4a-880c-6c94bc60c808 · outbound

This paper cites Tamformer: Multi-modal transformer with learned attention mask for early intent prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tamformer: Multi-modal transformer with learned attention mask for early intent prediction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.600197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.600197Z digest=sha256:838ec3830fb4e4ea19131d73e380c0795bcc2872c8e5e7ba4e1aec7e7f0db02c

Observation c47b113e-39f4-4286-8073-9c025b0ab892 · outbound

This paper cites GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding GTransPDM: A Graph-embedded Transformer with Positional Decoupling for Pedestrian Crossing Intention Prediction

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:49:55.192546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:36.704505Z digest=sha256:51ff20f286618bc06c772dc01d2b37334989cb46e5cf95f237f0f5a8a241657b

Observation 27815572-f287-4085-a7ad-e23add9bdce2 · outbound

This paper cites Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.860276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.860276Z digest=sha256:e4da0da3d67f28b56b9434442baa5af5063701d4b9fbf5b65f91858191ce9323

Observation e65c5890-52c5-41db-8aec-26f07361ddfc · outbound

This paper cites Incorporating physics principles for precise human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Incorporating physics principles for precise human motion prediction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:36.992085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:36.992085Z digest=sha256:0b5fb10d435aef3e63ed88b18b09135581cdb74f6cf14abe480ba49c442f79a4

Observation f19f81f0-32c3-4616-976b-4e360e651e53 · outbound

This paper cites Behavioral intention prediction in driving scenes: A survey.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Behavioral intention prediction in driving scenes: A survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.104896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.104896Z digest=sha256:190cf65bedc41339f750f478580bdaf13a6046cd810fc220bde1bb34eb4b30cb

Observation 668c5fa7-85d0-4f6d-a0aa-a5a34ede1a94 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Scalability in perception for autonomous driving: Waymo open dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.233375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.233375Z digest=sha256:504b60e3df68b297e73757bf8495e1ee8ef784e1c4b86f1a733662f3db0e86da

Observation 8f10b2ba-4ba7-4d84-bdfa-d7cb13c856f6 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding nuscenes: A multimodal dataset for autonomous driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.369537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.369537Z digest=sha256:6d708b9f7d2679b9a50e4de4ce3f953292df8031a087e6cdea09a142b066b288

Observation 6b0d5c9a-88d9-44fe-b3ba-60febd80c9e9 · outbound

This paper cites Vision meets robotics: The kitti dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Vision meets robotics: The kitti dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.511468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.511468Z digest=sha256:b2185ebeb7f745d313d88f6b226d45f5359c4bd4e69693e3040e4767daf3b796

Observation 6a6153d1-418d-4584-bf64-aab75b599cfb · outbound

This paper cites Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.661119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.661119Z digest=sha256:3c1d3e9ad576b3c43a037318155343422b34534a9bd953f627ea35345d28dddb

Observation b50f0152-fb57-4c86-824d-22574cc8200a · outbound

This paper cites Euro-pvi: Pedestrian vehicle interactions in dense urban centers.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Euro-pvi: Pedestrian vehicle interactions in dense urban centers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.815286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.815286Z digest=sha256:d08504ff2d6293f9c2a3ba0357f18039edfa009ad8a1702b3d9c5d4a7fbb2dd2

Observation 75c4dd4a-6aa8-4f10-b74c-b9cbea0c36f2 · outbound

This paper cites Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:37.957541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:37.957541Z digest=sha256:ce7ed9cdfce56562d40ccdacffc9db8644ba621ed15adcb1b603469763a188d8

Observation 2c03a62f-90fb-4846-abd9-0525de31d5b6 · outbound

This paper cites Recovering accurate 3d human pose in the wild using imus and a moving camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recovering accurate 3d human pose in the wild using imus and a moving camera

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.083776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.083776Z digest=sha256:a971e2e28db6c7a77f36fd8fd20fd5cb0044fc2d1dd32038ec745da89262140b

Observation de699b7c-0835-46f1-9b34-031d2a002a56 · outbound

This paper cites Pedestrian motion reconstruction: A large-scale benchmark via mixed reality rendering with multiple perspectives and modalities.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian motion reconstruction: A large-scale benchmark via mixed reality rendering with multiple perspectives and modalities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.220473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.220473Z digest=sha256:ea4d11ed07ec8478d13029d8f91a2c5179f9ad8dc84912a3fd1ae5e4ba0d76d8

Observation 35a10f2d-3953-4c15-83fd-a9fd04b6a633 · outbound

This paper cites Text to blind motion.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Text to blind motion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.336758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.336758Z digest=sha256:e799182c4421b370ca3c73e4db223a0aff00c2eafcdc1dbfc0f09adca64423fe

Observation 522147a0-1c2e-4da7-b5dc-f7fc850ea232 · outbound

This paper cites Learning to generate diverse pedestrian move- ments from web videos with noisy labels.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Learning to generate diverse pedestrian move- ments from web videos with noisy labels

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.508404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.508404Z digest=sha256:10a4c34699feb14cb69a6e01b300c7e83e26ac78712a286133b5149d41358ee8

Observation 5f264ace-4e2b-475a-8332-4e25d805e24b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.658154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.658154Z digest=sha256:7dd6560ceab3af78f279c77c562691fc64c6f83768aeed7d90e53f9054025619

Observation c37d2d69-251c-4188-a3a1-f199219b0a40 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Vila: On pre-training for visual language models, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.780959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.780959Z digest=sha256:9db03f5148df95e8ef462436fe30f7ade9b33c37583389016ef4ddbe986af457

Observation 7ba56018-4ad9-4ac4-91cf-f28fe1dadac0 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Longvila: Scaling long-context visual language models for long videos, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:38.897379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:38.897379Z digest=sha256:8613f5f47353e43c1a9edb17e2b2055a4a13c587fbff4de44edb230f9350df30

Observation 693cdbeb-5477-427d-9c0c-ccac7a52ccb5 · outbound

This paper cites Nvila: Efficient frontier visual language models, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Nvila: Efficient frontier visual language models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.033356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.033356Z digest=sha256:f580d829e8d545d9b074818357ae9eb8fa7977b180dd9eda5c005da9b39c50f5

Observation 297c659b-7536-4a3d-ad99-c81590c0b5f8 · outbound

This paper cites Visual instruction tuning, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.144588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.144588Z digest=sha256:8902391508d40041224b97589625ff6f27ce9364b2ff61ffb5785c96059f0123

Observation a63ec538-4950-44c5-b9a0-10b1c4594b7f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.289917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.289917Z digest=sha256:a8eef118e2193871527544b2307258a22e1403391f933ef5adba2a2a0cb3ff63

Observation 1453a2f1-8bde-4cfd-9337-264e70df223f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Improved baselines with visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.431858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.431858Z digest=sha256:c18cbc536faa41b64d24da81ae95a4d147a279ea80afd57fcdb2cd8921c787ea

Observation 2eb5d404-b758-49a9-b3a4-66d32a56798c · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Video instruction tuning with synthetic data, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.607726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.607726Z digest=sha256:e104b865bd48c5a54c0005421251750360b0ba3a6c525b544df37d54b27e3da8

Observation aea0b38f-be94-4ce3-8094-dbf9f85e5196 · outbound

This paper cites Qwen Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.726727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.726727Z digest=sha256:22bb637e4ac43bcf339edc6523cd3b3822df2c5efab2359207000391ca7347f1

Observation a666b112-4d18-43ab-b4ba-28345d4a0e29 · outbound

This paper cites 2d human pose estimation calibration and keypoint visibility classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding 2d human pose estimation calibration and keypoint visibility classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:39.878547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:39.878547Z digest=sha256:7d2afcc4c00bb7ad66e3ab21d37b31d9d4692d42ae307a67b2fc4c7ec412d8e2

Observation c57d3b76-d08b-4d33-b7f5-c64d243f7602 · outbound

This paper cites Recurrent human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recurrent human pose estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:40.039779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:40.039779Z digest=sha256:cd06b3d4b53695dcb75dc8c109f642fb38906c2a33d563925889f45473bb10ba

Observation fa480358-192c-4714-9253-d3f1016384a7 · outbound

This paper cites Rethinking the heatmap regression for bottom-up human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rethinking the heatmap regression for bottom-up human pose estimation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:04.784210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:40.230695Z digest=sha256:5dd8c8d0dbdb7c831c20a1c71a21f1b858bdd2a5a026d83f86c06e2c3f338248

Observation 892c2586-5580-411a-8265-3ab02a313167 · outbound

This paper cites Tokenpose: Learning keypoint tokens for human pose estimation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tokenpose: Learning keypoint tokens for human pose estimation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:04.659736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:40.355038Z digest=sha256:a435d0ba3bc9684ee546ca1911d3224fdb395f77561534c37a4538359d213c52

Observation 8a758109-a470-4bd3-a647-71bfc7302460 · outbound

This paper cites Whole- body human pose estimation in the wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Whole- body human pose estimation in the wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:41.266257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:41.266257Z digest=sha256:28763042d922e6a898fe92ce03af0069c4f5c513913919b7848e901d4fb242af

Observation 94e049d9-f517-4a89-a27f-b739d48817cb · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:04.442700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:42.533122Z digest=sha256:c55004f349f308393ed1c7a6d9c4217e4d29c741b24d8bc693f54a1d431a6212

Observation 9b883b61-44b7-437b-bcc5-0cd0cc3b1c58 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:04.210756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:42.642434Z digest=sha256:78c0ce41d94457fe58210654fe153a9f964b974539f73b11e372087218992549

Observation 59899439-f75b-4e85-a161-55b8cac73f39 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:50:03.926076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:43.103542Z digest=sha256:f36da55c666423299d27e2e9788ab7867a55d65b9c7e8cf647b3e9dda6b59205

Observation 1c010a3d-66b7-4b03-9a71-7c68846f1581 · outbound

This paper cites an unresolved cited work.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:45.787654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:45.787654Z digest=sha256:41578ad4180226a3b5ee84ab9484fb0aeec89f7fa1361f3a0807cd3e5320cc1f

Observation 0a50da58-49dc-4d31-8ab4-40816ba7ab2e · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:45.980246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:45.980246Z digest=sha256:c806e0376d937499597652e385a65ac3816b529ee8515face836dfa49e83546c

Observation 27b6e82b-430e-42c8-9144-895cccf7a313 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Amass: Archive of motion capture as surface shapes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:46.171972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:46.171972Z digest=sha256:bfd300e80fa21cf2a2470e16470b53cd32d220232074f2bd5363737dbc10eb63

Observation f4e72f73-3dac-49b6-a2cc-7bbff6b27495 · outbound

This paper cites Motion- x: A large-scale 3d expressive whole-body human motion dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motion- x: A large-scale 3d expressive whole-body human motion dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.653273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:46.305846Z digest=sha256:817daf9383422a010af5ad87c75737f7495e829a7b20d178af7ee0a9c8522739

Observation d30fa8f4-654e-4a01-9c9a-f20182624064 · outbound

This paper cites Dynamic multi-person mesh recovery from uncalibrated multi-view cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dynamic multi-person mesh recovery from uncalibrated multi-view cameras

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.383774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:46.407522Z digest=sha256:82161469514deb2ccf9f8c4978f6745edf7716236a09c4758912595b13334fe6

Observation 867a74ba-8797-4300-a5cb-9d6c5274a6f9 · outbound

This paper cites Motion capture from internet videos.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motion capture from internet videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:03.152948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:46.555723Z digest=sha256:d5527dfc9c26625ca1a6a462a7a6052aa2322f0d910a32ae54c05af118b5c665

Observation 8d9a3d61-079c-4639-a8bb-838233d8e2f8 · outbound

This paper cites Scene-aware 3d multi-human motion capture from a single camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Scene-aware 3d multi-human motion capture from a single camera

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:46.756807Z digest=sha256:eea97d59f883621fadf8634dfb2b4befe9158f2c036a32855beedc458027ef82

Observation 2291b9e1-8e3d-4ee4-afc5-36e6e2ef7029 · outbound

This paper cites D &d: Learning human dynamics from dynamic camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding D &d: Learning human dynamics from dynamic camera

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.746385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:46.912517Z digest=sha256:0975687718cdd8e6bdf84a37b8308732da22131b07258ae777c9ddbeaa0476dd

Observation c096dc11-32ef-427e-8d4b-02f99abb36f5 · outbound

This paper cites Decoupling human and camera motion from videos in the wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Decoupling human and camera motion from videos in the wild

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.536495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:47.066608Z digest=sha256:bdb85a8d8b7caf74e79389e5fd08e86db7bb62bd06c29d261e927d1c93525271

Observation 6aefdd6d-1981-4a44-a1f1-0dddfaf1523e · outbound

This paper cites Glamr: Global occlusion-aware human mesh recovery with dynamic cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Glamr: Global occlusion-aware human mesh recovery with dynamic cameras

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.319830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:47.201746Z digest=sha256:34470120a34a3f3e1ce1bf891fc2b89688c414d148a7ac9f02994301d62beb0c

Observation 9820c467-ba5a-4440-aa1b-2bf2e57f7dc5 · outbound

This paper cites Implicit neural representations for variable length human motion generation.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Implicit neural representations for variable length human motion generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:02.098770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:47.271744Z digest=sha256:6971bf58eee3498acd69da78da34b6b4c13785341fb2b5c0ecc09d2b88058e52

Observation 376615c3-cc43-41f0-9c2c-c1b895599a2c · outbound

This paper cites Human Motion Diffusion Model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Human Motion Diffusion Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.444311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.444311Z digest=sha256:218be8ebb63185b6a857139813c9454524320767b45193aaebd2abca2803dd2f

Observation 03f7e6d3-2a41-4896-8caa-e2297dc1c84a · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motiondiffuse: Text-driven human motion generation with diffusion model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.619490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.619490Z digest=sha256:6df193a37610c0870ebd5d65fb3fd7b4c9da0f93dc3465bbcda06a848de8511a

Observation 9ea47c99-d57c-4a00-904f-4631a1ac504c · outbound

This paper cites Home action genome: Cooperative compositional action understanding.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Home action genome: Cooperative compositional action understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:47.780696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:47.780696Z digest=sha256:24387c7aa99e283b05db0769f65bdbc3380ed5b67a3dfd385cd2aedc0ba44600

Observation b4758825-c618-4953-bb34-bb8a00c56d74 · outbound

This paper cites Babel: Bodies, action and behavior with english labels.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Babel: Bodies, action and behavior with english labels

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.923957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:47.926998Z digest=sha256:c34dcb0dd96ae3276acb292eeb4ab9a2de8cd9875f9fbe8645507cda26d90348

Observation be6c8ef9-7a1f-4292-a862-7d67243e69d2 · outbound

This paper cites Mining actionlet ensemble for action recognition with depth cameras.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Mining actionlet ensemble for action recognition with depth cameras

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.710745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:48.054445Z digest=sha256:329e66938c83e0191615d14b089e4c418d24685f6e92f98e5f017a54bf893836

Observation 92a5830b-19b3-47c5-a2b6-c13a1da481bb · outbound

This paper cites Action recognition based on a bag of 3d points.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Action recognition based on a bag of 3d points

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.495068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:48.182938Z digest=sha256:19358b2148642a7bae6d87a5d3d0bd8afeb7e34ffe801e4f258b7ad4a684f1ed

Observation 6a2104d7-8a74-41c3-be0e-22d7490af9a9 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:48.439467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:48.439467Z digest=sha256:2a714828f70db28e9c81457b7f5b51460710454df6649a08497a80b03833c0a2

Observation d4bca55f-f114-4a84-8c6b-b03b8097020a · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Modeling temporal structure of decomposable motion segments for activity classification

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.314815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:48.662758Z digest=sha256:1154899e89a8475e314f8b756ebb9c80a653e9712ae9f0ce1fe0ba5a5bb0486a

Observation c58ba53b-e930-46e4-8461-78f9e80d71dd · outbound

This paper cites The Kinetics Human Action Video Dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding The Kinetics Human Action Video Dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:48.737903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:48.737903Z digest=sha256:30fe6854a87be66713eef9603b7361a7a11b3b4823b1a85430271b812494e824

Observation 92b440a4-e734-4821-a45c-8212ea0819ed · outbound

This paper cites Actions in context.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Actions in context

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:01.161871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:48.924740Z digest=sha256:017d7a9d7506b63ef046feb84c99300edc561a9d850959e4f4d48072cc20d160

Observation 164e9ab0-c904-4bf8-8d68-3c8a9b41626b · outbound

This paper cites Auxformer: Robust approach to audiovisual emotion recognition.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Auxformer: Robust approach to audiovisual emotion recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.946957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.076536Z digest=sha256:283019ab694b76192a9e8a7b07d16cbab082551507eaced81daf06e3eae71a28

Observation 7ead9bb5-7390-4fe1-8aaa-50dc25b53977 · outbound

This paper cites Context-based inter- pretable spatio-temporal graph convolutional network for human motion forecasting.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Context-based inter- pretable spatio-temporal graph convolutional network for human motion forecasting

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.804751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.207068Z digest=sha256:4c0756692dca68d5c74c421668a02d97b5beab25cf0e979e9f5704c1aab5e9e7

Observation c0391ee2-d79c-45bd-99b2-d74667c2bae5 · outbound

This paper cites Gcnext: Towards the unity of graph convo- lutions for human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gcnext: Towards the unity of graph convo- lutions for human motion prediction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.618553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.340159Z digest=sha256:0c3c3f6ecb07719859071615813882043cbecc20b52cca78a143e00ebc3333b2

Observation 2f2d564a-098b-42ee-927d-698371054ade · outbound

This paper cites Back to mlp: A simple baseline for human motion prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Back to mlp: A simple baseline for human motion prediction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.443157Z digest=sha256:3bb2347516a6b69ae940c4d0ba43b9177670e2b3253bba52c5c6723966e0cd85

Observation ed0b00e1-56de-4771-950a-999aac11ce18 · outbound

This paper cites History repeats itself: Human motion prediction via motion attention.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding History repeats itself: Human motion prediction via motion attention

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:50:00.194250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.578637Z digest=sha256:b98fcf27c8c9c4b3b4d999b13ff270c96351a70d04a3a812313a657af87274ef

Observation 85d4f0b6-6f2d-4329-bc6b-378e3b4f8379 · outbound

This paper cites Trep: Transformer-based evidential prediction for pedestrian intention with uncertainty.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Trep: Transformer-based evidential prediction for pedestrian intention with uncertainty

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.936136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.698402Z digest=sha256:2162ceee2e051b8013622c704baf780ba6ec7371b56fba8d26553b5ab64672be

Observation 2e0955c3-2a8a-4fe0-bb23-cbb3bf827409 · outbound

This paper cites Visual–motion–interaction-guided pedestrian intention prediction framework.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Visual–motion–interaction-guided pedestrian intention prediction framework

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.759578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.805932Z digest=sha256:dac330378e64a01f418a8bfc805a902f7d786527f71e17168493fa77c85bb7d4

Observation f0fe421d-6901-4329-b2f8-aab7dc2edacd · outbound

This paper cites Benchmark for evaluating pedestrian action prediction.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Benchmark for evaluating pedestrian action prediction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.547880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:49.841996Z digest=sha256:0dd754e9cb795277b527eb57be0dbbd3d7600f174076b9c48f26c9419ed089b4

Observation 8ceeceaf-1327-4ba8-bea8-02410bca2679 · outbound

This paper cites Agreeing to cross: How drivers and pedestrians communicate.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Agreeing to cross: How drivers and pedestrians communicate

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.335344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.076568Z digest=sha256:12e94cd4fce9ed7eaccffb359cb01b668a3b951f838de4d1704a576c4cc7f396

Observation 179f464a-4f06-48cc-8735-37ee93e2c362 · outbound

This paper cites Pedestrian intention prediction based on dynamic fuzzy automata for vehicle driving at nighttime.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian intention prediction based on dynamic fuzzy automata for vehicle driving at nighttime

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:59.134669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.298586Z digest=sha256:a2273534947cfc13d51802095256484a48f4b7500bb5eed8eaa7fa8e6c100662

Observation 7fa810cf-0bc2-4a02-b7fd-52823e3d55f3 · outbound

This paper cites Pedestrian intention and pose prediction through dynamical models and behaviour classification.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian intention and pose prediction through dynamical models and behaviour classification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.977044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.432971Z digest=sha256:812290510d8e845e1934e8a817aadf8792f78ca45977f68357d73b9c6b65e2f5

Observation bea57fc8-781b-4767-99e4-422fe8c50f5d · outbound

This paper cites Pedestrian path prediction with recursive bayesian filters: A comparative study.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pedestrian path prediction with recursive bayesian filters: A comparative study

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.799103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.558560Z digest=sha256:919ce1cca05e292a50ac92423d02a8ecba6f14669f0efb1a9d1e17a1cfd56fd2

Observation 51bebc89-e5e6-4884-b173-a92fa96ceabd · outbound

This paper cites Dolphins: Multimodal language model for driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dolphins: Multimodal language model for driving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:50.712847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:50.712847Z digest=sha256:22d5ef9fb2a08f47e69c793434009ddc4aa6b657829d4ab982c3d8dfbb24c6d7

Observation 2d51aa2e-fb38-43af-8dac-c2b5fcf3ba25 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Drivelm: Driving with graph visual question answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.596482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.803207Z digest=sha256:9ba65e7c9bfdade0e7314bb6a4bdb7f5ab94127f62e30cd12c5a652bb0006222

Observation 10676d0d-7943-42c1-9a00-4e7cbe2b0cd8 · outbound

This paper cites Tem-adapter: Adapting image-text pretraining for video question answer.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Tem-adapter: Adapting image-text pretraining for video question answer

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.389195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:50.879641Z digest=sha256:5e3471f6451bfe7be58c5cf228bc5d428738e1c6aa402e7173302dd81728cfdd

Observation e6c44bc0-0547-4c9c-bdfb-b981eedf1829 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Lmdrive: Closed-loop end-to-end driving with large language models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:50.978757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:50.978757Z digest=sha256:3bbdcecd898447d8aa77a8b1933a6878e095c8e6fdea59505c708f11e2aee13b

Observation 0d2be7b6-6737-4ed4-ae54-7934b12a7f6f · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.079501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.079501Z digest=sha256:bdf96e0b8a91066d7157c40fda0e970b1dc29246d65728a82499356394515960

Observation d39b9996-86f6-481e-8701-3389039ebcf5 · outbound

This paper cites Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi- modal large language model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.158814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.158814Z digest=sha256:0c03efa8db7ecd4895a22541ad0d627018737aceee27d837c4ad738b4d79493a

Observation 68838d8b-e6f7-43f1-872d-d54892d9d338 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:58.167373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:51.254950Z digest=sha256:ebddf1d8d81b043e67af6a81ae1ac1ab868ecde80156e21326106380dd32bbd7

Observation da9bb143-f317-4cbe-8537-f40b3002c5e2 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Lingoqa: Visual question answering for autonomous driving

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.984400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:51.438008Z digest=sha256:5e26dd6d2c1d284a9001099640afe82aabc2672e537b27c259238fc329fd601c

Observation fd6b48ab-d441-4973-8c40-d5b1556c37ac · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Covla: Comprehensive vision-language-action dataset for autonomous driving

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:51.633873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:51.633873Z digest=sha256:f9b99a30d57987e78fa44c0b7240696ff88305a8f088037c45d68a617cf7002b

Observation 3e784a95-c9bb-45c9-9b7d-d40f50b91555 · outbound

This paper cites Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Rea- son2drive: Towards interpretable and chain-based reasoning for autonomous driving

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.751162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:51.737098Z digest=sha256:8afc8f9f4612fe64f9f4e9f3d4e4d6a33df189db366fa9ea9fa44511f7241d21

Observation 607d30e9-3a00-4aaf-9820-30d0b87c37e2 · outbound

This paper cites Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.500537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:51.910116Z digest=sha256:570d1afb93b394a574ec58fc9a20a05c013287769215d0aa9ac9f9e2645b2e9b

Observation e5f7465c-63d7-4549-83db-691b68c40cc8 · outbound

This paper cites NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.051367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.051367Z digest=sha256:4833799d1cb113501db2c7b7ad3df28ab8ab22b6fdce29a15ccf44e2c7a7291c

Observation a79d6fa2-8fd4-4d0e-bd84-197aed7768f3 · outbound

This paper cites 1 year, 1000 km: The oxford robotcar dataset.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding 1 year, 1000 km: The oxford robotcar dataset

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.288537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.224282Z digest=sha256:bc862f254622fe6bc14688e1044df4beaf637a4306a7b9390a859f7254b3c7f6

Observation 9bb8aba5-a4b5-4106-9d54-f1f3eba09c6d · outbound

This paper cites Dada: Driver attention prediction in driving accident scenarios.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Dada: Driver attention prediction in driving accident scenarios

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:57.106647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.296701Z digest=sha256:5acc70ff6617c8d145c9b8f18f5deb1c45e3f2a60fe5b18d512465afaf94e9dd

Observation ef3aa4fc-2030-4b75-91b9-11c50c2c4fe4 · outbound

This paper cites Snow removal in video: A new dataset and a novel method.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Snow removal in video: A new dataset and a novel method

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.933123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.391187Z digest=sha256:79b0bc12d7f17254a293b7d346057922af08a4cb81d187ff979d441e886dc6bc

Observation 633e92b7-df81-4138-bd02-74811c27fd65 · outbound

This paper cites Semantic foggy scene understanding with synthetic data.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Semantic foggy scene understanding with synthetic data

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.711029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.460177Z digest=sha256:206f599487df5d55ac5456f143bbdcd65ad24afe9fed62a560f846d7fadebdae

Observation 14e4f89f-40d7-4f5b-9d28-723b02c3b296 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.501706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.544376Z digest=sha256:7dc50479ca8794dfeccf15baebdc7fbe81c2f2ff363442403d9796243ca08821

Observation 5dfdc58d-34c0-4e05-afb3-b59c6fd130fe · outbound

This paper cites Recov- ering accurate 3d human pose in the wild using imus and a moving camera.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Recov- ering accurate 3d human pose in the wild using imus and a moving camera

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.335057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.648254Z digest=sha256:d4b966319ca0d4c62e5740649e52006f072143cca8da813733db014b92d46302

Observation d3a61028-8328-42fc-adfe-ee4e7bb0957b · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Drama: Joint risk localization and captioning in driving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.765409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.765409Z digest=sha256:6302f554a3ba70cea526f54f2dde83b8f6e85c28eb18b3484cecb409363b84ea

Observation a7cc4c9c-ef71-4152-930d-8a4d0a52d6f3 · outbound

This paper cites Posescript: 3d human poses from natural language.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Posescript: 3d human poses from natural language

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:49:56.146203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:49:52.848079Z digest=sha256:6b466d6e45271b18139a4a6de601060d91ca3560c733882a6be1002e253bc5d4

Observation d09a7f04-0b4f-465b-b0ce-31e679d39e3b · outbound

This paper cites Motiongpt: Human motion as a foreign language.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Motiongpt: Human motion as a foreign language

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:52.954611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:52.954611Z digest=sha256:d9e5ad3bfdcb6cfe72b00c1c04b22a8dfe532dc14d07c6ec49e434bfda254dba

Observation d6efac6a-c297-4a31-abc6-8c9401a22094 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.064565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.064565Z digest=sha256:ebfaa4168dc5e92f3ecc0966bfb5dc97355589ef79c3ea84acfecc2486410d62

Observation b8fa220a-9394-478e-ae1a-55896c6ea69d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.159318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.159318Z digest=sha256:c0223a93815aa7d271335f413978f283cdda021f046e53e38597fd61490e8944

Observation 716f75ef-6c74-4cd4-9f30-f5613f70e6bd · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.272808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.272808Z digest=sha256:ee537bed8150cadf15ae45d697fb216fe2ac98d626137873faf4cd7f456c93c0

Observation 7200b3b0-b926-4848-bd33-5343c9a6842c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.371557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.371557Z digest=sha256:80bffef78c1207c44abd3a8e0493ac3e0a95524fe801a148a7ea2e53b48a93cf

Observation a5b2b250-4156-402e-98a8-275f3f32c366 · outbound

This paper cites Qwen2.5-VL Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Qwen2.5-VL Technical Report

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.451178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.451178Z digest=sha256:9862dc8ad01338420edb03fff938f02ac741f766e118d5f2088051df76289911

Observation 26395a82-440b-4d02-9634-af4b864fd32f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.541236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.541236Z digest=sha256:1a6ae1e229fc6e2b8ffa55b78258c432cfbeda45558cc3815229a0f494f0c5d4

Observation 68ea932b-35ff-451a-91b8-44938f6df4d9 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.632371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.632371Z digest=sha256:f128b08f37a2204f58e498a90e26bfab99d5f074719c1eebdcdc543061e35782

Observation f3e330ce-1d74-48eb-a971-be76d1ae37d0 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions., 2024.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Building and better understanding vision-language models: insights and future directions., 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.709943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.709943Z digest=sha256:1a6fae244e6e8939a72e5440f3d1853a91b89c37e1b65a185c2a9024de361c99

Observation de930206-53e1-4297-90c8-e1418f1cbca7 · outbound

This paper cites Pixtral 12B.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Pixtral 12B

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.785315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.785315Z digest=sha256:c4c2a111e0180871d56706dd564ff6773ba3c0bd6334a6f3d52f7bf6a521e880

Observation 31e8e320-eaac-4ab4-8bc5-b55eee0c7e62 · outbound

This paper cites Gemma 3 Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Gemma 3 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.884739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.884739Z digest=sha256:e22c1740179d3b7c9f417ef2e97fbb9211c455ec419e4b6ae6118df76915ced1

Observation f4b662be-748a-4797-b619-ddf3794da39d · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:53.990949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:53.990949Z digest=sha256:bfd25322ee3d8e3b184d9145e893b739fcd34409c19a12893760e0fa09c40d9e

Observation 84beb5f3-69ff-4066-9f02-b08cae3daa5f · outbound

This paper cites Kimi-VL Technical Report.

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Kimi-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:54.072518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:54.072518Z digest=sha256:d06e0913591fd24268b996116aa154c92d4416f95ae37e3487adf085d43ebaa6

Observation f1010b81-99b5-464e-8195-5401d9bd769b · outbound

This paper cites Joint Attention in Autonomous Driving (JAAD).

MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding Joint Attention in Autonomous Driving (JAAD)

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:54.185917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:54.185917Z digest=sha256:d70418b69fd0c9c03f8638e97a0fb643305575498403c8f7b76b941c322c2394

Pith citing papers

No inbound Pith citation observations are available.