Pith. sign in

Paper Citation Record · LEDGER

VUDG: A Dataset for Video Understanding Domain Generalization

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.24346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24346 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:13.889873Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac25e71-3ae8-449b-9bd8-4267519269bd · outbound

This paper cites Mm-vit: Multi-modal video transformer for compressed video action recognition.

VUDG: A Dataset for Video Understanding Domain Generalization Mm-vit: Multi-modal video transformer for compressed video action recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.898598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.169622Z digest=sha256:e1b9e770e7b43193a345c0dacadc18995b6b9c9a207525f737f4e472ca1c5b24

Observation 408e9a2d-3b76-43e1-8d19-c0db8a1f8b40 · outbound

This paper cites Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Mar: Masked autoencoders for efficient action recognition.IEEE Transactions on Multimedia, 26:218–233, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.765665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.271553Z digest=sha256:7e0dbe1d40dd9892f62f9dca471bcb7be09b5630dbef9d350f5be3797db84592

Observation d3ca7c93-d783-4fc9-8c8f-d91f16364647 · outbound

This paper cites Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Mnv3-mfae: A lightweight network for video action recognition.Electronics, 14(5):981, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.342343Z digest=sha256:3ba60f11e2eb7abf5e05f85f493327275f7bbc28b60b279fa82238dd33bacf7d

Observation 0f4a7285-0452-42b3-a448-a88371dc3454 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Swinbert: End-to-end transformers with sparse attention for video captioning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.266711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.421238Z digest=sha256:23c5d869325bd36c2c595a20c610149f6fb67802b3766eb41de8455e8b666d40

Observation d18ed918-683d-4e62-9802-8732f95e455e · outbound

This paper cites End-to-end generative pretraining for multimodal video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization End-to-end generative pretraining for multimodal video captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:18.135728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.512771Z digest=sha256:128fd38437b67151bc6629519a5ef11b7a737fe3c06312d11b320925f02ffece

Observation 2dd63dcc-8a37-4eba-92da-5b710be2637a · outbound

This paper cites Text with knowledge graph augmented transformer for video captioning.

VUDG: A Dataset for Video Understanding Domain Generalization Text with knowledge graph augmented transformer for video captioning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.924672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.605822Z digest=sha256:061e8a01e746c85967fa8934f06eac21dfe0d44fbe147c018d3f556a7bc4918a

Observation 3c44d19c-02c3-41b2-b73a-b3873a620ef5 · outbound

This paper cites Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Automatic video captioning using tree hierarchical deep convolutional neural network and asrnn-bi-directional lstm.Computing, 106(11):3691–3709, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.778516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.711402Z digest=sha256:ff0832a285d8cd2d5368b79684f1c027b7591a4a13a3a42a48eee7355b2b545e

Observation da4843da-867d-484e-a70b-f61e44a77040 · outbound

This paper cites Invariant grounding for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Invariant grounding for video question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.551486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.807715Z digest=sha256:3ab33e62b3b50206eefc728c4889faf137c6441ab0bbe0ee54ff48e41c949c80

Observation a94c8fd1-48a4-4d3b-b56e-525276fc14cb · outbound

This paper cites From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

VUDG: A Dataset for Video Understanding Domain Generalization From representation to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:17.195284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:09.885422Z digest=sha256:5e929c893f25645f9e3c74a830792bd076d4c6957e2bbe9aa2f071270737cf67

Observation 39c95ccd-7480-4c1c-947a-a64ad30d46af · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Morevqa: Exploring modular reasoning models for video question answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:09.991658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:09.991658Z digest=sha256:e2eadc2b93d34013d146fd9949b11f6fa0b0e97d5fbeb0b5d7b9045833d57fa3

Observation 719fcd8c-da12-49b7-952c-787b07e52646 · outbound

This paper cites Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Diversifying spatial-temporal perception for video domain generalization.Advances in Neural Information Processing Systems, 36:56012–56026, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.929225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.103329Z digest=sha256:2f8bd3b32c92362f1baf319bfc471900247f0081a6ce7a225f93d321624083b0

Observation c6ae223e-649a-476b-b129-640cd0b54210 · outbound

This paper cites A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization A multi-modal egocentric activity recognition approach towards video domain generalization.Sensors, 24(8):2491, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.172317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.172317Z digest=sha256:41cb4797bd02776420ab8bf1de81f4b4c059ee77b92e22dff6afdaf6b39ef238

Observation 9aa4dada-0a55-4abf-a65a-93b597e49f3f · outbound

This paper cites Meta-causal learning for single domain generalization.

VUDG: A Dataset for Video Understanding Domain Generalization Meta-causal learning for single domain generalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.764934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.246214Z digest=sha256:268857d8c0dbaaae661c3845255e623ba1254c84b73967e54055481c17e2ef2e

Observation 5de17fab-d8c3-44db-9614-aade6df7d16b · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.645225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.308652Z digest=sha256:5adf23247c8918400ba14fb96db2a4ba34079ae5db04842b4e55ec35683567d4

Observation 3642f322-b4d7-420c-99d3-6c0c4b4e89ee · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VUDG: A Dataset for Video Understanding Domain Generalization Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.406120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.406120Z digest=sha256:8e453ae572eb401a821f33cd25e134a281c86b6030d1f29008677daa595813d4

Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.535033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.535033Z digest=sha256:28a5b3b174e37db7e2a0d7c71ce3b0769c0e11ffa17211d75dff5ee859216fbb

Observation 4d7565d3-43ba-40ea-bf00-649e69218ad5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VUDG: A Dataset for Video Understanding Domain Generalization Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.659981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.659981Z digest=sha256:5d74643d2915d607acbaec977eb875c1cafc718b452ddb2dd5f78f6022cbc5d5

Observation 04193e7e-8181-459f-9af9-e29e13903829 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

VUDG: A Dataset for Video Understanding Domain Generalization Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.532575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.737475Z digest=sha256:3462cb2ba72f6ddb2c52eea005f088381ab4f9817405d36f2bec8a5f61d2bc87

Observation bfef86a0-5c3d-43d5-acfa-ccba42c555b1 · outbound

This paper cites Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:32:14.306658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.839221Z digest=sha256:375f35d989926d90fa291e733e7de928ed1b0952fc47aa50fe3f1d6d1445a77e

Observation f001e191-3bef-465f-873e-5f0544e86273 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

VUDG: A Dataset for Video Understanding Domain Generalization Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.400235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:10.912842Z digest=sha256:95258855cb2ad246392ee61b909ece96d8f7ba0562fa83df7e72dcf287ff7304

Observation ca3c054d-30ba-46a5-a16e-d55863fa1781 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.012551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.012551Z digest=sha256:389ab729320cca974473be2dd0077b37bbe17ee309904902898bba1c5f94b239

Observation aa8fcc48-396c-45eb-8395-952d6e62d9e5 · outbound

This paper cites Qwen2.5-VL Technical Report.

VUDG: A Dataset for Video Understanding Domain Generalization Qwen2.5-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.135409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.135409Z digest=sha256:f936863b913f26601cde143403a36f0d04491a71a8ebd24e9b79a7a4dcfebe41

Observation 47e1aad9-c1a1-46fb-9736-104483cb5b61 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet: A large-scale video benchmark for human activity understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.277723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.277723Z digest=sha256:a5acdd2f9b500304e8a1138ae7961868bed40a0b2c143dca40a4b9d7d2777f40

Observation 0f30c54c-a707-413a-acfd-957ef53fdc5a · outbound

This paper cites The Kinetics Human Action Video Dataset.

VUDG: A Dataset for Video Understanding Domain Generalization The Kinetics Human Action Video Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.404740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.404740Z digest=sha256:1ad6d821d7657d5bea50e078ba20bfd6d383da244e7cdaabd088a2bb8771d6a4

Observation 12e75b38-7b5e-4f43-91bf-3494471c519f · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

VUDG: A Dataset for Video Understanding Domain Generalization Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.537568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.537568Z digest=sha256:3cd93c3c4f4ea8de05b31aaec9a17affecfd62b2dfe027a43591aab386aa1fc0

Observation 3daff457-2f87-4fc7-b2e4-9a7c19385452 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

VUDG: A Dataset for Video Understanding Domain Generalization TVQA: Localized, Compositional Video Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.633968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.633968Z digest=sha256:ab93630eaf8afc5e39c0c44b23c369172219f36613fb0fb86688c84280c35de2

Observation 3abd516f-781c-4d5d-bd2e-5341b2b2387d · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VUDG: A Dataset for Video Understanding Domain Generalization Video question answering via gradually refined attention over appearance and motion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.713809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.713809Z digest=sha256:84449368cb7fca353302cede7b6a553fc4833740680d91d477fe536b3e1271e2

Observation fe10182d-2607-4e6b-a8cc-aa468c18279e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

VUDG: A Dataset for Video Understanding Domain Generalization Msr-vtt: A large video description dataset for bridging video and language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:11.808333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:11.808333Z digest=sha256:05aeba093c4e002825dcf89c558cbff237f115ebfa794b9af879841fd29b5fba

Observation a479c01a-fbd6-4be4-90ed-34515963f71f · outbound

This paper cites Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Domain generalization for video anomaly detection considering diverse anomaly types.Signal, Image and Video Processing, 18(4):3691–3704, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.181783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:11.907963Z digest=sha256:27e923e1dab01144e5b05a8a3de1d2a1bc5d56abe1ec8ed4e7d7ba0366968e03

Observation afd06ccc-503f-4e77-90a8-fad64cd4d8f5 · outbound

This paper cites Video-audio domain generalization via confounder disentanglement.

VUDG: A Dataset for Video Understanding Domain Generalization Video-audio domain generalization via confounder disentanglement

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:16.076518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:11.958263Z digest=sha256:0266862021fb170fc20511684ec6a6b694361778abafc9758a981dfbb1f05559

Observation e8ae2171-134a-421e-a12a-96bcb2751ee4 · outbound

This paper cites Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021.

VUDG: A Dataset for Video Understanding Domain Generalization Videodg: General- izing temporal relations in videos to novel domains.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7989–8004, 2021

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.938790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:12.051541Z digest=sha256:458c83edc44f63cdfdb67a29bd06d6e2ae57e6aabc9f2213c3298f7df1b20dc4

Observation 83944137-ba67-43fa-80a5-680fe3cc3151 · outbound

This paper cites Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Ani-gifs: A benchmark dataset for domain generalization of action recognition from gifs.Frontiers in Computer Science, 4:876846, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.819775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:12.169976Z digest=sha256:8e92d0b9c098fbdb518b4b6a91bac0923009632fdb9730df5d9b6bf0308de41e

Observation 88bec4d3-31fa-4d79-bf07-42b717047a3a · outbound

This paper cites What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations.

VUDG: A Dataset for Video Understanding Domain Generalization What can a cook in italy teach a mechanic in india? action recognition generalisation over scenarios and locations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.731642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:12.279011Z digest=sha256:178725977bf50056c4b94c252b5955d3f68a52b1a5f89a7a87d02ee46137a0b0

Observation 3173452b-aaee-44b7-8dec-102394f24a74 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VUDG: A Dataset for Video Understanding Domain Generalization Ego4d: Around the world in 3,000 hours of egocentric video

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.349303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.349303Z digest=sha256:2244a9ec5366445eebb474c448bf4e53087550ef59a771bdafa6602830f10a1e

Observation 50943904-c01d-455a-9c31-7a1f84084fe4 · outbound

This paper cites Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection.

VUDG: A Dataset for Video Understanding Domain Generalization Multimodal motion conditioned diffusion model for skeleton- based video anomaly detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.579291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:12.469345Z digest=sha256:82bc314844dfc0c417f7313a7aea084206d5d2af4ffd12d8fe9a838f8fe8f970

Observation ea7fe981-ba3d-46fc-b507-43ec8e51e7b8 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VUDG: A Dataset for Video Understanding Domain Generalization Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.589851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.589851Z digest=sha256:e1cd29b789ad71bb120557bff28239c7f3e44674272a4b0b451693bd57276f9c

Observation 1acd7fa0-0319-4d83-a723-8ff7b95ac313 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

VUDG: A Dataset for Video Understanding Domain Generalization Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.660743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.660743Z digest=sha256:147637cc1cc1bff3eec3d1ffa922c90f2b0eaa25df5c6ba044daabd5923106d0

Observation dec74c2f-daa6-4af3-993d-154d4e839773 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.770007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.770007Z digest=sha256:873a0e8347389010f0f1e9deadcbf51c8df2305b08cf2a4be4718fbd07de2c9e

Observation 57437514-aca4-4203-9fdb-6bd9ed8012ea · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VUDG: A Dataset for Video Understanding Domain Generalization TempCompass: Do Video LLMs Really Understand Videos?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.883906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.883906Z digest=sha256:43a18df82ce02f3fa1ae1af65dbcd6b2a2ff160f93c89b9f0b662aee9244ed1e

Observation e93b29a3-6889-478a-9775-24288b4fe2fe · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.990057Z digest=sha256:03350ea1417d3cade55cbfd3d1bbe82f7ceacda2b46381cf5ff630660f572002

Observation b0e01d63-607a-4dd7-86f7-e68e3e54acb9 · outbound

This paper cites Poem: polarization of embeddings for domain-invariant representations.

VUDG: A Dataset for Video Understanding Domain Generalization Poem: polarization of embeddings for domain-invariant representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.456435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.106658Z digest=sha256:229da78ceae988e6ca42ad74ce836accc5f23fd25a5e89209be27c70db9e56c7

Observation de7d5e04-b281-46b1-8e1c-6643da3feda6 · outbound

This paper cites Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning.

VUDG: A Dataset for Video Understanding Domain Generalization Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.297134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.180703Z digest=sha256:c13824a519a4aba35c76c6e97fb00e376b2a4978826ab5261509b866be68ba7a

Observation f3122d34-d08f-4471-b2fb-b56761562758 · outbound

This paper cites Clifton, and Jie Chen.

VUDG: A Dataset for Video Understanding Domain Generalization Clifton, and Jie Chen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:15.126826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.233194Z digest=sha256:92d33bb990cbd283d478b42d81b5b068b4e06fa7d1eff9452c1cd5007e1f7a55

Observation 88d004b5-88ee-426f-90fa-505684e4b023 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VUDG: A Dataset for Video Understanding Domain Generalization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.326003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.326003Z digest=sha256:35c41cb6e0f5d35d0be231847436b23b057cb5247b17c664e85632627393bff5

Observation b324e589-68a1-4714-8d08-3deb4559a04d · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.398642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.398642Z digest=sha256:4bf631cb25e169cdffa802bea8a2a4d6fcf8b3aac034de62ba8ac7597a1e9e45

Observation 7ddc9ed9-fd2d-4530-ba5f-289c8eae72c8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VUDG: A Dataset for Video Understanding Domain Generalization VideoChat: Chat-Centric Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.451229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.451229Z digest=sha256:e3b8b331ed9dcd714c74d7f3a4a254bb7d08441b7374dfea5fbe5d354f5610a9

Observation 52315005-d983-45bb-afc4-a0374560c0ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VUDG: A Dataset for Video Understanding Domain Generalization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.505372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.505372Z digest=sha256:e755d4e7b0871f8dc1511978742fa171d97632893f19bd538384eb4546973cbd

Observation f53ff2bb-efdb-4858-b175-7b01a95b1168 · outbound

This paper cites mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models.

VUDG: A Dataset for Video Understanding Domain Generalization mPLUG-owl3: Towards long image-sequence understanding in multi-modal large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.916081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.548397Z digest=sha256:8d666f598fa7b471fbe82d175d3092a775ffc26f64f19ecde184c82684e5d0a4

Observation 56e612ab-1b36-4961-b76d-e4097da982a5 · outbound

This paper cites Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024.

VUDG: A Dataset for Video Understanding Domain Generalization Video-ccam: Enhancing video-language understanding with causal cross-attention masks for short and long videos, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.721883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.607365Z digest=sha256:fa168aee0212da78a513d1decda064559f320368f6c21b51a34a9b242cb87488

Observation e5057a45-94c8-41c0-9308-1e4fcef91d8a · outbound

This paper cites Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025.

VUDG: A Dataset for Video Understanding Domain Generalization Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.692363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.692363Z digest=sha256:16250e4a2451ad89582e80883adeb7411d165e38c53fbbecd2e18a606a89a83a

Observation 1cd6c106-ca4c-4be1-9780-1baf1543045a · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

VUDG: A Dataset for Video Understanding Domain Generalization Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.775272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.775272Z digest=sha256:c9cbadc5be5ba48c5be469675277b4402fb6ad8ebeff081b0b6ca0e1950c87a0

Observation 40ee116f-83bc-4a1e-8d60-984fb2908a3f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VUDG: A Dataset for Video Understanding Domain Generalization An image is worth 16x16 words: Transformers for image recognition at scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.822308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.822308Z digest=sha256:f39b64dea7f77f84e5c98273fa077935a4c04b057b9c90ec6b9f1a1bbde2c4c1

Observation b7f291b6-cb27-4868-aac0-081effaf01d9 · outbound

This paper cites C- pack: Packed resources for general chinese embeddings.

VUDG: A Dataset for Video Understanding Domain Generalization C- pack: Packed resources for general chinese embeddings

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:32:14.508279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:32:13.889873Z digest=sha256:1329cca01da6f529c8c5050ba8343986e220908617e038792a6c791a49f73c52

Pith citing papers

No inbound Pith citation observations are available.