Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:12:06.172447Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.20964.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:12:06.172447Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 580181f6-c219-467d-8a29-3675c4e7e910 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Parallel Vertex Diffusion for Unified Visual Grounding,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2caa0d2-1939-4c13-b5b1-490a56ac8513 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Align and Prompt: Video-and-Language Pre-training with Entity Prompts,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e24e2e46-65b9-48d4-b2b7-2d2faf469f31 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d8a163-76f5-466d-913b-9e3a768d01db · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning FreestyleRet: Retrieving Images from Style-Diversified Queries,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1aaa9f2-4362-4779-9054-8aa9998607cd · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ef21c2-3c08-41e0-80e4-ea9205a68081 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dual Encoding for Video Retrieval by Text,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8abaa5a-d856-47b7-ad00-344a54ca18ef · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Temporal Alignment Networks for Long-term Video,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a402677-488d-4e60-b5c4-0642ff5be8a1 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadebf4e-91ec-441e-8b23-2b5dbf0299c1 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An axiomatic approach to the concept of interaction among players in cooperative games,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd351e5b-f8c8-4aeb-9b84-439383506c24 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weighted Banzhaf power and interaction indexes through weighted approximations of games,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4cecc72-b5a1-4a92-9794-f533cd60633d · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc667b79-58bb-48a1-b6d1-09c9a51b3d9e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8257d3ac-c9d2-4e89-9ffd-6045984ec663 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dense- Captioning Events in Videos,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce22814-90f7-495f-9113-89221bb66181 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Localizing Moments in Video with Natural Lan- guage,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3403c7-4884-4a80-9f49-c8750cb0a0bc · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering via Gradually Refined Attention over Appearance and Motion,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b274597a-7c10-46af-8fe2-9eb2b24caf6b · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f155bde-624d-470e-ab62-1c829b685aab · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Universal Weight- ing Metric Learning for Cross-Modal Retrieval,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 826b2159-8ae8-475c-9961-5f410fb49b6c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 187ba5fb-4864-40ca-860e-08adcb7f871e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Revisiting the ‘Video’ in Video-Language Understand- ing,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5131f6e0-6939-480c-90d9-d4f522313cd6 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Fine-Grained Semantically Aligned Vision-Language Pre-Training,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ed08a169-cf2c-4d57-a05a-1acbe1425b0c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be87c5ae-cc08-40cc-b5e0-d85657143a3c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba5899eb-b8d5-44db-9f26-2564ea2e10f4 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Decoupled peak property learning for efficient and interpretable ecd spectra prediction,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 85c7ebb0-03aa-4346-af36-21c2b3c6941c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0084ff6b-a2bd-4067-9791-3afa3253ddd3 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac6680f-dac9-4815-81fc-e833b5061aed · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233a77f0-5372-4513-8faf-0789cd268a65 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1457285d-4191-414b-bd55-2cec95ab0033 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Next Patch Prediction for Autoregressive Visual Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d71f36-46ef-48fe-8df3-c76f22863b9a · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning the Best Pooling Strategy for Visual Semantic Embedding,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c620472-2588-440c-b775-3ea6db513a70 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60c8fc4d-61b7-47e8-8edf-ed2d186a6c97 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning Transferable Visual Models From Natural Language Supervision,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46fe8fa6-37fc-4019-8dcd-858642da83cd · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc757e8e-5304-499b-b4f7-8b5485bbea67 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aba342a8-9f5d-4cfe-88f6-5fa7c2d1eac4 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Hierarchical Con- ditional Relation Networks for Video Question Answering,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d05302cd-1063-4a5e-9a41-d0d265532e07 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering: Datasets, Algorithms and Challenges,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cecd60ef-9e7a-4f28-9cc1-9663220d202e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3aee8a22-e517-41c6-8159-c847b1d072cd · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering with Iterative Video-Text Co- Tokenization,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d53e862-8086-4691-8574-2ff90815e13e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f94017d-14a8-43b1-9f46-e0e0993c33f0 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06dea930-9905-4701-b330-7647b8ec12b1 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Localizing and Describing Events for Dense Video Captioning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72c9e196-afda-4e0b-b89a-ba78e686ab17 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Captioning with Transferred Semantic Attributes,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30e5a4da-b26c-40e9-8b55-4ed14d965516 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Modeling Embedding and Translation to Bridge Video and Language,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1c5967c-c4d2-4621-b14f-94ede0299011 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ade93c9-c123-4725-b194-f045060e46e9 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07426d4a-e689-4008-9b1a-5ba4e3deb0bb · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Optical-model po- tential in finite nuclei from Reid’s hard core interaction,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01e9f7a3-a58c-4079-a4e6-1bb56496094c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a32ecb05-ee33-4dca-90a3-2989dc7ab9d9 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d89eefbc-6df6-41e0-9283-65bb44626049 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d12f0efa-572a-49ea-964e-9618a659463d · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4435afe-cd1d-43ca-aee0-7ed9b362b0b1 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b280c4e-95ac-4ae6-acd2-1ccd5c61d464 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Kullback, Information Theory and Statistics
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e03b7460-3f52-4367-9457-1b99a1dcf37c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Study on density peaks clustering based on k-nearest neighbors and principal component analysis,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee34466-8277-4045-8893-f4b3911909ad · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3afa91d-1e33-4046-8e5e-f679205bacfd · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba9f6136-6bcf-4314-ba69-d67834dda4f0 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cross Modal Retrieval with Querybank Normalisation,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 997268e8-2bf2-473c-b96c-eae4f7ab6974 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-modal Transformer for Video Retrieval,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8cce65b0-6d35-4cfb-98b6-760e8e60abd7 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70980822-f817-4cbb-88e1-50106cc2be1d · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e55b30e8-aa31-416c-8cce-50fef8ba3127 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Support-set bottlenecks for video- text representation learning,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd42f802-5095-4e2c-80d9-bb264a531cfc · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cb85fc88-9c83-4ee4-9f5d-ed059dc9c89e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 57368048-0dcc-483a-884b-ccffb54fb35d · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 133ea5e1-da19-45f1-a8d2-e39daae0b236 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UATVR: Uncertainty-Adaptive Text-Video Retrieval,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 776c9f9a-afb4-46d5-8000-cf06ec459f02 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa1b7248-99e4-48db-94ae-e21fd7fd86ed · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da4aa689-5234-4760-b1f2-2ba1ce0e7b18 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Representation Learning with Contrastive Predictive Coding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0fb8e42-135b-425f-b127-0967ae37d3cb · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Use What You Have: Video Retrieval Using Representations From Collaborative Experts,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 79c510ef-24e5-44a9-a331-f0b7d7605a6c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 427b624c-1c82-47b8-8a34-34b1837428ce · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning A Joint Sequence Fusion Model for Video Question Answering and Retrieval,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b00737e7-0ff0-4a41-a683-c3b4ddc93247 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning BLEU: a Method for Automatic Evaluation of Machine Translation,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f6780819-c2ea-4a2c-ac1b-75b072106f9e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2608b2a6-8f9d-4e4e-aa70-44dc6abe8f24 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc57212a-8e67-4c84-9fdd-f2fe437e7b9f · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CIDEr: Consensus-based Image Description Evaluation,
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c880074f-2595-4c3a-8a4b-a8fa0607d1bc · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Adam: A Method for Stochastic Opti- mization,
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d62a870f-b1c9-4696-8aba-208a5cf6ce28 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d72011aa-5ba8-428b-93d9-d17126a0a052 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Approximating power indices: theoretical and empirical analysis,
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a5e0918-8873-4beb-9cb1-3a340a0a809d · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning All in One: Exploring Unified Video-Language Pre-training,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a4b19c3e-4df0-4fe8-a0b4-93250bc684a1 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 59788f69-d0ae-421c-b8f0-aad0845698ad · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-Granularity Interaction and Integration Network for Video Question Answering,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9dceb0e7-669a-4a6c-9a09-f5cdbcb84564 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Invariant Grounding for Video Question Answering,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 03e73832-5573-4e38-a027-897810212dc9 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning to Answer Visual Questions from Web Videos
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6bbc903e-83ad-4f64-a7e0-0a5b23d8fcc3 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering With Semantic Disentanglement and Reasoning,
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 87fc9f50-ec22-48ff-8463-251123d092c2 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SViTT: Temporal Learning of Sparse Video-Text Transformers,
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1503c78-c126-43f8-8679-bb33c5a59b99 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TG-VQA: Ternary Game of Video Question Answering,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eef74904-de05-42e9-a437-f6a71b29b991 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a4ac961-06e3-4e5a-b2be-fe544deeec82 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning End-to-end Generative Pretraining for Multimodal Video Captioning,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e5f9d9c6-66ce-4ac3-b6ef-64bcf0c1f83a · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Motion Guided Region Message Passing for Video Captioning,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 49ad36ec-6710-4eed-85f8-b9f6df6b9bd7 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Open-book Video Captioning with Retrieve-Copy-Generate Net- work,
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cffd5d3-dc94-43ff-802b-23f791be790c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Attentive Visual Semantic Specialized Network for Video Captioning,
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6bcd28a6-c000-49e9-98c4-a62a84bc9b9c · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Global semantic enhancement network for video captioning,
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b245349-1528-467a-a873-cb780712c22e · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Accurate and Fast Compressed Video Captioning,
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5356a222-97a4-4788-8cd2-56edf440d795 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Emotional Video Captioning with Vision-based Emotion Interpretation Network,
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 989e1414-4faa-433e-81f7-9d43588b5c27 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2b4da795-3272-4e61-94e2-545790861b51 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Caption: CLIP for Video Caption,
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 44183826-f439-4dc4-baa9-53623641cdf6 · outbound
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Visualizing data using t-SNE,
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.