Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

As of 13 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.20964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20964 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:12:06.172447Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact1
  • verified fuzzy71
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 580181f6-c219-467d-8a29-3675c4e7e910 · outbound

This paper cites Parallel Vertex Diffusion for Unified Visual Grounding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Parallel Vertex Diffusion for Unified Visual Grounding,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.723321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.723321Z digest=sha256:bc59639d47763c6ef2e3be5cf25a39042f71b0269cae868d7adbc29fb3b1d98b

Observation e2caa0d2-1939-4c13-b5b1-490a56ac8513 · outbound

This paper cites Align and Prompt: Video-and-Language Pre-training with Entity Prompts,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Align and Prompt: Video-and-Language Pre-training with Entity Prompts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.729502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.729502Z digest=sha256:8e893d6b81cc557c94c8c742fc04817dbb5bd3d38c3b1b101b04b1ca844d1aef

Observation e24e2e46-65b9-48d4-b2b7-2d2faf469f31 · outbound

This paper cites Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.734526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.734526Z digest=sha256:d242522986d4e740083822fb06c69df9910923b1d4718ef19ced496e6f622d38

Observation 01d8a163-76f5-466d-913b-9e3a768d01db · outbound

This paper cites FreestyleRet: Retrieving Images from Style-Diversified Queries,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning FreestyleRet: Retrieving Images from Style-Diversified Queries,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.740405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.740405Z digest=sha256:bb9497086ad59f792658d5155858b5a526f5c8305b8f4c5b862607a9e33d6424

Observation b1aaa9f2-4362-4779-9054-8aa9998607cd · outbound

This paper cites Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.746306Z digest=sha256:d50135d9fe057f0cf1f1f9699cf79404e78a4de1336929a7626cef3ad37751bc

Observation 29ef21c2-3c08-41e0-80e4-ea9205a68081 · outbound

This paper cites Dual Encoding for Video Retrieval by Text,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dual Encoding for Video Retrieval by Text,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.751937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.751937Z digest=sha256:49ad52952e4e6b2c17264269743beafac5dd8a95d14f980d39da32ec3acb8579

Observation b8abaa5a-d856-47b7-ad00-344a54ca18ef · outbound

This paper cites Temporal Alignment Networks for Long-term Video,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Temporal Alignment Networks for Long-term Video,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.757710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.757710Z digest=sha256:d65f7705bcfd914e27315d69e2883e231697731b2ff0441e0e71cbb56678d5f2

Observation 9a402677-488d-4e60-b5c4-0642ff5be8a1 · outbound

This paper cites Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.762521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.762521Z digest=sha256:7ac462ba91ebead072a4cc842b4ecb90798402a59ddbc2df07acf6b346fe00b5

Observation fadebf4e-91ec-441e-8b23-2b5dbf0299c1 · outbound

This paper cites An axiomatic approach to the concept of interaction among players in cooperative games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An axiomatic approach to the concept of interaction among players in cooperative games,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.767544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.767544Z digest=sha256:a9494131f66478fe0063439c69ec501e4a7435d82cdd527be6fb798ae7dc4a83

Observation dd351e5b-f8c8-4aeb-9b84-439383506c24 · outbound

This paper cites Weighted Banzhaf power and interaction indexes through weighted approximations of games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weighted Banzhaf power and interaction indexes through weighted approximations of games,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.772384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.772384Z digest=sha256:9dbee6c781fcb34e83a04a5520624b08909c80ccbbc96e8ce09f698dd144b4cc

Observation f4cecc72-b5a1-4a92-9794-f533cd60633d · outbound

This paper cites Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.776983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.776983Z digest=sha256:1e20fb85e0592d669618d695e21404158c7667de88928a4843be8fd71152b92e

Observation dc667b79-58bb-48a1-b6d1-09c9a51b3d9e · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.782331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.782331Z digest=sha256:d686b27a9e3c63f87e27e0b5d9922bf52c2f5a807ff2d784c79cdbf944aebd88

Observation 8257d3ac-c9d2-4e89-9ffd-6045984ec663 · outbound

This paper cites Dense- Captioning Events in Videos,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dense- Captioning Events in Videos,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.787122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.787122Z digest=sha256:9ae6bf701e073cd3998c1134b716c874b3801b1c54cf0af7d5d5ce956192b97e

Observation fce22814-90f7-495f-9113-89221bb66181 · outbound

This paper cites Localizing Moments in Video with Natural Lan- guage,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Localizing Moments in Video with Natural Lan- guage,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.791722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.791722Z digest=sha256:025647a6e47f1affb7b025561cf61fe3c58bb977b23efdf9ad1944e51294d789

Observation de3403c7-4884-4a80-9f49-c8750cb0a0bc · outbound

This paper cites Video Question Answering via Gradually Refined Attention over Appearance and Motion,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering via Gradually Refined Attention over Appearance and Motion,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.796656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.796656Z digest=sha256:74c86c67548d344aec503aa76738c77c0436e569049962882b4baf40c3a8e4be

Observation b274597a-7c10-46af-8fe2-9eb2b24caf6b · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.459256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.801152Z digest=sha256:bc15ed38d5788f80c7966671d7d07e9288637d813dadebff22f5fd9c5102144d

Observation 0f155bde-624d-470e-ab62-1c829b685aab · outbound

This paper cites Universal Weight- ing Metric Learning for Cross-Modal Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Universal Weight- ing Metric Learning for Cross-Modal Retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.442279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.805550Z digest=sha256:e8f2e3be92b31e4833154d52b93471c1bfd10de96e0444b145b875581bacc6ae

Observation 826b2159-8ae8-475c-9961-5f410fb49b6c · outbound

This paper cites Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.426202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.810678Z digest=sha256:d3ff9ec2b1a43ba4749581be1a50e2658db6517f8930625cad7515854173b16a

Observation 187ba5fb-4864-40ca-860e-08adcb7f871e · outbound

This paper cites Revisiting the ‘Video’ in Video-Language Understand- ing,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Revisiting the ‘Video’ in Video-Language Understand- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.409207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.815304Z digest=sha256:98af51736e3ac8f69280c0909e23c7d69e96c12e5c948f911d3088961d09615a

Observation 5131f6e0-6939-480c-90d9-d4f522313cd6 · outbound

This paper cites Fine-Grained Semantically Aligned Vision-Language Pre-Training,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Fine-Grained Semantically Aligned Vision-Language Pre-Training,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.391727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.820300Z digest=sha256:8c52fca250ed04bb44dbb58932ef9492130ccfbd39dfaecf0a22770c3813e76e

Observation ed08a169-cf2c-4d57-a05a-1acbe1425b0c · outbound

This paper cites Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.376631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.824893Z digest=sha256:00da54a6c07fdcef2510355eca43190ea52d835aeffc57120846881ed5f5be87

Observation be87c5ae-cc08-40cc-b5e0-d85657143a3c · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.360014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.829969Z digest=sha256:c4a3568d422ba23eb565a4fdd18f9a121a122ab7136860befc053ccc21094639

Observation ba5899eb-b8d5-44db-9f26-2564ea2e10f4 · outbound

This paper cites Decoupled peak property learning for efficient and interpretable ecd spectra prediction,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Decoupled peak property learning for efficient and interpretable ecd spectra prediction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.342699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.835414Z digest=sha256:9feab421cb029f774f79219d1b600b159103cb6d4c5831b209691b052a49b3d0

Observation 85c7ebb0-03aa-4346-af36-21c2b3c6941c · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.839893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.839893Z digest=sha256:2217a633658841fcf5278d37d03ac015b0be344b12ee090c073c6693afe2f1dd

Observation 0084ff6b-a2bd-4067-9791-3afa3253ddd3 · outbound

This paper cites EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.845207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.845207Z digest=sha256:6dadb987295fe59db67991e09e2f6954687b9e2e911fce14cbbf31e0f9649d90

Observation 7ac6680f-dac9-4815-81fc-e833b5061aed · outbound

This paper cites Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.850062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.850062Z digest=sha256:27ead6831b3c168bd046c61bbcedb5a2580eb434629c5285e0083449d4d5d278

Observation 233a77f0-5372-4513-8faf-0789cd268a65 · outbound

This paper cites Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.326839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.854989Z digest=sha256:8ab2713f55e1617f6c720d2d9ce07bf0e19251303ad309108cbb52710f2da0be

Observation 1457285d-4191-414b-bd55-2cec95ab0033 · outbound

This paper cites Next Patch Prediction for Autoregressive Visual Generation.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Next Patch Prediction for Autoregressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.860105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.860105Z digest=sha256:105cfe6feb3c100a9b5696ce50edbfc75a8b0d0160b16e3f644687d3a52d26d4

Observation 06d71f36-46ef-48fe-8df3-c76f22863b9a · outbound

This paper cites Learning the Best Pooling Strategy for Visual Semantic Embedding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning the Best Pooling Strategy for Visual Semantic Embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.310826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.864794Z digest=sha256:b8ffc2be937a54188089ae26a35e81f3faac267e8d659b9b9c48367fdba39764

Observation 5c620472-2588-440c-b775-3ea6db513a70 · outbound

This paper cites DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.294215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.869660Z digest=sha256:48a7e96b1007c1bdf3b1231a4f334c961daa65c5f63d8782edc7114eaf076dcb

Observation 60c8fc4d-61b7-47e8-8edf-ed2d186a6c97 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning Transferable Visual Models From Natural Language Supervision,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.278911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.873998Z digest=sha256:db52f767ba899d097fa2c95026f26911f8af49910e3ebbd74a04db67a69dc453

Observation 46fe8fa6-37fc-4019-8dcd-858642da83cd · outbound

This paper cites ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.263624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.878311Z digest=sha256:a51421a7f00d3664a204c82e632f615defd2758e0600b01f559575f0c89ae960

Observation bc757e8e-5304-499b-b4f7-8b5485bbea67 · outbound

This paper cites SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.247617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.883444Z digest=sha256:57cebb7a360b3953c563b1f3e564ace769aee7e548f8fdbab6d7478addb58b20

Observation aba342a8-9f5d-4cfe-88f6-5fa7c2d1eac4 · outbound

This paper cites Hierarchical Con- ditional Relation Networks for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Hierarchical Con- ditional Relation Networks for Video Question Answering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.232752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.887935Z digest=sha256:cf1cd5afb255bcdec2dcc2dfa8e7df20a60dfc9e3740dfb0d590f569d4a889e7

Observation d05302cd-1063-4a5e-9a41-d0d265532e07 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering: Datasets, Algorithms and Challenges,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.217463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.892548Z digest=sha256:60ec832f61d092c4b5aa008d41103981c81cd755d104673f72d9cf9afee44510

Observation cecd60ef-9e7a-4f28-9cc1-9663220d202e · outbound

This paper cites Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.200092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.897229Z digest=sha256:93ba7575e9e9fd0d7e9f5e413b168eb935695ffd87aae13e84197471f7d46ad8

Observation 3aee8a22-e517-41c6-8159-c847b1d072cd · outbound

This paper cites Video Question Answering with Iterative Video-Text Co- Tokenization,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering with Iterative Video-Text Co- Tokenization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.183820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.901654Z digest=sha256:8d0e2ea6f0bfb12ca69483fc84484bbba03b8064fc6c74d7a867e53fb4f2094e

Observation 3d53e862-8086-4691-8574-2ff90815e13e · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.167383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.906018Z digest=sha256:c6924b41789d96e278100982e7f3786d1ca6489546632c6ee77600808e08f53e

Observation 6f94017d-14a8-43b1-9f46-e0e0993c33f0 · outbound

This paper cites Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.151121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.910639Z digest=sha256:f0693dcfac1f1f85bff490444330beb93355714b0d33a684035b158034e27f25

Observation 06dea930-9905-4701-b330-7647b8ec12b1 · outbound

This paper cites Jointly Localizing and Describing Events for Dense Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Localizing and Describing Events for Dense Video Captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.136079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.915159Z digest=sha256:fa26e9bf861eb79387b9ac449ffaa04b89acbd40c42205496014df10747940df

Observation 72c9e196-afda-4e0b-b89a-ba78e686ab17 · outbound

This paper cites Video Captioning with Transferred Semantic Attributes,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Captioning with Transferred Semantic Attributes,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.121726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.919598Z digest=sha256:d79cb27606a6c2b8099d94d1ae13c33c3037aee7123ea5b5ff776e0cc4973ab0

Observation 30e5a4da-b26c-40e9-8b55-4ed14d965516 · outbound

This paper cites Jointly Modeling Embedding and Translation to Bridge Video and Language,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Jointly Modeling Embedding and Translation to Bridge Video and Language,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.106098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.924199Z digest=sha256:7c2c1818ea6fb1f30bc675fa92e51608fa56dcca1e2ae6c218d329393213f049

Observation b1c5967c-c4d2-4621-b14f-94ede0299011 · outbound

This paper cites Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.090476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.928690Z digest=sha256:9c02fc1ed5444973c19080db9645270f9ec1b76a515b97007db002e3a2030042

Observation 7ade93c9-c123-4725-b194-f045060e46e9 · outbound

This paper cites an unresolved cited work.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:12:07.074915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.933482Z digest=sha256:c702ce82ed88c33f783a335f26e75745d433cc5eeecc7fa2550bdf2b47224e29

Observation 07426d4a-e689-4008-9b1a-5ba4e3deb0bb · outbound

This paper cites Optical-model po- tential in finite nuclei from Reid’s hard core interaction,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Optical-model po- tential in finite nuclei from Reid’s hard core interaction,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.059636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.937927Z digest=sha256:09df04e844dfc7a734633e44f6668b42077497446cb3f9e7f26f11b2e64977da

Observation 01e9f7a3-a58c-4079-a4e6-1bb56496094c · outbound

This paper cites Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.044075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.942379Z digest=sha256:224f7450147e7c7c36ec5cc524a6c5a06eb27f45ed289c35a9dc958b0258bf3d

Observation a32ecb05-ee33-4dca-90a3-2989dc7ab9d9 · outbound

This paper cites VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.026810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.947661Z digest=sha256:ff69ff263241fac273124750c44da452b39f26795d0d6a520280ffdf82afd2ff

Observation d89eefbc-6df6-41e0-9283-65bb44626049 · outbound

This paper cites Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:07.010019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.952180Z digest=sha256:7e9b5cabb6bb519c374567b01f0bab009d2221a1292858c5c625ad8642cee3f7

Observation d12f0efa-572a-49ea-964e-9618a659463d · outbound

This paper cites Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.993715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.956766Z digest=sha256:7fe5ed64e5b9060d2282437ef5c9ffb6bd77fba9aee6d1f3861613100718ed02

Observation b4435afe-cd1d-43ca-aee0-7ed9b362b0b1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.976856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.961335Z digest=sha256:66a92d586abb4495f46047ba65041d81a821af02802b3890fa82377e3abd8c5d

Observation 6b280c4e-95ac-4ae6-acd2-1ccd5c61d464 · outbound

This paper cites Kullback, Information Theory and Statistics.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Kullback, Information Theory and Statistics

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.960734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.966063Z digest=sha256:ea7851742e302c2d505c39b4538925d601c15fafeb8d57c6a3168914505cd735

Observation e03b7460-3f52-4367-9457-1b99a1dcf37c · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Study on density peaks clustering based on k-nearest neighbors and principal component analysis,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.971077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.971077Z digest=sha256:92280b82e44d8090f2a85c0ab40aa9a01e7d28bd3cf52e06e2124fb0dfa451ae

Observation 5ee34466-8277-4045-8893-f4b3911909ad · outbound

This paper cites ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.934023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.975596Z digest=sha256:9d1acbe91cb7199a42a87a477d3a02192b36b765b5ebf0b41a7a01fe2d130783

Observation e3afa91d-1e33-4046-8e5e-f679205bacfd · outbound

This paper cites Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.917159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.980369Z digest=sha256:394128bea062584d920aea6d014fd150111c6c470b2f6238a70275efbce82cdf

Observation ba9f6136-6bcf-4314-ba69-d67834dda4f0 · outbound

This paper cites Cross Modal Retrieval with Querybank Normalisation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Cross Modal Retrieval with Querybank Normalisation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.901496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.984936Z digest=sha256:178cfa5b52b4b9273b9fbc5f3ff71a350f37e815200a28a7a76d9127fbf7e849

Observation 997268e8-2bf2-473c-b96c-eae4f7ab6974 · outbound

This paper cites Multi-modal Transformer for Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-modal Transformer for Video Retrieval,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.885504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.990474Z digest=sha256:ef2e728f95ff21cf4881df714e005c1e8cfc33d742ed693587f18f5e0845955e

Observation 8cce65b0-6d35-4cfb-98b6-760e8e60abd7 · outbound

This paper cites T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.870428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.994958Z digest=sha256:f6fadd6055f20a7e44e27671efc7c9317a3b50f22925f2a12226f3f03a2d45d2

Observation 70980822-f817-4cbb-88e1-50106cc2be1d · outbound

This paper cites TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.854606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:05.999483Z digest=sha256:6455c8bd2f88d6c8a62381c495a7a2181e907fa703ef4df7e226ace99c96da25

Observation e55b30e8-aa31-416c-8cce-50fef8ba3127 · outbound

This paper cites Support-set bottlenecks for video- text representation learning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Support-set bottlenecks for video- text representation learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.839480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.004645Z digest=sha256:b8191d1d8f2f36d8d5eb14afb4210271d2140173e8c32e13f53ecf56870b787a

Observation bd42f802-5095-4e2c-80d9-bb264a531cfc · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.824809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.009501Z digest=sha256:399420e9ee5ae018e0f1f2b35eabf9fbe109179d15fff5f5fd3ce7ce6626b3ed

Observation cb85fc88-9c83-4ee4-9f5d-ed059dc9c89e · outbound

This paper cites X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.808890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.013862Z digest=sha256:76ad39a464dbcaccc3fd205b03724716de963f5418dd1c3ec652df173faa0b44

Observation 57368048-0dcc-483a-884b-ccffb54fb35d · outbound

This paper cites TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.793165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.018307Z digest=sha256:3a4db0321d48eac4968b403f3a148ae303abac13a8e06d42beb066fdc29bb247

Observation 133ea5e1-da19-45f1-a8d2-e39daae0b236 · outbound

This paper cites UATVR: Uncertainty-Adaptive Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UATVR: Uncertainty-Adaptive Text-Video Retrieval,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.778433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.022809Z digest=sha256:80ff953cfec073d7d46bf83e351cf396040c7f4bbc39800da66a2222a137734f

Observation 776c9f9a-afb4-46d5-8000-cf06ec459f02 · outbound

This paper cites Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.762238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.027184Z digest=sha256:54661d0ebb01ecaf98adecddda6ec3e3f111ceedc7d0bcf7fed6f9fe53fce9ed

Observation aa1b7248-99e4-48db-94ae-e21fd7fd86ed · outbound

This paper cites CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.746220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.031875Z digest=sha256:2d60b39ef2d7fbef92ae57a4efd7736d24f1ef90acc1bd7641e22449f8c70e52

Observation da4aa689-5234-4760-b1f2-2ba1ce0e7b18 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Representation Learning with Contrastive Predictive Coding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:06.036476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:06.036476Z digest=sha256:917759495cb63ea0bfb884c4e6dd46e7ff7a5ef6a8b3bcd6faad4240384913ed

Observation c0fb8e42-135b-425f-b127-0967ae37d3cb · outbound

This paper cites Use What You Have: Video Retrieval Using Representations From Collaborative Experts,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Use What You Have: Video Retrieval Using Representations From Collaborative Experts,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.731010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.041258Z digest=sha256:d879f1351c571191ae1654aa3c2895867d0828ed9bb41ac12354d7c295004884

Observation 79c510ef-24e5-44a9-a331-f0b7d7605a6c · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.715584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.045870Z digest=sha256:aaf922e4bc5ac8a2c22cb5095cf65769fbfb6c763705ba9384309462755bf3b8

Observation 427b624c-1c82-47b8-8a34-34b1837428ce · outbound

This paper cites A Joint Sequence Fusion Model for Video Question Answering and Retrieval,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning A Joint Sequence Fusion Model for Video Question Answering and Retrieval,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.700240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.050286Z digest=sha256:1fc1d6df8e352e6381eb7febbaad05979436fbec710286c8721c44b9b385e27f

Observation b00737e7-0ff0-4a41-a683-c3b4ddc93247 · outbound

This paper cites BLEU: a Method for Automatic Evaluation of Machine Translation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning BLEU: a Method for Automatic Evaluation of Machine Translation,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.684401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.055538Z digest=sha256:468654d8084ca049a4554ffebe2481ed5a6060d1ed08ef2e3c0e897fd8946abb

Observation f6780819-c2ea-4a2c-ac1b-75b072106f9e · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:06.060027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:06.060027Z digest=sha256:e0ffb71b356e549e4c60ca6eafadfad4d075ba59d24964d300a7b4b88ca59a00

Observation 2608b2a6-8f9d-4e4e-aa70-44dc6abe8f24 · outbound

This paper cites METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.669554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.065222Z digest=sha256:c9f387404afaafed11f81a994e807d1a80ee196c2e22917c289eacd1dbde2dc0

Observation cc57212a-8e67-4c84-9fdd-f2fe437e7b9f · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CIDEr: Consensus-based Image Description Evaluation,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.653153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.069878Z digest=sha256:cc5d45928e98574ccd38c2282235a27d34de589edc8f8234733dd748840ff3e5

Observation c880074f-2595-4c3a-8a4b-a8fa0607d1bc · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Adam: A Method for Stochastic Opti- mization,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.637445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.074370Z digest=sha256:97ca14417f41d5f1ff0e55f5a15cdd8478c276e602839d99171df752f7ab8e26

Observation d62a870f-b1c9-4696-8aba-208a5cf6ce28 · outbound

This paper cites Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.622128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.079203Z digest=sha256:776c541c660fc16bbd287d030858a0effc0e759f60bc0e061c99cd7452e8f337

Observation d72011aa-5ba8-428b-93d9-d17126a0a052 · outbound

This paper cites Approximating power indices: theoretical and empirical analysis,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Approximating power indices: theoretical and empirical analysis,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.607028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.084651Z digest=sha256:f451d300f382651d1d2af8dd8ae18d8db44e63d301815cd3588cb426f8c17b12

Observation 8a5e0918-8873-4beb-9cb1-3a340a0a809d · outbound

This paper cites All in One: Exploring Unified Video-Language Pre-training,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning All in One: Exploring Unified Video-Language Pre-training,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.592171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.089292Z digest=sha256:95a22e90ff8703ca2254b51360462abcbeb5b42b6b634ed77b373bcd5db36091

Observation a4b19c3e-4df0-4fe8-a0b4-93250bc684a1 · outbound

This paper cites Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.577862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.093859Z digest=sha256:5435b9a9f5eed7bb4f00af8519353c7bfc64ed6aa225b899cbbaa95f7432f537

Observation 59788f69-d0ae-421c-b8f0-aad0845698ad · outbound

This paper cites Multi-Granularity Interaction and Integration Network for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Multi-Granularity Interaction and Integration Network for Video Question Answering,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.562922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.098477Z digest=sha256:6333d94364642dcb8f910c6e54ea3cb68b05a6d05bb4dfea66eaa5516b6a4b9d

Observation 9dceb0e7-669a-4a6c-9a09-f5cdbcb84564 · outbound

This paper cites Invariant Grounding for Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Invariant Grounding for Video Question Answering,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.547762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.103019Z digest=sha256:9a47f20a96cd61a298c20f56396540d648abe23d088659b801b20b93b809cbf5

Observation 03e73832-5573-4e38-a027-897810212dc9 · outbound

This paper cites Learning to Answer Visual Questions from Web Videos.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Learning to Answer Visual Questions from Web Videos

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:12:06.218997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.108321Z digest=sha256:7484d42103d519c43d93d1cc4bc0c29d94f9feae6bee63e0f46d3c5e53d84328

Observation 6bbc903e-83ad-4f64-a7e0-0a5b23d8fcc3 · outbound

This paper cites Video Question Answering With Semantic Disentanglement and Reasoning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Video Question Answering With Semantic Disentanglement and Reasoning,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.532572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.113224Z digest=sha256:6bc34f6f8a190bc590444f04d3067e4a8408f8fab273585cda9634f9442f3ec6

Observation 87fc9f50-ec22-48ff-8463-251123d092c2 · outbound

This paper cites SViTT: Temporal Learning of Sparse Video-Text Transformers,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SViTT: Temporal Learning of Sparse Video-Text Transformers,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.516916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.117745Z digest=sha256:94226a944ccf3236a7964e6167e863418de5c19898c3c7c2ce455333c9dbbd1b

Observation b1503c78-c126-43f8-8679-bb33c5a59b99 · outbound

This paper cites TG-VQA: Ternary Game of Video Question Answering,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning TG-VQA: Ternary Game of Video Question Answering,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.501948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.122401Z digest=sha256:b4d228631f6b66ecde75198859484e5656c2a8d421b6fed17a7e587d4db80fd9

Observation eef74904-de05-42e9-a437-f6a71b29b991 · outbound

This paper cites SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.486822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.127080Z digest=sha256:439b306dd22e83975d71b13de6193279f3152cad1bb222468b4cc62afc31de05

Observation 0a4ac961-06e3-4e5a-b2be-fe544deeec82 · outbound

This paper cites End-to-end Generative Pretraining for Multimodal Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning End-to-end Generative Pretraining for Multimodal Video Captioning,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.471147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.131507Z digest=sha256:23e2d62e0f67ca16d5477f2c3e4876d413fcf7d7f1dc52c6fd174f29e5188365

Observation e5f9d9c6-66ce-4ac3-b6ef-64bcf0c1f83a · outbound

This paper cites Motion Guided Region Message Passing for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Motion Guided Region Message Passing for Video Captioning,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.454914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.136069Z digest=sha256:b097a0838d51c3197624a51793395c4ca115fa535ab4ab9136b87c249e4e7bcd

Observation 49ad36ec-6710-4eed-85f8-b9f6df6b9bd7 · outbound

This paper cites Open-book Video Captioning with Retrieve-Copy-Generate Net- work,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Open-book Video Captioning with Retrieve-Copy-Generate Net- work,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.439692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.140645Z digest=sha256:70ffea1179d1e0724f943f0ec15a8885f16055326cd2c08f40774cd59dcb62d6

Observation 7cffd5d3-dc94-43ff-802b-23f791be790c · outbound

This paper cites Attentive Visual Semantic Specialized Network for Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Attentive Visual Semantic Specialized Network for Video Captioning,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.424464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.145153Z digest=sha256:04a13caa598a5607e7fe369629a14aa21fd43a2740e052980edfcfee1fc8bbe0

Observation 6bcd28a6-c000-49e9-98c4-a62a84bc9b9c · outbound

This paper cites Global semantic enhancement network for video captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Global semantic enhancement network for video captioning,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.408778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.149625Z digest=sha256:571eb5debbb26dc8dc1c9ebacff99bfc9cae0c007137fe2200cd25db3394cf54

Observation 5b245349-1528-467a-a873-cb780712c22e · outbound

This paper cites Accurate and Fast Compressed Video Captioning,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Accurate and Fast Compressed Video Captioning,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.391481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.153955Z digest=sha256:5f7cded2abd49318fc06b32e2e77e1b634868cbc27bb9ed028f3cc5739a08463

Observation 5356a222-97a4-4788-8cd2-56edf440d795 · outbound

This paper cites Emotional Video Captioning with Vision-based Emotion Interpretation Network,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Emotional Video Captioning with Vision-based Emotion Interpretation Network,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.375787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.158851Z digest=sha256:79654cf24836d298d07c1d4800e0a06c95f6144ca6d9b8768a84241d38dba4f4

Observation 989e1414-4faa-433e-81f7-9d43588b5c27 · outbound

This paper cites Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.360802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.163381Z digest=sha256:d6e4ba626c6cf6a5549a6048a0ec432dbd13853c5b025a676d061e195058301a

Observation 2b4da795-3272-4e61-94e2-545790861b51 · outbound

This paper cites CLIP4Caption: CLIP for Video Caption,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning CLIP4Caption: CLIP for Video Caption,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.345385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.167875Z digest=sha256:f91a581dbb60712596643785d960f3f4ff1b52ebf956af832d119583434b8538

Observation 44183826-f439-4dc4-baa9-53623641cdf6 · outbound

This paper cites Visualizing data using t-SNE,.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Visualizing data using t-SNE,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:12:06.330034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:12:06.172447Z digest=sha256:a39fcbbe9253f486b51b4ff90b9a34fd04b00b4c18f95099efcc29b20dbd2e09

Pith citing papers

No inbound Pith citation observations are available.