REVIEW 3 major objections 5 minor 95 references
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Fine-grained video-language alignment can be learned without manual annotations by treating frames and words as cooperative game players and using Banzhaf Interaction values as training targets.
desk verdict Solid empirical extension of the authors' CVPR 2023 HBI work, but the game-theoretic framing rests on an unverified assumption about the characteristic function. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Banzhaf Interaction index in Eq. 1: for a coalition $\{i,j\}$, it averages over all subsets $C$ of the remaining players the difference $\phi(C \cup [\{i,j\}]) + \phi(C) - \phi(C \cup \{i\}) - \phi(C \cup \{j\})$, so a high value means the pair cooperates more than their separate contributions would suggest. The paper sets the characteristic function $\phi$ to the cross-modal similarity $S$ of Eq. 12, a weighted average of per-frame maximum frame-word alignment scores, and it uses this index as the training target for a prediction head whose output $R$ is matched to the index by a KL-divergence loss. Around this object, the machinery consists of a representation reconstruction module that blends single-modal and cross-modal features with learnable weights (Eqs. 2-5), and a token-merge module that clusters tokens via density-peak K-nearest-neighbor search and attention, so the same interaction is computed at entity, action, and event levels.
What would settle it
Compute $S$ from Eq. 12 on real batches and check the three inequalities in Section 3.2.1 for clear positive pairs (a frame containing the object named by the word) and clear negative pairs; if a positive pair fails to make the payoff difference negative or a negative pair fails to make it positive, the premise is violated. A cheaper decisive probe is to replace the Banzhaf targets in Eq. 8 with permuted interaction values and compare downstream scores; if performance survives, the specific game-theoretic semantics are not what carries the gain.
Extended reading notes
Core claim
The central claim is that the uncertainty in fine-grained video-text correspondence, which frame pairs with which word, at what granularity, and with what intensity, can be handled by formulating the correspondence as a multivariate cooperative game. Video frames and text words are the players; the cross-modal similarity function is the characteristic function; and the Banzhaf Interaction of a coalition measures the coalition's incremental contribution to the total similarity score beyond what the players contribute separately. HBI V2 trains a prediction head to output this interaction index, using KL divergence to match the exact index computed from the similarity function, and because exact computation is NP-hard it learns a tiny estimator of the index for speed. The paper also claims that a reconstructed representation, formed as a learnable blend of single-modal and cross-modal encodings, reduces bias in the Banzhaf values, and that stacking token-merge modules produces entity-, action-, and event-level interactions. On three retrieval benchmarks, three video-QA benchmarks, and one captioning benchmark, the method reports consistent gains over its predecessor HBI and over task-specific methods.
Load-bearing premise
The load-bearing premise is that the cross-modal similarity $S$ used as the characteristic function actually rewards strongly matched frame-word pairs and punishes irrelevant ones in the sense required by conditions (a)-(c); if $S$ violates those inequalities, the Banzhaf Interaction values used as training targets do not encode the claimed semantic correspondence.
Editorial extensions
If this is right
- Fine-grained frame-word alignment can be supervised without any manual annotation, because the Banzhaf index turns coarse video-text pair labels into dense soft targets.
- The interaction loss is dropped at inference time, so the method's interpretability and alignment gains do not slow down deployment.
- The same encoder with task-specific heads covers text-video retrieval, video question answering, and video captioning, suggesting the interaction objective is a general video-language learning signal rather than a retrieval-specific trick.
- Because the exact index is NP-hard, the paper's use of a learned estimator means the method trades a small approximation error for tractability; the reported inference time on MSRVTT is only about one second more than the baseline on the test set.
- The hierarchical visualization shows that coalitions of words and clips have higher semantic similarity than individual frame-word pairs, which the paper offers as evidence that the action- and event-level interactions capture coarser correspondences.
Reading between the lines
- A direct empirical check of the three characteristic-function assumptions would settle whether the training targets literally encode semantic correspondence: sample frame-word pairs, compute $S$ from Eq. 12, and test the inequalities in Section 3.2.1; this probe is within reach and is not reported in the paper.
- The same Banzhaf supervision scheme should transfer to other paired modalities such as image-text, audio-text, or video-audio, since the mechanism only requires a tokenization and a differentiable cross-modal similarity, so a natural extension is to apply HBI-style reconstruction and hierarchy there.
- The reported gains depend on the choice of pretrained vision-language backbone and its token granularity; ablating the reconstruction weights with frozen versus finetuned encoders would separate the game-theoretic signal from representation quality, which the paper does not do.
- The hierarchical interaction visualization could be validated as an interpretability tool by comparing the model's frame-word interaction weights against human-annotated alignments on a small probe set, turning the qualitative figures into a quantitative benchmark.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes HBI V2, an extension of the authors' previous HBI method, for video-language representation learning. The method models video frames and text words as players in a cooperative game, defines a Banzhaf Interaction index over the cross-modal similarity function, and uses it as an auxiliary training target via a prediction head and KL divergence. It also introduces a representation reconstruction module that combines single-modal and cross-modal features, a hierarchical token merging scheme (entity/action/event levels), deep supervision, and self-distillation, together with task-specific heads for text-video retrieval, video question answering, and video captioning. Experiments on MSRVTT, ActivityNet Captions, DiDeMo, MSRVTT-QA, MSVD-QA, ActivityNet-QA, and MSRVTT captioning report consistent gains over prior methods, and ablations show each component contributes.
Significance. If the claimed effects are genuine, the paper would offer a broadly applicable recipe for adding fine-grained interaction modeling to contrastive video-language learning, with gains across retrieval, QA, and captioning. The code release, the multi-task evaluation, the ablations, and the efficiency measurements are concrete strengths. The main caveat is that the game-theoretic target is not yet shown to have the semantic content attributed to it: the defining conditions on the characteristic function are asserted rather than verified, and the computational approximation used in training is not validated. These are fixable in revision, but without such validation the 'Banzhaf Interaction' serving as fine-grained alignment signal could be just an additional self-supervision head.
major comments (3)
- [Section 3.2.1, Eq. (1) and Eq. (12)] The characteristic function φ is taken to be the similarity S of Eq. (12), but S is defined only for a complete video–text pair; it is never specified for an arbitrary coalition C⊆N or for the merged coalition N\{i,j}∪{[{i,j}]}. Consequently the Banzhaf Interaction I([{v_i,t_j}]) in Eq. (1) is not well-defined with this choice of φ. Likewise, conditions (a)–(c) in Section 3.2.1 are asserted without proof or empirical check; for the max-over-frames form of Eq. (12), merging a strongly matched pair need not increase S (the maximum can be attained elsewhere), so condition (a) can fail. Because Eq. (8) optimizes against these I values, the central claim that the auxiliary loss provides fine-grained semantic alignment is not supported. The authors should either specify a coalition-level φ satisfying (a)–(c), or provide an empirical validation that the I values correlate with human-annotated or otherwise ground-truth frame–word correspondences.
- [Section 4.1, Implementation Details] The pretrained tiny model that approximates Banzhaf Interaction is described in two sentences, with no information about how its training targets are generated (exact computation for small N? sampling?), how many instances, or what approximation error it achieves. Since the Banzhaf loss in Eq. (8) is computed from this approximation, the observed gains in Table 4 could in principle come from the extra prediction head rather than from the intended game-theoretic signal. Please report the approximation accuracy and, if feasible, an ablation using sampling-based estimates instead of the learned approximator.
- [Section 4.2, Tables 1–3] The paper reports a single run for each configuration, and many of the headline improvements are small (e.g., MSRVTT text-to-video R@1 49.4 vs. 48.6 for HBI; MSRVTT-QA accuracy 46.4 vs. 46.2). Without standard deviations across seeds or a paired test, the claim that HBI V2 consistently surpasses existing methods is not statistically grounded. Please report variance across at least 3 seeds, or provide significance tests for the main comparisons.
minor comments (5)
- [Section 3.1, Eq. (2)–(5)] The notation for the reconstructed representations is ambiguous; in particular, V_c^f and T_c^w appear to be sequences of repeated copies of a single cross-modal vector, and the dimensions of the MLP outputs for γ and δ are not stated. Please clarify whether γ and δ are scalar or per-token and specify the output shapes.
- [Section 2.2, reference [45]] The citation for 'core interaction' points to a nuclear physics paper (Jeukenne et al.), which does not appear to be a cooperative-game-theory source; please verify and replace.
- [Section 4.1, Implementation Details] The exact CLIP variant (e.g., ViT-B/32 vs. ViT-L/14, input resolution) is not specified; please state it for reproducibility.
- [Section 3.3, Task-Specific Prediction Heads] The claim that the fine-grained alignment from HBI V2 allows a 'simplified answer prediction head' is not directly supported, since no comparison with a stronger head on the same features is provided.
- [Section 4.4, Fig. 9] The discussion of the γ and δ convergence curves is qualitative; the statement that the model 'adaptively reduces cross-modal information to widen the feature distribution' should be backed by quantitative analysis or a controlled experiment.
Circularity Check
No significant circularity: Banzhaf targets are self-supervised auxiliary labels, not fitted final predictions, and the reported gains are measured on external benchmarks.
full rationale
The paper's training signal is the Banzhaf Interaction I computed from the model's own cross-modal similarity S (Eqs. 1, 6, 12), and the prediction head R is trained to match I via KL divergence (Eq. 8). This is a self-referential auxiliary objective, but it is not a circular derivation: I is a nonlinear second-order statistic of S, not equal to S by construction; R is an auxiliary head that is removed at inference; and the reported retrieval, VideoQA, and captioning results are evaluated on held-out test splits against external baselines, so the main performance claim is not a restatement of the training loss. The characteristic-function conditions (a)-(c) in Sec. 3.2.1 are asserted rather than proved for the specific S in Eq. 12, and the exact Banzhaf calculation is replaced by an unvalidated tiny-model approximation (Sec. 4.1); these are correctness and validity risks, not circularity. Self-citations to the authors' prior HBI [11] and related works are used for context and comparison, not as the justification of the current results. No equation in the paper reduces by construction to its own input, and no fitted parameter is renamed as a prediction. Hence no significant circularity.
Assumptions & free parameters
free parameters (5)
- alpha (Banzhaf interaction loss weight) =
1.0 (retrieval/captioning), 2.0 (VideoQA)
- beta (self-distillation loss weight) =
1.0
- lambda (task loss weight) =
2.5 (VideoQA), 3.3 (captioning)
- cluster counts {N_a_v, N_o_v, N_a_t, N_o_t} =
{6, 2, 16, 4}
- temperature tau =
0.01
assumptions (4)
- standard math Banzhaf Interaction in Eq. 1 is a valid interaction measure from cooperative game theory.
- ad hoc to paper The cross-modal similarity S can serve as characteristic function phi and satisfies conditions (a)-(c) in Section 3.2.1.
- domain assumption DPC-KNN clustering produces semantically meaningful token groups at each hierarchy level.
- ad hoc to paper The pre-trained tiny Banzhaf estimator approximates exact Banzhaf Interaction accurately enough to provide useful supervision.
Cite this review
Pith. "Pith review of Hierarchical Banzhaf Interaction for General Video-Language Representation Learning." pith.science (2026). https://pith.science/paper/CKR35S2Q
@misc{pith2026241220964,
author = {Pith},
title = {Pith review of: Hierarchical Banzhaf Interaction for General Video-Language Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/CKR35S2Q}},
note = {Machine review of arXiv:2412.20964}
}
read the original abstract
Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning representations using global semantic interactions between pre-defined video-text pairs. However, to enhance and refine such coarse-grained global interactions, more detailed interactions are necessary for fine-grained multimodal learning. In this study, we introduce a new approach that models video-text as game players using multivariate cooperative game theory to handle uncertainty during fine-grained semantic interactions with diverse granularity, flexible combination, and vague intensity. Specifically, we design the Hierarchical Banzhaf Interaction to simulate the fine-grained correspondence between video clips and textual words from hierarchical perspectives. Furthermore, to mitigate the bias in calculations within Banzhaf Interaction, we propose reconstructing the representation through a fusion of single-modal and cross-modal components. This reconstructed representation ensures fine granularity comparable to that of the single-modal representation, while also preserving the adaptive encoding characteristics of cross-modal representation. Additionally, we extend our original structure into a flexible encoder-decoder framework, enabling the model to adapt to various downstream tasks. Extensive experiments on commonly used text-video retrieval, video-question answering, and video captioning benchmarks, with superior performance, validate the effectiveness and generalization of our method.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[11]
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,
P . Jin, J. Huang, P . Xiong, S. Tian, C. Liu, X. Ji, L. Yuan, and J. Chen, “Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2472–2482
2023
-
[1]
Parallel Vertex Diffusion for Unified Visual Grounding,
Z. Cheng, K. Li, P . Jin, S. Li, X. Ji, L. Yuan, C. Liu, and J. Chen, “Parallel Vertex Diffusion for Unified Visual Grounding,” in Pro- ceedings of the AAAI Conference on Artificial Intelligence , 2024, pp. 1326–1334
2024
-
[2]
Align and Prompt: Video-and-Language Pre-training with Entity Prompts,
D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and Prompt: Video-and-Language Pre-training with Entity Prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4953–4963
2022
-
[3]
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,
P . Jin, J. Huang, F. Liu, X. Wu, S. Ge, G. Song, D. A. Clifton, and J. Chen, “Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,” in Proceedings of the Annual Conference on Neural Information Processing Systems , 2022, pp. 30 291–30 306
2022
-
[4]
FreestyleRet: Retrieving Images from Style-Diversified Queries,
H. Li, Y. Jia, P . Jin, Z. Cheng, K. Li, J. Sui, C. Liu, and L. Yuan, “FreestyleRet: Retrieving Images from Style-Diversified Queries,” in Proceedings of the European Conference on Computer Vision , 2025, pp. 258–274
2025
-
[5]
Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,
W. Wang, J. Gao, X. Yang, and C. Xu, “Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,” IEEE Transactions on Multimedia, vol. 25, pp. 2661– 2674, 2022
2022
-
[6]
Dual Encoding for Video Retrieval by Text,
J. Dong, X. Li, C. Xu, X. Yang, G. Yang, X. Wang, and M. Wang, “Dual Encoding for Video Retrieval by Text,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 8, pp. 4065– 4080, 2021
2021
-
[7]
Temporal Alignment Networks for Long-term Video,
T. Han, W. Xie, and A. Zisserman, “Temporal Alignment Networks for Long-term Video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2906–2916
2022
Show all 95 references
-
[8]
Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,
P . Jin, H. Li, Z. Cheng, K. Li, X. Ji, C. Liu, L. Yuan, and J. Chen, “Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2470–2481
2023
-
[9]
An axiomatic approach to the concept of interaction among players in cooperative games,
M. Grabisch and M. Roubens, “An axiomatic approach to the concept of interaction among players in cooperative games,” In- ternational Journal of Game Theory, vol. 28, pp. 547–565, 1999
1999
-
[10]
Weighted Banzhaf power and interaction indexes through weighted approximations of games,
J.-L. Marichal and P . Mathonet, “Weighted Banzhaf power and interaction indexes through weighted approximations of games,” European Journal of Operational Research, vol. 211, no. 2, pp. 352–358, 2011
2011
-
[12]
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,
J. Xu, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016, pp. 5288–5296
2016
-
[13]
Dense- Captioning Events in Videos,
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles, “Dense- Captioning Events in Videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 706–715
2017
-
[14]
Localizing Moments in Video with Natural Lan- guage,
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing Moments in Video with Natural Lan- guage,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 5803–5812
2017
-
[15]
Video Question Answering via Gradually Refined Attention over Appearance and Motion,
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video Question Answering via Gradually Refined Attention over Appearance and Motion,” in Proceedings of the ACM International Conference on Multimedia, 2017, pp. 1645–1653
2017
-
[16]
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,” in Proceedings of the AAAI Con- ference on Artificial Intelligence, 2019, pp. 9127–9134
2019
-
[17]
Universal Weight- ing Metric Learning for Cross-Modal Retrieval,
J. Wei, Y. Yang, X. Xu, X. Zhu, and H. T. Shen, “Universal Weight- ing Metric Learning for Cross-Modal Retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 10, pp. 6534–6545, 2021
2021
-
[18]
Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,
H. Li, J. Huang, P . Jin, G. Song, Q. Wu, and J. Chen, “Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,” IEEE Transactions on Image Processing , vol. 32, pp. 3367–3382, 2023
2023
-
[19]
Revisiting the ‘Video’ in Video-Language Understand- ing,
S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, and J. C. Niebles, “Revisiting the ‘Video’ in Video-Language Understand- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2917–2927
2022
-
[20]
Fine-Grained Semantically Aligned Vision-Language Pre-Training,
J. Li, X. He, L. Wei, L. Qian, L. Zhu, L. Xie, Y. Zhuang, Q. Tian, and S. Tang, “Fine-Grained Semantically Aligned Vision-Language Pre-Training,” in Proceedings of the Annual Conference on Neural Information Processing Systems, 2022, pp. 7290–7303
2022
-
[21]
Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,
P . Jin, R. Takanobu, W. Zhang, X. Cao, and L. Yuan, “Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 700–13 710
2024
-
[22]
LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,
B. Zhu, B. Lin, M. Ning, Y. Yan, J. Cui, W. HongFa, Y. Pang, W. Jiang, J. Zhang, Z. Li, C. W. Zhang, Z. Li, W. Liu, and L. Yuan, “LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,” in Proceedings of the International Confer...
2024
-
[23]
Decoupled peak property learning for efficient and interpretable ecd spectra prediction,
H. Li, D. Long, L. Yuan, Y. Wang, Y. Tian, X. Wang, and F. Mo, “Decoupled peak property learning for efficient and interpretable ecd spectra prediction,” in Nature Computational Science, 2024
2024
-
[24]
LLaVA-o1: Let Vision Language Models Reason Step-by-Step,
G. Xu, P . Jin, L. Hao, Y. Song, L. Sun, and L. Yuan, “LLaVA-o1: Let Vision Language Models Reason Step-by-Step,” arXiv preprint arXiv:2411.10440, 2024
2024 arXiv
-
[25]
Evagaussians: Event stream assisted gaussian splatting from blurry images,
W. Yu, C. Feng, J. Tang, X. Jia, L. Yuan, and Y. Tian, “Evagaussians: Event stream assisted gaussian splatting from blurry images,” arXiv preprint arXiv:2405.20224, 2024
2024 arXiv
-
[26]
Cycle3d: High-quality and consistent image-to- 3d generation via generation-reconstruction cycle,
Z. Tang, J. Zhang, X. Cheng, W. Yu, C. Feng, Y. Pang, B. Lin, and L. Yuan, “Cycle3d: High-quality and consistent image-to- 3d generation via generation-reconstruction cycle,” arXiv preprint arXiv:2407.19548, 2024
2024 arXiv
-
[27]
Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,
J. Zhang, Z. Tang, Y. Pang, X. Cheng, P . Jin, Y. Wei, X. Zhou, M. Ning, and L. Yuan, “Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,” in Proceedings of the European Conference on Computer Vision , 2025, pp. 303–320
2025
-
[28]
Next patch prediction for autoregressive visual generation,
Y. Pang, P . Jin, S. Yang, B. Lin, B. Zhu, Z. Tang, L. Chen, F. E. Tay, S.-N. Lim, H. Yang et al., “Next patch prediction for autoregressive visual generation,” arXiv preprint arXiv:2412.15321, 2024
2024 arXiv
-
[29]
Learning the Best Pooling Strategy for Visual Semantic Embedding,
J. Chen, H. Hu, H. Wu, Y. Jiang, and C. Wang, “Learning the Best Pooling Strategy for Visual Semantic Embedding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 789–15 798
2021
-
[30]
DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,
X. Yang, L. Zhu, X. Wang, and Y. Yang, “DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024, pp. 6540–6548
2024
-
[31]
Learning Transferable Visual Models From Natural Language Supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in Proceedings of the International Confer- ence on Machine L...
2021
-
[32]
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,
M. Cheng, Y. Sun, L. Wang, X. Zhu, K. Yao, J. Chen, G. Song, J. Han, J. Liu, E. Ding, and J. Wang, “ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5184–5193
2022
-
[33]
SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,
L. Xu, H. Huang, and J. Liu, “SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9878–9888
2021
-
[34]
Hierarchical Con- ditional Relation Networks for Video Question Answering,
T. M. Le, V . Le, S. Venkatesh, and T. Tran, “Hierarchical Con- ditional Relation Networks for Video Question Answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9972–9981
2020
-
[35]
Video Question Answering: Datasets, Algorithms and Challenges,
Y. Zhong, W. Ji, J. Xiao, Y. Li, W. Deng, and T.-S. Chua, “Video Question Answering: Datasets, Algorithms and Challenges,” in Proceedings of the Conference on Empirical Methods in Natural Lan- guage Processing, 2022, pp. 6439–6455. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MA...
2022
-
[36]
Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7331–7341
2021
-
[37]
Video Question Answering with Iterative Video-Text Co- Tokenization,
A. Piergiovanni, K. Morton, W. Kuo, M. S. Ryoo, and A. An- gelova, “Video Question Answering with Iterative Video-Text Co- Tokenization,” inProceedings of the European Conference on Computer Vision, 2022, pp. 76–94
2022
-
[38]
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1728–1738
2021
-
[39]
Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,
P .-Y. Huang, M. Patrick, J. Hu, G. Neubig, F. Metze, and A. Haupt- mann, “Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,” in Proceedings of the Conference of the North American Chapter of the Association for Computational...
2021
-
[40]
Jointly Localizing and Describing Events for Dense Video Captioning,
Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei, “Jointly Localizing and Describing Events for Dense Video Captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 7492–7500
2018
-
[41]
Video Captioning with Transferred Semantic Attributes,
Y. Pan, T. Yao, H. Li, and T. Mei, “Video Captioning with Transferred Semantic Attributes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 6504–6512
2017
-
[42]
Jointly Modeling Embedding and Translation to Bridge Video and Language,
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly Modeling Embedding and Translation to Bridge Video and Language,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016, pp. 4594–4602
2016
-
[43]
Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,
J. Chen, Y. Pan, Y. Li, T. Yao, H. Chao, and T. Mei, “Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,” ACM Transactions on Multimedia Computing, Commu- nications and Applications, vol. 19, no. 1s, pp. 1–24, 2023
2023
-
[44]
T. S. Ferguson, A Course in Game Theory. World Scientific, 2020
2020
-
[45]
Optical-model po- tential in finite nuclei from Reid’s hard core interaction,
J.-P . Jeukenne, A. Lejeune, and C. Mahaux, “Optical-model po- tential in finite nuclei from Reid’s hard core interaction,” Physical Review C, vol. 16, no. 1, p. 80, 1977
1977
-
[46]
Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,
J. Sun, H. Yu, G. Zhong, J. Dong, S. Zhang, and H. Yu, “Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,” IEEE Transactions on Cybernetics , vol. 52, no. 1, pp. 205–214, 2020
2020
-
[47]
VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,
E. Aflalo, M. Du, S.-Y. Tseng, Y. Liu, C. Wu, N. Duan, and V . Lal, “VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 21 406–21 415
2022
-
[48]
Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,
A. Datta, S. Sen, and Y. Zick, “Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,” in IEEE Symposium on Security and Privacy , 2016, pp. 598–617
2016
-
[49]
Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,
P . Jin, H. Li, Z. Cheng, J. Huang, Z. Wang, L. Yuan, C. Liu, and J. Chen, “Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,” in Proceedings of the International Joint Conference on Artificial Intelligence, 2023, pp. 938–946
2023
-
[50]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in Proceedings of the International Con...
2021
-
[51]
Kullback, Information Theory and Statistics
S. Kullback, Information Theory and Statistics. Courier Corporation, 1997
1997
-
[52]
Study on density peaks clustering based on k-nearest neighbors and principal component analysis,
M. Du, S. Ding, and H. Jia, “Study on density peaks clustering based on k-nearest neighbors and principal component analysis,” Knowledge-Based Systems, vol. 99, pp. 135–145, 2016
2016
-
[53]
ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,
K. Li, Z. Wang, Z. Cheng, R. Yu, Y. Zhao, G. Song, C. Liu, L. Yuan, and J. Chen, “ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 7162–7172
2023
-
[54]
Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,
Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,” in Proceedings of the Annual Conference on Neural Information Processing Systems, 2021, pp. 13 937–13 949
2021
-
[55]
Cross Modal Retrieval with Querybank Normalisation,
S.-V . Bogolin, I. Croitoru, H. Jin, Y. Liu, and S. Albanie, “Cross Modal Retrieval with Querybank Normalisation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5194–5205
2022
-
[56]
Multi-modal Transformer for Video Retrieval,
V . Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal Transformer for Video Retrieval,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 214–229
2020
-
[57]
T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,
X. Wang, L. Zhu, and Y. Yang, “T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5079–5088
2021
-
[58]
TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,
I. Croitoru, S.-V . Bogolin, M. Leordeanu, H. Jin, A. Zisserman, S. Albanie, and Y. Liu, “TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 583–11 593
2021
-
[59]
Support-set bottlenecks for video- text representation learning,
M. Patrick, P .-Y. Huang, Y. Asano, F. Metze, A. G. Hauptmann, J. F. Henriques, and A. Vedaldi, “Support-set bottlenecks for video- text representation learning,” in Proceedings of the International Conference on Learning Representations, 2021
2021
-
[60]
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,” Neurocomputing, vol. 508, pp. 293–304, 2022
2022
-
[61]
X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,
S. K. Gorti, N. Vouitsis, J. Ma, K. Golestan, M. Volkovs, A. Garg, and G. Yu, “X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5006–5015
2022
-
[62]
TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,
Y. Liu, P . Xiong, L. Xu, S. Cao, and Q. Jin, “TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,” inProceedings of the European Conference on Computer Vision, 2022, pp. 319–335
2022
-
[63]
UATVR: Uncertainty-Adaptive Text-Video Retrieval,
B. Fang, W. Wu, C. Liu, Y. Zhou, Y. Song, W. Wang, X. Shu, X. Ji, and J. Wang, “UATVR: Uncertainty-Adaptive Text-Video Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 13 723–13 733
2023
-
[64]
Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,
C. Deng, Q. Chen, P . Qin, D. Chen, and Q. Wu, “Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 648–15 658
2023
-
[65]
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,
S. Zhao, L. Zhu, X. Wang, and Y. Yang, “CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,” in Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 2022, pp. 970–981
2022
-
[66]
Representation Learning with Contrastive Predictive Coding,
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation Learning with Contrastive Predictive Coding,” arXiv preprint arXiv:1807.03748 , 2018
2018 arXiv
-
[67]
Use What You Have: Video Retrieval Using Representations From Collaborative Experts,
Y. Liu, S. Albanie, A. Nagrani, and A. Zisserman, “Use What You Have: Video Retrieval Using Representations From Collaborative Experts,” in Proceedings of the British Machine Vision Conference , 2019
2019
-
[68]
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 2630–2640
2019
-
[69]
A Joint Sequence Fusion Model for Video Question Answering and Retrieval,
Y. Yu, J. Kim, and G. Kim, “A Joint Sequence Fusion Model for Video Question Answering and Retrieval,” in Proceedings of the European Conference on Computer Vision, 2018, pp. 471–487
2018
-
[70]
BLEU: a Method for Automatic Evaluation of Machine Translation,
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proceedings of the Annual Meeting on Association for Computational Linguistics , 2002, pp. 311–318
2002
-
[71]
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation,
H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou, “UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation,” arXiv preprint arXiv:2002.06353, 2020
2002 arXiv
-
[72]
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,
S. Banerjee and A. Lavie, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Eval- uation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
-
[73]
CIDEr: Consensus-based Image Description Evaluation,
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based Image Description Evaluation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2015, pp. 4566–4575
2015
-
[74]
Adam: A Method for Stochastic Opti- mization,
D. P . Kingma and J. Ba, “Adam: A Method for Stochastic Opti- mization,” in Proceedings of the International Conference on Learning Representations, 2015
2015
-
[75]
Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,
Y. Matsui and T. Matsui, “Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,” Theoretical Computer Science, vol. 263, no. 1, pp. 305–310, 2001
2001
-
[76]
Approximating power indices: theoretical and empirical analysis,
Y. Bachrach, E. Markakis, E. Resnick, A. D. Procaccia, J. S. Rosen- schein, and A. Saberi, “Approximating power indices: theoretical and empirical analysis,” Autonomous Agents and Multi-Agent Sys- tems, vol. 20, pp. 105–122, 2010
2010
-
[77]
All in One: Exploring Unified Video-Language Pre-training,
J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shan, X. Qie, and M. Z. Shou, “All in One: Exploring Unified Video-Language Pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6598–6608
2023
-
[78]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,” in Proceedings of the Annual Conference on Neural Infor- mation Processing Systems, 2022, pp. 124–141
2022
-
[79]
Multi-Granularity Interaction and Integration Network for Video Question Answering,
Y. Wang, M. Liu, J. Wu, and L. Nie, “Multi-Granularity Interaction and Integration Network for Video Question Answering,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 12, pp. 7684–7695, 2023
2023
-
[80]
Invariant Grounding for Video Question Answering,
Y. Li, X. Wang, J. Xiao, W. Ji, and T.-S. Chua, “Invariant Grounding for Video Question Answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 2928–2937
2022
-
[81]
Learning to Answer Visual Questions from Web Videos,
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Learning to Answer Visual Questions from Web Videos,” arXiv preprint arXiv:2205.05019, 2022
2022 arXiv
-
[82]
Video Question Answering With Semantic Disentanglement and Reasoning,
J. Liu, G. Wang, J. Xie, F. Zhou, and H. Xu, “Video Question Answering With Semantic Disentanglement and Reasoning,”IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 5, pp. 3663–3673, 2024
2024
-
[83]
SViTT: Temporal Learning of Sparse Video-Text Transformers,
Y. Li, K. Min, S. Tripathi, and N. Vasconcelos, “SViTT: Temporal Learning of Sparse Video-Text Transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 919–18 929
2023
-
[84]
TG-VQA: Ternary Game of Video Question Answering,
H. Li, P . Jin, Z. Cheng, S. Zhang, K. Chen, Z. Wang, C. Liu, and J. Chen, “TG-VQA: Ternary Game of Video Question Answering,” in Proceedings of the International Joint Conference on Artificial Intelli- gence, 2023, pp. 1044–1052
2023
-
[85]
SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,
K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y. Lu, and L. Wang, “SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 949–17 958
2022
-
[86]
End-to-end Generative Pretraining for Multimodal Video Captioning,
P . H. Seo, A. Nagrani, A. Arnab, and C. Schmid, “End-to-end Generative Pretraining for Multimodal Video Captioning,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 959–17 968
2022
-
[87]
Motion Guided Region Message Passing for Video Captioning,
S. Chen and Y.-G. Jiang, “Motion Guided Region Message Passing for Video Captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1543–1552
2021
-
[88]
Open-book Video Captioning with Retrieve-Copy-Generate Net- work,
Z. Zhang, Z. Qi, C. Yuan, Y. Shan, B. Li, Y. Deng, and W. Hu, “Open-book Video Captioning with Retrieve-Copy-Generate Net- work,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9837–9846
2021
-
[89]
Attentive Visual Semantic Specialized Network for Video Captioning,
J. Perez-Martin, B. Bustos, and J. P ´erez, “Attentive Visual Semantic Specialized Network for Video Captioning,” in Proceedings of the International Conference on Pattern Recognition, 2021, pp. 5767–5774
2021
-
[90]
Global semantic enhancement network for video captioning,
X. Luo, X. Luo, D. Wang, J. Liu, B. Wan, and L. Zhao, “Global semantic enhancement network for video captioning,” Pattern Recognition, vol. 145, p. 109906, 2024
2024
-
[91]
Accurate and Fast Compressed Video Captioning,
Y. Shen, X. Gu, K. Xu, H. Fan, L. Wen, and L. Zhang, “Accurate and Fast Compressed Video Captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 558–15 567
2023
-
[92]
Emotional Video Captioning with Vision-based Emotion Interpretation Network,
P . Song, D. Guo, X. Yang, S. Tang, and M. Wang, “Emotional Video Captioning with Vision-based Emotion Interpretation Network,” IEEE Transactions on Image Processing, vol. 33, pp. 1122–1135, 2024
2024
-
[93]
Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,
J. Perez-Martin, B. Bustos, and J. Perez, “Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 3039–3049
2021
-
[94]
CLIP4Caption: CLIP for Video Caption,
M. Tang, Z. Wang, Z. Liu, F. Rao, D. Li, and X. Li, “CLIP4Caption: CLIP for Video Caption,” in Proceedings of the ACM International Conference on Multimedia, 2021, pp. 4858–4862
2021
-
[95]
Visualizing data using t-SNE,
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 11, pp. 2579–2605, 2008. Peng Jin received the BE degree in the School of Electronic Information from Tsinghua Univer- sity, China, in 2021. He is currently work...
2008
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.