Pith. sign in

REVIEW 3 major objections 5 minor 95 references

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Fine-grained video-language alignment can be learned without manual annotations by treating frames and words as cooperative game players and using Banzhaf Interaction values as training targets.

desk verdict Solid empirical extension of the authors' CVPR 2023 HBI work, but the game-theoretic framing rests on an unverified assumption about the characteristic function. read the letter →

arxiv 2412.20964 v1 pith:CKR35S2Q submitted 2024-12-30 cs.CV

classification cs.CV
keywords video-languagerepresentationlearningBanzhafinteractioncooperativegametheoryfine-grainedcross-modalalignmenttext-videoretrievalvideoquestionansweringcaptioningcontrastive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Video-language models usually align whole videos to whole sentences through contrastive learning, which cannot say which frame matches which word. This paper argues that fine-grained matches can be learned without manual labels by modelling frames and words as players in a cooperative game, with the cross-modal similarity score as the payoff and the Banzhaf Interaction index as the measure of how much a frame-word coalition contributes to that payoff. The proposed HBI V2 (Hierarchical Banzhaf Interaction V2) uses these interaction values as soft training targets for a small prediction head, repeats this supervision at entity, action, and event levels by merging tokens into coalitions, and blends single-modal and cross-modal features to reduce bias in the interaction computation. If the claim is right, the same objective improves text-video retrieval, video question answering, and video captioning, and the interaction values also serve as a visualization of which words bind to which frames.

What carries the argument

The load-bearing object is the Banzhaf Interaction index in Eq. 1: for a coalition $\{i,j\}$, it averages over all subsets $C$ of the remaining players the difference $\phi(C \cup [\{i,j\}]) + \phi(C) - \phi(C \cup \{i\}) - \phi(C \cup \{j\})$, so a high value means the pair cooperates more than their separate contributions would suggest. The paper sets the characteristic function $\phi$ to the cross-modal similarity $S$ of Eq. 12, a weighted average of per-frame maximum frame-word alignment scores, and it uses this index as the training target for a prediction head whose output $R$ is matched to the index by a KL-divergence loss. Around this object, the machinery consists of a representation reconstruction module that blends single-modal and cross-modal features with learnable weights (Eqs. 2-5), and a token-merge module that clusters tokens via density-peak K-nearest-neighbor search and attention, so the same interaction is computed at entity, action, and event levels.

What would settle it

Compute $S$ from Eq. 12 on real batches and check the three inequalities in Section 3.2.1 for clear positive pairs (a frame containing the object named by the word) and clear negative pairs; if a positive pair fails to make the payoff difference negative or a negative pair fails to make it positive, the premise is violated. A cheaper decisive probe is to replace the Banzhaf targets in Eq. 8 with permuted interaction values and compare downstream scores; if performance survives, the specific game-theoretic semantics are not what carries the gain.

Watch

Extended reading notes

Core claim

The central claim is that the uncertainty in fine-grained video-text correspondence, which frame pairs with which word, at what granularity, and with what intensity, can be handled by formulating the correspondence as a multivariate cooperative game. Video frames and text words are the players; the cross-modal similarity function is the characteristic function; and the Banzhaf Interaction of a coalition measures the coalition's incremental contribution to the total similarity score beyond what the players contribute separately. HBI V2 trains a prediction head to output this interaction index, using KL divergence to match the exact index computed from the similarity function, and because exact computation is NP-hard it learns a tiny estimator of the index for speed. The paper also claims that a reconstructed representation, formed as a learnable blend of single-modal and cross-modal encodings, reduces bias in the Banzhaf values, and that stacking token-merge modules produces entity-, action-, and event-level interactions. On three retrieval benchmarks, three video-QA benchmarks, and one captioning benchmark, the method reports consistent gains over its predecessor HBI and over task-specific methods.

Load-bearing premise

The load-bearing premise is that the cross-modal similarity $S$ used as the characteristic function actually rewards strongly matched frame-word pairs and punishes irrelevant ones in the sense required by conditions (a)-(c); if $S$ violates those inequalities, the Banzhaf Interaction values used as training targets do not encode the claimed semantic correspondence.

Editorial extensions

If this is right

  • Fine-grained frame-word alignment can be supervised without any manual annotation, because the Banzhaf index turns coarse video-text pair labels into dense soft targets.
  • The interaction loss is dropped at inference time, so the method's interpretability and alignment gains do not slow down deployment.
  • The same encoder with task-specific heads covers text-video retrieval, video question answering, and video captioning, suggesting the interaction objective is a general video-language learning signal rather than a retrieval-specific trick.
  • Because the exact index is NP-hard, the paper's use of a learned estimator means the method trades a small approximation error for tractability; the reported inference time on MSRVTT is only about one second more than the baseline on the test set.
  • The hierarchical visualization shows that coalitions of words and clips have higher semantic similarity than individual frame-word pairs, which the paper offers as evidence that the action- and event-level interactions capture coarser correspondences.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct empirical check of the three characteristic-function assumptions would settle whether the training targets literally encode semantic correspondence: sample frame-word pairs, compute $S$ from Eq. 12, and test the inequalities in Section 3.2.1; this probe is within reach and is not reported in the paper.
  • The same Banzhaf supervision scheme should transfer to other paired modalities such as image-text, audio-text, or video-audio, since the mechanism only requires a tokenization and a differentiable cross-modal similarity, so a natural extension is to apply HBI-style reconstruction and hierarchy there.
  • The reported gains depend on the choice of pretrained vision-language backbone and its token granularity; ablating the reconstruction weights with frozen versus finetuned encoders would separate the game-theoretic signal from representation quality, which the paper does not do.
  • The hierarchical interaction visualization could be validated as an interpretability tool by comparing the model's frame-word interaction weights against human-annotated alignments on a small probe set, turning the qualitative figures into a quantitative benchmark.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper proposes HBI V2, an extension of the authors' previous HBI method, for video-language representation learning. The method models video frames and text words as players in a cooperative game, defines a Banzhaf Interaction index over the cross-modal similarity function, and uses it as an auxiliary training target via a prediction head and KL divergence. It also introduces a representation reconstruction module that combines single-modal and cross-modal features, a hierarchical token merging scheme (entity/action/event levels), deep supervision, and self-distillation, together with task-specific heads for text-video retrieval, video question answering, and video captioning. Experiments on MSRVTT, ActivityNet Captions, DiDeMo, MSRVTT-QA, MSVD-QA, ActivityNet-QA, and MSRVTT captioning report consistent gains over prior methods, and ablations show each component contributes.

Significance. If the claimed effects are genuine, the paper would offer a broadly applicable recipe for adding fine-grained interaction modeling to contrastive video-language learning, with gains across retrieval, QA, and captioning. The code release, the multi-task evaluation, the ablations, and the efficiency measurements are concrete strengths. The main caveat is that the game-theoretic target is not yet shown to have the semantic content attributed to it: the defining conditions on the characteristic function are asserted rather than verified, and the computational approximation used in training is not validated. These are fixable in revision, but without such validation the 'Banzhaf Interaction' serving as fine-grained alignment signal could be just an additional self-supervision head.

major comments (3)
  1. [Section 3.2.1, Eq. (1) and Eq. (12)] The characteristic function φ is taken to be the similarity S of Eq. (12), but S is defined only for a complete video–text pair; it is never specified for an arbitrary coalition C⊆N or for the merged coalition N\{i,j}∪{[{i,j}]}. Consequently the Banzhaf Interaction I([{v_i,t_j}]) in Eq. (1) is not well-defined with this choice of φ. Likewise, conditions (a)–(c) in Section 3.2.1 are asserted without proof or empirical check; for the max-over-frames form of Eq. (12), merging a strongly matched pair need not increase S (the maximum can be attained elsewhere), so condition (a) can fail. Because Eq. (8) optimizes against these I values, the central claim that the auxiliary loss provides fine-grained semantic alignment is not supported. The authors should either specify a coalition-level φ satisfying (a)–(c), or provide an empirical validation that the I values correlate with human-annotated or otherwise ground-truth frame–word correspondences.
  2. [Section 4.1, Implementation Details] The pretrained tiny model that approximates Banzhaf Interaction is described in two sentences, with no information about how its training targets are generated (exact computation for small N? sampling?), how many instances, or what approximation error it achieves. Since the Banzhaf loss in Eq. (8) is computed from this approximation, the observed gains in Table 4 could in principle come from the extra prediction head rather than from the intended game-theoretic signal. Please report the approximation accuracy and, if feasible, an ablation using sampling-based estimates instead of the learned approximator.
  3. [Section 4.2, Tables 1–3] The paper reports a single run for each configuration, and many of the headline improvements are small (e.g., MSRVTT text-to-video R@1 49.4 vs. 48.6 for HBI; MSRVTT-QA accuracy 46.4 vs. 46.2). Without standard deviations across seeds or a paired test, the claim that HBI V2 consistently surpasses existing methods is not statistically grounded. Please report variance across at least 3 seeds, or provide significance tests for the main comparisons.
minor comments (5)
  1. [Section 3.1, Eq. (2)–(5)] The notation for the reconstructed representations is ambiguous; in particular, V_c^f and T_c^w appear to be sequences of repeated copies of a single cross-modal vector, and the dimensions of the MLP outputs for γ and δ are not stated. Please clarify whether γ and δ are scalar or per-token and specify the output shapes.
  2. [Section 2.2, reference [45]] The citation for 'core interaction' points to a nuclear physics paper (Jeukenne et al.), which does not appear to be a cooperative-game-theory source; please verify and replace.
  3. [Section 4.1, Implementation Details] The exact CLIP variant (e.g., ViT-B/32 vs. ViT-L/14, input resolution) is not specified; please state it for reproducibility.
  4. [Section 3.3, Task-Specific Prediction Heads] The claim that the fine-grained alignment from HBI V2 allows a 'simplified answer prediction head' is not directly supported, since no comparison with a stronger head on the same features is provided.
  5. [Section 4.4, Fig. 9] The discussion of the γ and δ convergence curves is qualitative; the statement that the model 'adaptively reduces cross-modal information to widen the feature distribution' should be backed by quantitative analysis or a controlled experiment.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Banzhaf targets are self-supervised auxiliary labels, not fitted final predictions, and the reported gains are measured on external benchmarks.

full rationale

The paper's training signal is the Banzhaf Interaction I computed from the model's own cross-modal similarity S (Eqs. 1, 6, 12), and the prediction head R is trained to match I via KL divergence (Eq. 8). This is a self-referential auxiliary objective, but it is not a circular derivation: I is a nonlinear second-order statistic of S, not equal to S by construction; R is an auxiliary head that is removed at inference; and the reported retrieval, VideoQA, and captioning results are evaluated on held-out test splits against external baselines, so the main performance claim is not a restatement of the training loss. The characteristic-function conditions (a)-(c) in Sec. 3.2.1 are asserted rather than proved for the specific S in Eq. 12, and the exact Banzhaf calculation is replaced by an unvalidated tiny-model approximation (Sec. 4.1); these are correctness and validity risks, not circularity. Self-citations to the authors' prior HBI [11] and related works are used for context and comparison, not as the justification of the current results. No equation in the paper reduces by construction to its own input, and no fitted parameter is renamed as a prediction. Hence no significant circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central method rests on the standard Banzhaf Interaction formula plus several task-specific assumptions. The free parameters are hyperparameters and cluster counts selected on validation sets. The most fragile assumption is that S satisfies the cooperative-game conditions, since the Banzhaf labels are computed directly from S. No new physical or conceptual entities are introduced; the hypothetical merged player is standard game theory.

free parameters (5)
  • alpha (Banzhaf interaction loss weight) = 1.0 (retrieval/captioning), 2.0 (VideoQA)
    Set by validation search in Fig. 8(a,c,e); balances contrastive loss L_C and Banzhaf loss L_I in Eq. 14.
  • beta (self-distillation loss weight) = 1.0
    Set by validation search in Fig. 8(b,d,f); balances deep supervision and self-distillation in Eq. 16.
  • lambda (task loss weight) = 2.5 (VideoQA), 3.3 (captioning)
    Set by validation search in Fig. 8(g,h); used in Eq. 16 for VideoQA and captioning.
  • cluster counts {N_a_v, N_o_v, N_a_t, N_o_t} = {6, 2, 16, 4}
    Selected from Table 6; controls the number of action-level and event-level visual and text tokens.
  • temperature tau = 0.01
    Fixed hyperparameter in the contrastive loss Eq. 13; chosen by hand and not ablated.
assumptions (4)
  • standard math Banzhaf Interaction in Eq. 1 is a valid interaction measure from cooperative game theory.
    Taken from Grabisch and Roubens [9] and used without modification.
  • ad hoc to paper The cross-modal similarity S can serve as characteristic function phi and satisfies conditions (a)-(c) in Section 3.2.1.
    Asserted in Section 3.2.1 but never proven; Eq. 12 is not checked against the three inequalities.
  • domain assumption DPC-KNN clustering produces semantically meaningful token groups at each hierarchy level.
    Needed for the hierarchical token merge in Algorithm 1; no cluster-quality evaluation is provided beyond downstream R@1.
  • ad hoc to paper The pre-trained tiny Banzhaf estimator approximates exact Banzhaf Interaction accurately enough to provide useful supervision.
    Mentioned in Section 4.1; training data, sample counts, and approximation error are not reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hierarchical Banzhaf Interaction for General Video-Language Representation Learning." pith.science (2026). https://pith.science/paper/CKR35S2Q

@misc{pith2026241220964,
  author       = {Pith},
  title        = {Pith review of: Hierarchical Banzhaf Interaction for General Video-Language Representation Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CKR35S2Q}},
  note         = {Machine review of arXiv:2412.20964}
}
read the original abstract

Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning representations using global semantic interactions between pre-defined video-text pairs. However, to enhance and refine such coarse-grained global interactions, more detailed interactions are necessary for fine-grained multimodal learning. In this study, we introduce a new approach that models video-text as game players using multivariate cooperative game theory to handle uncertainty during fine-grained semantic interactions with diverse granularity, flexible combination, and vague intensity. Specifically, we design the Hierarchical Banzhaf Interaction to simulate the fine-grained correspondence between video clips and textual words from hierarchical perspectives. Furthermore, to mitigate the bias in calculations within Banzhaf Interaction, we propose reconstructing the representation through a fusion of single-modal and cross-modal components. This reconstructed representation ensures fine granularity comparable to that of the single-modal representation, while also preserving the adaptive encoding characteristics of cross-modal representation. Additionally, we extend our original structure into a flexible encoder-decoder framework, enabling the model to adapt to various downstream tasks. Extensive experiments on commonly used text-video retrieval, video-question answering, and video captioning benchmarks, with superior performance, validate the effectiveness and generalization of our method.

Figures

Figures reproduced from arXiv: 2412.20964 by the authors.

Figure 1
Figure 1. (a) Previous methods only learn a global semantic in [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The intuition of employing Banzhaf Interaction in video-language representation learning. When certain players (frames and words) form a coalition, it entails the exclusion of these players from potential coalitions with others, rendering them mutually exclusive from the target coalition. Banzhaf Interaction quantifies the disparity between the benefits derived from the coalition and the costs incurred due to the lo… view at source ↗
Figure 3
Figure 3. Performance comparisons on text-video retrieval, video-question answering, and video captioning. Our proposed framework, HBI V2, designed for general video-language representation learning, demonstrates superior performance consistently. Notably, HBI V2 not only surpasses the previous HBI, but also outperforms existing task-specific methods. • To the best of our knowledge, we are the first to in￾troduce the multivar… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Overview of our proposed HBI V2 framework. We employ a dual-stream encoder to extract features for video tokens and text tokens. Subsequently, we reconstruct the original representation by merging single-modal and cross￾modal components. We propose a novel proxy traini…
Figure 5
Figure 5. Figure 5: The representation reconstruction module. To address the bias in calculations within Banzhaf Interaction, we reconstruct both video and text representation as a fusion of single-modal and cross-modal components. The representation reconstruction module maintains the gr…
Figure 6
Figure 6. Figure 6: The token merge module. “1D-Conv” denotes the one-dimensional convolutional layer. M input tokens with D channels are first clustered into N clusters. Subsequently, we feed the merged tokens as queries Q and the original tokens as keys K and values V into an attention …
Figure 7
Figure 7. Figure 7: The task-specific prediction heads. To enhance the versatility of HBI V2 across various downstream tasks, we expand our original structure into a flexible encoder-decoder framework comprising an encoder and a task-specific decoder. module as features at a higher semant…
Figure 8
Figure 8. Figure 8: Parameter sensitivity. α, β, and λ are the trade-off hyper-parameters in Eq. 14 and Eq. 15. (a) and (b) are the effects of hyper-parameters α and β on the MSRVTT dataset for the text-video retrieval task. (c) and (d) are the effects of hyper-parameters α and β on the M…
Figure 9
Figure 9. Figure 9: The convergence curves of learnable video weight γ and text weight δ for both the text-to-video retrieval and VideoQA tasks. The video weight γ stably converges across both entity and action levels, while the text weight δ initially fluctuates and converges to differen…
Figure 10
Figure 10. Figure 10: Visualization (t-SNE [95]) of the distribution of original and reconstructed representations on the MSRVTT. We use consistent colors to denote tokens originating from the same text and video. The reconstructed representation maintains the distinctiveness of cross-moda…
Figure 11
Figure 11. Figure 11: Visualization of the hierarchical interaction. We take a representative sample from the MSRVTT dataset as an example. Here, the degree of confidence from high to low is represented by red, orange, green, and blue lines, respectively. video and text modalities, with ea…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 73 canonical work pages

  1. [11]

    Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,

    P . Jin, J. Huang, P . Xiong, S. Tian, C. Liu, X. Ji, L. Yuan, and J. Chen, “Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2472–2482

  2. [1]

    Parallel Vertex Diffusion for Unified Visual Grounding,

    Z. Cheng, K. Li, P . Jin, S. Li, X. Ji, L. Yuan, C. Liu, and J. Chen, “Parallel Vertex Diffusion for Unified Visual Grounding,” in Pro- ceedings of the AAAI Conference on Artificial Intelligence , 2024, pp. 1326–1334

  3. [2]

    Align and Prompt: Video-and-Language Pre-training with Entity Prompts,

    D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and Prompt: Video-and-Language Pre-training with Entity Prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4953–4963

  4. [3]

    Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,

    P . Jin, J. Huang, F. Liu, X. Wu, S. Ge, G. Song, D. A. Clifton, and J. Chen, “Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations,” in Proceedings of the Annual Conference on Neural Information Processing Systems , 2022, pp. 30 291–30 306

  5. [4]

    FreestyleRet: Retrieving Images from Style-Diversified Queries,

    H. Li, Y. Jia, P . Jin, Z. Cheng, K. Li, J. Sui, C. Liu, and L. Yuan, “FreestyleRet: Retrieving Images from Style-Diversified Queries,” in Proceedings of the European Conference on Computer Vision , 2025, pp. 258–274

  6. [5]

    Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,

    W. Wang, J. Gao, X. Yang, and C. Xu, “Many Hands Make Light Work: Transferring Knowledge from Auxiliary Tasks for Video- Text Retrieval,” IEEE Transactions on Multimedia, vol. 25, pp. 2661– 2674, 2022

  7. [6]

    Dual Encoding for Video Retrieval by Text,

    J. Dong, X. Li, C. Xu, X. Yang, G. Yang, X. Wang, and M. Wang, “Dual Encoding for Video Retrieval by Text,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 8, pp. 4065– 4080, 2021

  8. [7]

    Temporal Alignment Networks for Long-term Video,

    T. Han, W. Xie, and A. Zisserman, “Temporal Alignment Networks for Long-term Video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2906–2916

Show all 95 references
  1. [8]

    Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,

    P . Jin, H. Li, Z. Cheng, K. Li, X. Ji, C. Liu, L. Yuan, and J. Chen, “Dif- fusionRet: Generative Text-Video Retrieval with Diffusion Model,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2470–2481

  2. [9]

    An axiomatic approach to the concept of interaction among players in cooperative games,

    M. Grabisch and M. Roubens, “An axiomatic approach to the concept of interaction among players in cooperative games,” In- ternational Journal of Game Theory, vol. 28, pp. 547–565, 1999

  3. [10]

    Weighted Banzhaf power and interaction indexes through weighted approximations of games,

    J.-L. Marichal and P . Mathonet, “Weighted Banzhaf power and interaction indexes through weighted approximations of games,” European Journal of Operational Research, vol. 211, no. 2, pp. 352–358, 2011

  4. [12]

    MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,

    J. Xu, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016, pp. 5288–5296

  5. [13]

    Dense- Captioning Events in Videos,

    R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles, “Dense- Captioning Events in Videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 706–715

  6. [14]

    Localizing Moments in Video with Natural Lan- guage,

    L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing Moments in Video with Natural Lan- guage,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 5803–5812

  7. [15]

    Video Question Answering via Gradually Refined Attention over Appearance and Motion,

    D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video Question Answering via Gradually Refined Attention over Appearance and Motion,” in Proceedings of the ACM International Conference on Multimedia, 2017, pp. 1645–1653

  8. [16]

    ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,

    Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering,” in Proceedings of the AAAI Con- ference on Artificial Intelligence, 2019, pp. 9127–9134

  9. [17]

    Universal Weight- ing Metric Learning for Cross-Modal Retrieval,

    J. Wei, Y. Yang, X. Xu, X. Zhu, and H. T. Shen, “Universal Weight- ing Metric Learning for Cross-Modal Retrieval,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 10, pp. 6534–6545, 2021

  10. [18]

    Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,

    H. Li, J. Huang, P . Jin, G. Song, Q. Wu, and J. Chen, “Weakly- Supervised 3D Spatial Reasoning for Text-Based Visual Question Answering,” IEEE Transactions on Image Processing , vol. 32, pp. 3367–3382, 2023

  11. [19]

    Revisiting the ‘Video’ in Video-Language Understand- ing,

    S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, and J. C. Niebles, “Revisiting the ‘Video’ in Video-Language Understand- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2917–2927

  12. [20]

    Fine-Grained Semantically Aligned Vision-Language Pre-Training,

    J. Li, X. He, L. Wei, L. Qian, L. Zhu, L. Xie, Y. Zhuang, Q. Tian, and S. Tang, “Fine-Grained Semantically Aligned Vision-Language Pre-Training,” in Proceedings of the Annual Conference on Neural Information Processing Systems, 2022, pp. 7290–7303

  13. [21]

    Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,

    P . Jin, R. Takanobu, W. Zhang, X. Cao, and L. Yuan, “Chat- UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 700–13 710

  14. [22]

    LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,

    B. Zhu, B. Lin, M. Ning, Y. Yan, J. Cui, W. HongFa, Y. Pang, W. Jiang, J. Zhang, Z. Li, C. W. Zhang, Z. Li, W. Liu, and L. Yuan, “LanguageBind: Extending Video-Language Pretraining to N- modality by Language-based Semantic Alignment,” in Proceedings of the International Confer...

  15. [23]

    Decoupled peak property learning for efficient and interpretable ecd spectra prediction,

    H. Li, D. Long, L. Yuan, Y. Wang, Y. Tian, X. Wang, and F. Mo, “Decoupled peak property learning for efficient and interpretable ecd spectra prediction,” in Nature Computational Science, 2024

  16. [24]

    LLaVA-o1: Let Vision Language Models Reason Step-by-Step,

    G. Xu, P . Jin, L. Hao, Y. Song, L. Sun, and L. Yuan, “LLaVA-o1: Let Vision Language Models Reason Step-by-Step,” arXiv preprint arXiv:2411.10440, 2024

  17. [25]

    Evagaussians: Event stream assisted gaussian splatting from blurry images,

    W. Yu, C. Feng, J. Tang, X. Jia, L. Yuan, and Y. Tian, “Evagaussians: Event stream assisted gaussian splatting from blurry images,” arXiv preprint arXiv:2405.20224, 2024

  18. [26]

    Cycle3d: High-quality and consistent image-to- 3d generation via generation-reconstruction cycle,

    Z. Tang, J. Zhang, X. Cheng, W. Yu, C. Feng, Y. Pang, B. Lin, and L. Yuan, “Cycle3d: High-quality and consistent image-to- 3d generation via generation-reconstruction cycle,” arXiv preprint arXiv:2407.19548, 2024

  19. [27]

    Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,

    J. Zhang, Z. Tang, Y. Pang, X. Cheng, P . Jin, Y. Wei, X. Zhou, M. Ning, and L. Yuan, “Repaint123: Fast and high-quality one image to 3d generation with progressive controllable repainting,” in Proceedings of the European Conference on Computer Vision , 2025, pp. 303–320

  20. [28]

    Next patch prediction for autoregressive visual generation,

    Y. Pang, P . Jin, S. Yang, B. Lin, B. Zhu, Z. Tang, L. Chen, F. E. Tay, S.-N. Lim, H. Yang et al., “Next patch prediction for autoregressive visual generation,” arXiv preprint arXiv:2412.15321, 2024

  21. [29]

    Learning the Best Pooling Strategy for Visual Semantic Embedding,

    J. Chen, H. Hu, H. Wu, Y. Jiang, and C. Wang, “Learning the Best Pooling Strategy for Visual Semantic Embedding,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 789–15 798

  22. [30]

    DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,

    X. Yang, L. Zhu, X. Wang, and Y. Yang, “DGL: Dynamic Global- Local Prompt Tuning for Text-Video Retrieval,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024, pp. 6540–6548

  23. [31]

    Learning Transferable Visual Models From Natural Language Supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in Proceedings of the International Confer- ence on Machine L...

  24. [32]

    ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,

    M. Cheng, Y. Sun, L. Wang, X. Zhu, K. Yao, J. Chen, G. Song, J. Han, J. Liu, E. Ding, and J. Wang, “ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5184–5193

  25. [33]

    SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,

    L. Xu, H. Huang, and J. Liu, “SUTD-TrafficQA: A Question An- swering Benchmark and an Efficient Network for Video Reasoning over Traffic Events,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9878–9888

  26. [34]

    Hierarchical Con- ditional Relation Networks for Video Question Answering,

    T. M. Le, V . Le, S. Venkatesh, and T. Tran, “Hierarchical Con- ditional Relation Networks for Video Question Answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9972–9981

  27. [35]

    Video Question Answering: Datasets, Algorithms and Challenges,

    Y. Zhong, W. Ji, J. Xiao, Y. Li, W. Deng, and T.-S. Chua, “Video Question Answering: Datasets, Algorithms and Challenges,” in Proceedings of the Conference on Empirical Methods in Natural Lan- guage Processing, 2022, pp. 6439–6455. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MA...

  28. [36]

    Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,

    J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less Is More: ClipBERT for Video-and-Language Learning via Sparse Sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7331–7341

  29. [37]

    Video Question Answering with Iterative Video-Text Co- Tokenization,

    A. Piergiovanni, K. Morton, W. Kuo, M. S. Ryoo, and A. An- gelova, “Video Question Answering with Iterative Video-Text Co- Tokenization,” inProceedings of the European Conference on Computer Vision, 2022, pp. 76–94

  30. [38]

    Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1728–1738

  31. [39]

    Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,

    P .-Y. Huang, M. Patrick, J. Hu, G. Neubig, F. Metze, and A. Haupt- mann, “Multilingual Multimodal Pre-training for Zero-Shot Cross- Lingual Transfer of Vision-Language Models,” in Proceedings of the Conference of the North American Chapter of the Association for Computational...

  32. [40]

    Jointly Localizing and Describing Events for Dense Video Captioning,

    Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei, “Jointly Localizing and Describing Events for Dense Video Captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 7492–7500

  33. [41]

    Video Captioning with Transferred Semantic Attributes,

    Y. Pan, T. Yao, H. Li, and T. Mei, “Video Captioning with Transferred Semantic Attributes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 6504–6512

  34. [42]

    Jointly Modeling Embedding and Translation to Bridge Video and Language,

    Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly Modeling Embedding and Translation to Bridge Video and Language,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016, pp. 4594–4602

  35. [43]

    Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,

    J. Chen, Y. Pan, Y. Li, T. Yao, H. Chao, and T. Mei, “Retrieval Augmented Convolutional Encoder-Decoder Networks for Video Captioning,” ACM Transactions on Multimedia Computing, Commu- nications and Applications, vol. 19, no. 1s, pp. 1–24, 2023

  36. [44]

    T. S. Ferguson, A Course in Game Theory. World Scientific, 2020

  37. [45]

    Optical-model po- tential in finite nuclei from Reid’s hard core interaction,

    J.-P . Jeukenne, A. Lejeune, and C. Mahaux, “Optical-model po- tential in finite nuclei from Reid’s hard core interaction,” Physical Review C, vol. 16, no. 1, p. 80, 1977

  38. [46]

    Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,

    J. Sun, H. Yu, G. Zhong, J. Dong, S. Zhang, and H. Yu, “Random Shapley Forests: Cooperative Game Based Random Forests with Consistency,” IEEE Transactions on Cybernetics , vol. 52, no. 1, pp. 205–214, 2020

  39. [47]

    VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,

    E. Aflalo, M. Du, S.-Y. Tseng, Y. Liu, C. Wu, N. Duan, and V . Lal, “VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 21 406–21 415

  40. [48]

    Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,

    A. Datta, S. Sen, and Y. Zick, “Algorithmic Transparency via Quan- titative Input Influence: Theory and Experiments with Learning Systems,” in IEEE Symposium on Security and Privacy , 2016, pp. 598–617

  41. [49]

    Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,

    P . Jin, H. Li, Z. Cheng, J. Huang, Z. Wang, L. Yuan, C. Liu, and J. Chen, “Text-Video Retrieval with Disentangled Conceptualiza- tion and Set-to-Set Alignment,” in Proceedings of the International Joint Conference on Artificial Intelligence, 2023, pp. 938–946

  42. [50]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in Proceedings of the International Con...

  43. [51]

    Kullback, Information Theory and Statistics

    S. Kullback, Information Theory and Statistics. Courier Corporation, 1997

  44. [52]

    Study on density peaks clustering based on k-nearest neighbors and principal component analysis,

    M. Du, S. Ding, and H. Jia, “Study on density peaks clustering based on k-nearest neighbors and principal component analysis,” Knowledge-Based Systems, vol. 99, pp. 135–145, 2016

  45. [53]

    ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,

    K. Li, Z. Wang, Z. Cheng, R. Yu, Y. Zhao, G. Song, C. Liu, L. Yuan, and J. Chen, “ACSeg: Adaptive Conceptualization for Unsuper- vised Semantic Segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 7162–7172

  46. [54]

    Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,

    Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynam- icViT: Efficient Vision Transformers with Dynamic Token Sparsifi- cation,” in Proceedings of the Annual Conference on Neural Information Processing Systems, 2021, pp. 13 937–13 949

  47. [55]

    Cross Modal Retrieval with Querybank Normalisation,

    S.-V . Bogolin, I. Croitoru, H. Jin, Y. Liu, and S. Albanie, “Cross Modal Retrieval with Querybank Normalisation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5194–5205

  48. [56]

    Multi-modal Transformer for Video Retrieval,

    V . Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal Transformer for Video Retrieval,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 214–229

  49. [57]

    T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,

    X. Wang, L. Zhu, and Y. Yang, “T2VLAD: Global-Local Se- quence Alignment for Text-Video Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5079–5088

  50. [58]

    TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,

    I. Croitoru, S.-V . Bogolin, M. Leordeanu, H. Jin, A. Zisserman, S. Albanie, and Y. Liu, “TEACHTEXT: CrossModal General- ized Distillation for Text-Video Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 583–11 593

  51. [59]

    Support-set bottlenecks for video- text representation learning,

    M. Patrick, P .-Y. Huang, Y. Asano, F. Metze, A. G. Hauptmann, J. F. Henriques, and A. Vedaldi, “Support-set bottlenecks for video- text representation learning,” in Proceedings of the International Conference on Learning Representations, 2021

  52. [60]

    CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,

    H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval,” Neurocomputing, vol. 508, pp. 293–304, 2022

  53. [61]

    X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,

    S. K. Gorti, N. Vouitsis, J. Ma, K. Golestan, M. Volkovs, A. Garg, and G. Yu, “X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5006–5015

  54. [62]

    TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,

    Y. Liu, P . Xiong, L. Xu, S. Cao, and Q. Jin, “TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval,” inProceedings of the European Conference on Computer Vision, 2022, pp. 319–335

  55. [63]

    UATVR: Uncertainty-Adaptive Text-Video Retrieval,

    B. Fang, W. Wu, C. Liu, Y. Zhou, Y. Song, W. Wang, X. Shu, X. Ji, and J. Wang, “UATVR: Uncertainty-Adaptive Text-Video Retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 13 723–13 733

  56. [64]

    Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,

    C. Deng, Q. Chen, P . Qin, D. Chen, and Q. Wu, “Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 648–15 658

  57. [65]

    CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,

    S. Zhao, L. Zhu, X. Wang, and Y. Yang, “CenterCLIP: Token Clustering for Efficient Text-Video Retrieval,” in Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 2022, pp. 970–981

  58. [66]

    Representation Learning with Contrastive Predictive Coding,

    A. v. d. Oord, Y. Li, and O. Vinyals, “Representation Learning with Contrastive Predictive Coding,” arXiv preprint arXiv:1807.03748 , 2018

  59. [67]

    Use What You Have: Video Retrieval Using Representations From Collaborative Experts,

    Y. Liu, S. Albanie, A. Nagrani, and A. Zisserman, “Use What You Have: Video Retrieval Using Representations From Collaborative Experts,” in Proceedings of the British Machine Vision Conference , 2019

  60. [68]

    HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,

    A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 2630–2640

  61. [69]

    A Joint Sequence Fusion Model for Video Question Answering and Retrieval,

    Y. Yu, J. Kim, and G. Kim, “A Joint Sequence Fusion Model for Video Question Answering and Retrieval,” in Proceedings of the European Conference on Computer Vision, 2018, pp. 471–487

  62. [70]

    BLEU: a Method for Automatic Evaluation of Machine Translation,

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proceedings of the Annual Meeting on Association for Computational Linguistics , 2002, pp. 311–318

  63. [71]

    UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation,

    H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou, “UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation,” arXiv preprint arXiv:2002.06353, 2020

  64. [72]

    METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,

    S. Banerjee and A. Lavie, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Eval- uation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72

  65. [73]

    CIDEr: Consensus-based Image Description Evaluation,

    R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based Image Description Evaluation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2015, pp. 4566–4575

  66. [74]

    Adam: A Method for Stochastic Opti- mization,

    D. P . Kingma and J. Ba, “Adam: A Method for Stochastic Opti- mization,” in Proceedings of the International Conference on Learning Representations, 2015

  67. [75]

    Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,

    Y. Matsui and T. Matsui, “Np-completeness for calculating power IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE, VOL. XX,NO. XX, XXX. XXXX 15 indices of weighted majority games,” Theoretical Computer Science, vol. 263, no. 1, pp. 305–310, 2001

  68. [76]

    Approximating power indices: theoretical and empirical analysis,

    Y. Bachrach, E. Markakis, E. Resnick, A. D. Procaccia, J. S. Rosen- schein, and A. Saberi, “Approximating power indices: theoretical and empirical analysis,” Autonomous Agents and Multi-Agent Sys- tems, vol. 20, pp. 105–122, 2010

  69. [77]

    All in One: Exploring Unified Video-Language Pre-training,

    J. Wang, Y. Ge, R. Yan, Y. Ge, K. Q. Lin, S. Tsutsui, X. Lin, G. Cai, J. Wu, Y. Shan, X. Qie, and M. Z. Shou, “All in One: Exploring Unified Video-Language Pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6598–6608

  70. [78]

    Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,

    A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-Shot Video Question Answering via Frozen Bidirectional Language Models,” in Proceedings of the Annual Conference on Neural Infor- mation Processing Systems, 2022, pp. 124–141

  71. [79]

    Multi-Granularity Interaction and Integration Network for Video Question Answering,

    Y. Wang, M. Liu, J. Wu, and L. Nie, “Multi-Granularity Interaction and Integration Network for Video Question Answering,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 12, pp. 7684–7695, 2023

  72. [80]

    Invariant Grounding for Video Question Answering,

    Y. Li, X. Wang, J. Xiao, W. Ji, and T.-S. Chua, “Invariant Grounding for Video Question Answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 2928–2937

  73. [81]

    Learning to Answer Visual Questions from Web Videos,

    A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Learning to Answer Visual Questions from Web Videos,” arXiv preprint arXiv:2205.05019, 2022

  74. [82]

    Video Question Answering With Semantic Disentanglement and Reasoning,

    J. Liu, G. Wang, J. Xie, F. Zhou, and H. Xu, “Video Question Answering With Semantic Disentanglement and Reasoning,”IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 5, pp. 3663–3673, 2024

  75. [83]

    SViTT: Temporal Learning of Sparse Video-Text Transformers,

    Y. Li, K. Min, S. Tripathi, and N. Vasconcelos, “SViTT: Temporal Learning of Sparse Video-Text Transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 919–18 929

  76. [84]

    TG-VQA: Ternary Game of Video Question Answering,

    H. Li, P . Jin, Z. Cheng, S. Zhang, K. Chen, Z. Wang, C. Liu, and J. Chen, “TG-VQA: Ternary Game of Video Question Answering,” in Proceedings of the International Joint Conference on Artificial Intelli- gence, 2023, pp. 1044–1052

  77. [85]

    SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,

    K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y. Lu, and L. Wang, “SWINBERT: End-to-End Transformers with Sparse Attention for Video Captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 949–17 958

  78. [86]

    End-to-end Generative Pretraining for Multimodal Video Captioning,

    P . H. Seo, A. Nagrani, A. Arnab, and C. Schmid, “End-to-end Generative Pretraining for Multimodal Video Captioning,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 959–17 968

  79. [87]

    Motion Guided Region Message Passing for Video Captioning,

    S. Chen and Y.-G. Jiang, “Motion Guided Region Message Passing for Video Captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1543–1552

  80. [88]

    Open-book Video Captioning with Retrieve-Copy-Generate Net- work,

    Z. Zhang, Z. Qi, C. Yuan, Y. Shan, B. Li, Y. Deng, and W. Hu, “Open-book Video Captioning with Retrieve-Copy-Generate Net- work,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9837–9846

  81. [89]

    Attentive Visual Semantic Specialized Network for Video Captioning,

    J. Perez-Martin, B. Bustos, and J. P ´erez, “Attentive Visual Semantic Specialized Network for Video Captioning,” in Proceedings of the International Conference on Pattern Recognition, 2021, pp. 5767–5774

  82. [90]

    Global semantic enhancement network for video captioning,

    X. Luo, X. Luo, D. Wang, J. Liu, B. Wan, and L. Zhao, “Global semantic enhancement network for video captioning,” Pattern Recognition, vol. 145, p. 109906, 2024

  83. [91]

    Accurate and Fast Compressed Video Captioning,

    Y. Shen, X. Gu, K. Xu, H. Fan, L. Wen, and L. Zhang, “Accurate and Fast Compressed Video Captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 558–15 567

  84. [92]

    Emotional Video Captioning with Vision-based Emotion Interpretation Network,

    P . Song, D. Guo, X. Yang, S. Tang, and M. Wang, “Emotional Video Captioning with Vision-based Emotion Interpretation Network,” IEEE Transactions on Image Processing, vol. 33, pp. 1122–1135, 2024

  85. [93]

    Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,

    J. Perez-Martin, B. Bustos, and J. Perez, “Improving Video Cap- tioning with Temporal Composition of a Visual-Syntactic Em- bedding,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 3039–3049

  86. [94]

    CLIP4Caption: CLIP for Video Caption,

    M. Tang, Z. Wang, Z. Liu, F. Rao, D. Li, and X. Li, “CLIP4Caption: CLIP for Video Caption,” in Proceedings of the ACM International Conference on Multimedia, 2021, pp. 4858–4862

  87. [95]

    Visualizing data using t-SNE,

    L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 11, pp. 2579–2605, 2008. Peng Jin received the BE degree in the School of Electronic Information from Tsinghua Univer- sity, China, in 2021. He is currently work...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.