REVIEW 3 major objections 5 minor 60 references
Extract Free Dense Misalignment from CLIP
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Removing the ReLU from CLIP's attention attributions makes negative gradients expose misaligned caption words, giving state-of-the-art zero-shot dense misalignment detection and a CLIPScore replacement that better tracks human alignment…
desk verdict Removing ReLU from CLIP attribution exposes a usable misalignment signal; the paper's global-score claims outrun its calibration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the modified GAE attribution rule $R^h_l = \nabla A^h_l \odot A^h_l$, where $\nabla A^h_l$ is the gradient of the CLIP score with respect to an attention map and $A^h_l$ is the attention map itself; removing the ReLU lets negative values flow. Word attributions $w_j$ are obtained by averaging token-level attributions across layers and along the [EOS] row, and a word is predicted misaligned when $w_j$ falls below a single fixed threshold $\epsilon$. F-CLIPScore then combines the global CLIP score with the summed negative misaligned attributions, acting as a drop-in replacement for CLIPScore.
What would settle it
Take a held-out set of captions where each misaligned word is known, and plot the distribution of the word-level attribution $w_j$ for misaligned versus aligned words per layer group; if the two distributions do not separate below a fixed epsilon consistently across domains, or if the optimal epsilon varies by more than the reported search range when tuned per dataset, the central heuristic is falsified.
Extended reading notes
Core claim
The paper's central claim is that the negative entries of the gradient-attention product $R^h_l = \nabla A^h_l \odot A^h_l$ carry a consistent semantic signal: for a misaligned caption, the text tokens that contradict the image receive proportionally negative attribution, so the decision rule $\mathrm{mis}(w_j)=1$ iff $w_j < \epsilon$ picks out exactly those words. The claim rests on removing the ReLU that prior relevance-propagation methods apply to gradients, which had discarded negative values as noise. Supported by ablations showing full-gradient attributions outperform negative-only variants, the paper asserts that both positive and negative gradients are needed and that averaging attribution maps across the final layers preserves this signal. It further claims that aggregating only the negative misaligned attributions into $(1-\mathrm{score})\cdot\sum_j \mathrm{mis}(w_j)\cdot w_j$ yields a global score, F-CLIPScore, that better correlates with human alignment judgments than the plain CLIP similarity score.
Load-bearing premise
The claim breaks down if negative attribution values do not carry a uniform 'this word is wrong' signal across layers and domains, since the method relies on one fixed threshold epsilon to separate bad words from good words across all benchmarks.
Editorial extensions
If this is right
- Misalignment labels can be extracted from any frozen CLIP model in one backward pass, removing the need for reference captions, object detectors, or fine-tuned reward models.
- F-CLIPScore improves global image-text alignment estimation over CLIPScore, especially on hard negatives where added words inflate plain similarity, giving a cheap upgrade for captioning and text-to-image evaluation.
- The method scales with backbone quality: larger CLIP variants and pretraining on larger alt-text corpora improve both localization accuracy and misalignment classification, pointing to continued gains from better CLIP checkpoints.
- Because the same detector can run at tens of frames per second, it could be used as real-time feedback for generation loops or large-scale data cleaning rather than post-hoc analysis.
Reading between the lines
- Where the paper stops with a fixed epsilon, a natural extension is to calibrate the threshold per domain or per caption-length bucket; the benchmark-wide CLIPScores vary, so a per-domain epsilon might lift performance on very low-score inputs.
- The same 'negative attribution' reading could transfer to other contrastive dual encoders built on attention, not just CLIP, and to other tasks like VQA alignment or retrieval reranking where a token-level mismatch is the failure mode.
- The documented failure cases (backgrounds, small objects, adjectives) suggest the signal is biased toward salient foreground nouns; combining the attribution map with a vision backbone that is stronger on small objects, or with an upweighted token-level prior, is a testable path to fix the bias.
- F-CLIPScore's degraded behavior at very low CLIPScore, where gradients spread across tokens, implies that applying it to noisy alt-text (as opposed to well-aligned generated captions) would need a fallback or a gating function on the global score.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CLIP4DM, a zero-shot method for dense misalignment detection between images and text. The method computes per-token attribution scores by removing the ReLU from the gradient-weighted attention aggregation of GAE (Eq. 8), so that negative attribution values can flow. Word-level attributions are thresholded (Eq. 10) to flag misaligned words, and F-CLIPScore (Eq. 11) aggregates these into a global alignment score. The authors evaluate on FOIL, nocaps-FOIL, HAT, SeeTRUE-Feedback, Rich-HF, MMVP, and SugarCrepe, reporting state-of-the-art zero-shot localization accuracy (e.g., 0.836 LA on FOIL with ViT-H/14) and competitive global metrics, while also providing ablations of the attribution formulation, layer choice, and F-CLIPScore components. The central claim is that negative gradients of individual text tokens in a frozen CLIP model indicate misalignment, and that thresholding these attributions yields a general-purpose detector.
Significance. If the central heuristic holds, the paper contributes a cheap, training-free dense misalignment detector that avoids the heavy inference cost of foundation-model pipelines and the annotation cost of fine-tuned approaches. The experimental breadth is a genuine strength: the method is evaluated on multiple benchmarks spanning natural and generated images/text, single and multiple misalignments, and object, attribute, relation, and action errors. The ablations isolate the contribution of removing ReLU, the choice of layers, and the composition of F-CLIPScore, and the code is publicly released. The qualitative analyses honestly document both strengths (entity-level objects, intangible objects) and limitations (backgrounds, small objects, adjectives). However, the significance is conditional: the universal decision rule in Eq. (10) depends on a threshold that is tuned per benchmark and fails in a regime the paper itself identifies as important.
major comments (3)
- [Token Aggregation and F-CLIPScore (Eq. 10), Table 5] The central decision rule mis(w_j)=1 if w_j < epsilon assumes a single threshold separates misaligned from aligned words across all inputs. Yet the paper uses different thresholds on different benchmarks: epsilon=-0.00005 for FOIL/nocaps-FOIL (Section Experiments) and epsilon=-0.00001 for Rich-HF (Tables 5 and 6). Table 5 shows that changing epsilon from -0.00001 to -0.00005 on Rich-HF swings F1 from 0.427 to 0.314, with recall dropping from 0.516 to 0.231. This demonstrates that the reported state-of-the-art numbers rest on benchmark-specific threshold selection, not on a scale-invariant signal. The paper should either derive a principled way to set epsilon (e.g., input-dependent normalization or a statistical criterion) or explicitly scope the claims to the tuned setting; without this, the 'uniform signal' interpretation is not supported.
- [Appendix D and Figure 13] The paper's own analysis undermines the drop-in replacement claim for F-CLIPScore. In Appendix D, the authors state that when CLIPScore is extremely low, gradients distribute across tokens so that few fall below epsilon, causing F-CLIPScore to assign high alignment to clearly misaligned captions (e.g., 'A car an two men...' is ranked top 3% by F-CLIPScore). Figure 13a shows that F-CLIPScore's Pearson correlation with ground-truth alignment is worst in the [0.0, 0.2) group, which is precisely the regime a misalignment detector must handle. The manuscript suggests applying F-CLIPScore selectively to samples with typical CLIPScore values, which is a severe qualification of the general 'drop-in replacement' claim. The authors should quantify what fraction of real-world inputs fall in the failing regime and either fix the metric or clearly state the restricted applicability.
- [Allowing Negative Gradient Flow (Eq. 8), Ablation Table 8] The paper's core premise—that negative entries of R_l^h carry a uniform semantic signal for misalignment—is an empirical heuristic without a mechanistic derivation. Table 8 shows that using both positive and negative gradients outperforms retaining only negative gradients, which weakens the sign-specific interpretation: if negative gradients alone are the misalignment signal, it is unclear why discarding positive gradients hurts performance. The authors should provide a diagnostic (e.g., distributions of attribution values for aligned vs. misaligned words across score bins, or a per-sample analysis) to show that the negative-gradient signal is not confounded by gradient scale or by interactions with positive gradients. Without such evidence, the claim that negative gradients 'indicate misalignment' is only supported by benchmark aggregates under tuned thresholds.
minor comments (5)
- [Analysis (typo)] The section heading 'Comparsion with Baselines' is misspelled; it should be 'Comparison with Baselines'.
- [Table 1 and Table 11] The column labels 'Dense Misalign' and 'Global Misalign' are confusing: it is not immediately clear that they refer to the type of evaluation (dense detection vs. global score) rather than the number of misaligned words. Please clarify in the caption or use more descriptive names.
- [Table 5] The epsilon values are run into the model names (e.g., 'Ours (ViT-H/14)epsilon=-0.00001'); add a space or present epsilon as a separate column for readability.
- [Appendix D, Figure 15] The example 'A car an two men standing in front of it' contains a grammatical error that should be corrected if it is quoted verbatim; otherwise please mark it as a transcription of the dataset example.
- [Related Work] The related-work section would benefit from a brief discussion of whether the negative-gradient hypothesis has any precedent in gradient-based explanation methods beyond GAE, since the paper frames this as a novel interpretation.
Circularity Check
No significant circularity; the derivation is self-contained and evaluated on external benchmarks.
full rationale
The paper's derivation chain is empirically testable rather than definitionally circular. The method computes CLIP gradient-attribution maps (Eq. 5), removes the ReLU to retain negative gradients (Eq. 8), averages across layers (Eq. 9), thresholds word attributions (Eq. 10), and aggregates them into F-CLIPScore (Eq. 11). None of these equations defines the predicted misalignment label in terms of the benchmark ground-truth labels; the claim that negative attributions indicate misalignment is a heuristic validated against external human- or rule-labeled datasets such as FOIL, nocaps-FOIL, HAT, SeeTRUE-Feedback, and Rich-HF. The only tunable quantity, epsilon, is selected on development splits and reported on held-out test sets, which is standard hyperparameter selection rather than fitting the target labels into the method's definition. There is no load-bearing self-citation: the paper cites GAE as inspiration but explicitly deviates from it, and its ablations compare variants on external metrics. The Appendix D observation that F-CLIPScore degrades at extremely low CLIPScore is a stated limitation and a correctness/robustness concern, not evidence of circularity. The dataset-specific epsilon values (e.g., -0.00001 on Rich-HF vs. the default -0.00005) indicate that the decision threshold is not perfectly calibrated across regimes, but this is a hyperparameter-tuning caveat rather than a reduction of the prediction to its input by construction. Overall, the paper does not rename a known result, import a uniqueness claim from the authors' own prior work, or fit a parameter to the exact quantity it then claims to predict.
Assumptions & free parameters
free parameters (3)
- epsilon (misalignment threshold) =
-0.00005 default; -0.00001 for Rich-HF F1/correlation
- tilde_l (starting layer for attribution accumulation) =
10 (ViT-B/32), 22 (ViT-H/14)
- number of accumulated layers =
3
assumptions (4)
- domain assumption Negative values in grad(score) * attention carry a semantically meaningful misalignment signal
- domain assumption A single global epsilon can separate aligned from misaligned words across domains
- domain assumption Attribution averaging over layers preserves interpretability of signed gradients
- standard math Text tokenization and [EOS] pooling in CLIP provide a sufficient representation for alignment scoring
Cite this review
Pith. "Pith review of Extract Free Dense Misalignment from CLIP." pith.science (2026). https://pith.science/paper/SYB6H6M7
@misc{pith2026241218404,
author = {Pith},
title = {Pith review of: Extract Free Dense Misalignment from CLIP},
year = {2026},
howpublished = {\url{https://pith.science/paper/SYB6H6M7}},
note = {Machine review of arXiv:2412.18404}
}
read the original abstract
Recent vision-language foundation models still frequently produce outputs misaligned with their inputs, evidenced by object hallucination in captioning and prompt misalignment in the text-to-image generation model. Recent studies have explored methods for identifying misaligned elements, aiming not only to enhance interpretability but also to improve model performance. However, current approaches primarily rely on large foundation models in a zero-shot manner or fine-tuned models with human annotations, which limits scalability due to significant computational costs. This work proposes a novel approach, dubbed CLIP4DM, for detecting dense misalignments from pre-trained CLIP, specifically focusing on pinpointing misaligned words between image and text. We carefully revamp the gradient-based attribution computation method, enabling negative gradient of individual text tokens to indicate misalignment. We also propose F-CLIPScore, which aggregates misaligned attributions with a global alignment score. We evaluate our method on various dense misalignment detection benchmarks, covering various image and text domains and misalignment types. Our method demonstrates state-of-the-art performance among zero-shot models and competitive performance with fine-tuned models while maintaining superior efficiency. Our qualitative examples show that our method has a unique strength to detect entity-level objects, intangible objects, and attributes that can not be easily detected for existing works. We conduct ablation studies and analyses to highlight the strengths and limitations of our approach. Our code is publicly available at https://github.com/naver-ai/CLIP4DM.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Abnar, S.; and Zuidema, W. 2020. Quantifying Attention Flow in Transformers. In ACL, 4190--4197. Online: Association for Computational Linguistics
work page 2020
-
[2]
Arras, L.; Montavon, G.; M \"u ller, K.-R.; and Samek, W. 2017. Explaining Recurrent Neural Network Predictions in Sentiment Analysis. In Proceedings of the 8th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, 159--168. Copenhagen, Denmark: ACL
work page 2017
-
[3]
Bach, S.; Binder, A.; Montavon, G.; Klauschen, F.; M \"u ller, K.-R.; and Samek, W. 2015. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7): e0130140
work page 2015
-
[4]
R.; Angeli, G.; Potts, C.; and Manning, C
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015. A large annotated corpus for learning natural language inference. In EMNLP. Association for Computational Linguistics
work page 2015
-
[5]
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-end object detection with transformers. In ECCV, 213--229. Springer
2020
-
[6]
Castro, S.; Ignat, O.; and Mihalcea, R. 2023. Scalable Performance Analysis for Vision-Language Models. In Proceedings of the 12th Joint Conference on Lexical and Computational Semantics (* SEM 2023), 284--294
work page 2023
-
[7]
Chan, D.; Myers, A.; Vijayanarasimhan, S.; Ross, D.; and Canny, J. 2023. IC 3: Image Captioning by Committee Consensus. In EMNLP, 8975--9003. Singapore: Association for Computational Linguistics
work page 2023
-
[8]
Chefer, H.; Alaluf, Y.; Vinker, Y.; Wolf, L.; and Cohen-Or, D. 2023. Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models. ACM Trans. Graph., 42(4)
work page 2023
Show all 60 references
-
[9]
Chefer, H.; Gur, S.; and Wolf, L. 2021 a . Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers. In ICCV, 387--396. IEEE
2021
-
[10]
Chefer, H.; Gur, S.; and Wolf, L. 2021 b . Transformer Interpretability Beyond Attention Visualization. In CVPR, 782--791. Computer Vision Foundation / IEEE
2021
-
[11]
Chen, J.; Zhu, D.; Shen, X.; Li, X.; Liu, Z.; Zhang, P.; Krishnamoorthi, R.; Chandra, V.; Xiong, Y.; and Elhoseiny, M. 2023. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478
2023 arXiv
-
[12]
Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; and Jitsev, J. 2023. Reproducible scaling laws for contrastive language-image learning. In CVPR, 2818--2829
2023
-
[13]
M.; Garg, R.; Anderson, P.; Krishna, R.; Bansal, M.; Pont-Tuset, J.; and Wang, S
Cho, J.; Hu, Y.; Baldridge, J. M.; Garg, R.; Anderson, P.; Krishna, R.; Bansal, M.; Pont-Tuset, J.; and Wang, S. 2024. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation. In ICLR
2024
-
[14]
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In CVPR, 248--255. Ieee
2009
-
[15]
Gordon, B.; Bitton, Y.; Shafir, Y.; Garg, R.; Chen, X.; Lischinski, D.; Cohen-Or, D.; and Szpektor, I. 2024. Mismatch quest: Visual and textual feedback for image-text misalignment. In ECCV, 310--328. Springer
2024
-
[16]
Goyal, Y.; Mohapatra, A.; Parikh, D.; and Batra, D. 2016. Towards transparent ai systems: Interpreting visual question answering models. In ICML 2016 Workshop on Visualization for Deep Learning
2016
-
[17]
Gunjal, A.; Yin, J.; and Bas, E. 2024. Detecting and preventing hallucinations in large vision language models. In AAAI, volume 38, 18135--18143
2024
-
[18]
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021. CLIPS core: A Reference-free Evaluation Metric for Image Captioning. In EMNLP, 7514--7528. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics
2021
-
[19]
Hsieh, C.-Y.; Zhang, J.; Ma, Z.; Kembhavi, A.; and Krishna, R. 2023. SUGARCREPE: fixing hackable benchmarks for vision-language compositionality. In NeurIPS, 31096--31116
2023
-
[20]
Hu, Y.; Liu, B.; Kasai, J.; Wang, Y.; Ostendorf, M.; Krishna, R.; and Smith, N. A. 2023. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In ICCV, 20406--20417
2023
-
[21]
Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2023. Pick-a-Pic: an open dataset of user preferences for text-to-image generation. In NeurIPS, 36652--36663
2023
-
[22]
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020. BART : Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL, 7871--7880. Online: Association for Comp...
2020
-
[23]
Li, J.; Li, D.; Xiong, C.; and Hoi, S. C. H. 2022. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In ICML, volume 162 of Proceedings of Machine Learning Research, 12888--12900. PMLR
2022
-
[24]
X.; and Wen, J.-R
Li, Y.; Du, Y.; Zhou, K.; Wang, J.; Zhao, W. X.; and Wen, J.-R. 2023. Evaluating Object Hallucination in Large Vision-Language Models. In EMNLP, 292--305
2023
-
[25]
Liang, Y.; He, J.; Li, G.; Li, P.; Klimovskiy, A.; Carolan, N.; Sun, J.; Pont-Tuset, J.; Young, S.; Yang, F.; et al. 2024. Rich human feedback for text-to-image generation. In CVPR, 19401--19411
2024
-
[26]
J.; Wang, B.; Li, W.; and Shou, M
Lin, Y.; He, C.; Wang, A. J.; Wang, B.; Li, W.; and Shou, M. Z. 2024. Parrot captions teach clip to spot text. In ECCV, 368--385. Springer
2024
-
[27]
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H.; Zhu, J.; and Zhang, L. 2024. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection. In ECCV, 38--55. Cham: Springer Nature Switzerland. ISBN 978-3-031-72970-6
2024
-
[28]
M.; and Lee, S
Lundberg, S. M.; and Lee, S. 2017. A Unified Approach to Interpreting Model Predictions. In NeurIPS, 4765--4774
2017
-
[29]
O.; Gandhi, M.; Gao, I.; and Krishna, R
Ma, Z.; Hong, J.; Gul, M. O.; Gandhi, M.; Gao, I.; and Krishna, R. 2023. CREPE: Can Vision-Language Foundation Models Reason Compositionally? In CVPR, 10910--10921
2023
-
[30]
Montavon, G.; Lapuschkin, S.; Binder, A.; Samek, W.; and M \"u ller, K.-R. 2017. Explaining nonlinear classification decisions with deep taylor decomposition. Pattern recognition, 65: 211--222
2017
-
[31]
H.; and Lim, S.-N
Mukhoti, J.; Lin, T.-Y.; Poursaeed, O.; Wang, R.; Shah, A.; Torr, P. H.; and Lim, S.-N. 2023. Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning. In CVPR, 19413--19423. IEEE Computer Society
2023
-
[32]
Nikolaus, M.; Salin, E.; Ayache, S.; Fourtassi, A.; and Favre, B. 2022. Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies? In EMNLP, 1538--1555
2022
-
[33]
Paiss, R.; Ephrat, A.; Tov, O.; Zada, S.; Mosseri, I.; Irani, M.; and Dekel, T. 2023. Teaching CLIP to Count to Ten. In ICCV, 3147--3157. IEEE
2023
-
[34]
Petryk, S.; Chan, D.; Kachinthaya, A.; Zou, H.; Canny, J.; Gonzalez, J.; and Darrell, T. 2024. ALOHa: A New Measure for Hallucination in Captioning Models. In NAACL, 342--357
2024
-
[35]
Pezzelle, S. 2023. Dealing with Semantic Underspecification in Multimodal NLP. In ACL, 12098--12112
2023
-
[36]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML, volume 139, 8748--8763
2021
-
[37]
Rassin, R.; Ravfogel, S.; and Goldberg, Y. 2022. DALLE-2 is Seeing Double: Flaws in Word-to-Concept Mapping in Text2Image Models. In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 335--345
2022
-
[38]
Reimers, N.; and Gurevych, I. 2019. Sentence- BERT : Sentence Embeddings using S iamese BERT -Networks. In EMNLP, 3982--3992. Hong Kong, China: Association for Computational Linguistics
2019
-
[39]
T.; Singh, S.; and Guestrin, C
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016. ``Why Should I Trust You?": Explaining the Predictions of Any Classifier. In SIGKDD, 1135--1144. ACM
2016
-
[40]
A.; Burns, K.; Darrell, T.; and Saenko, K
Rohrbach, A.; Hendricks, L. A.; Burns, K.; Darrell, T.; and Saenko, K. 2018. Object Hallucination in Image Captioning. In EMNLP, 4035--4045. Brussels, Belgium: Association for Computational Linguistics
2018
-
[41]
W.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C. W.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022. LAION-5B: An open large-scale dataset for training next generation image-text models. In NeurIPS
2022
-
[42]
R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In ICCV, 618--626. IEEE Computer Society
2017
-
[43]
Shekhar, R.; Pezzelle, S.; Klimovich, Y.; Herbelot, A.; Nabi, M.; Sangineto, E.; and Bernardi, R. 2017. FOIL it! Find One mismatch between Image and Language caption. In ACL, 255--265. Vancouver, Canada: Association for Computational Linguistics
2017
-
[44]
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. In CVPR, 9568--9578
2024
-
[45]
N.; Kaiser, L.; and Polosukhin, I
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In NeurIPS, 5998--6008
2017
-
[46]
Wang, P.; Yang, A.; Men, R.; Lin, J.; Bai, S.; Li, Z.; Ma, J.; Zhou, C.; Zhou, J.; and Yang, H. 2022. OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. In ICML, volume 162 of Proceedings of Machine Learning Research, 2...
2022
-
[47]
G.; and Wilson, A
Wang, Y.; Rudner, T. G.; and Wilson, A. G. 2023. Visual explanations of image-text representations via multi-modal information bottleneck attribution. In NeurIPS, volume 36, 16009--16027
2023
-
[48]
Xiao, W.; Huang, Z.; Gan, L.; He, W.; Li, H.; Yu, Z.; Jiang, H.; Wu, F.; and Zhu, L. 2024. Detecting and mitigating hallucination in large vision language models via fine-grained ai feedback. arXiv preprint, arXiv:2404.14233
2024 arXiv
-
[49]
Yan, S.; Bai, M.; Chen, W.; Zhou, X.; Huang, Q.; and Li, L. E. 2024. ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling. In ECCV, 37--53. Cham: Springer Nature Switzerland. ISBN 978-3-031-73030-6
2024
-
[50]
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2022. FILIP: Fine-grained Interactive Language-Image Pre-Training. In ICLR
2022
-
[51]
Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B
Yu, J.; Xu, Y.; Koh, J. Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B. K.; Hutchinson, B.; Han, W.; Parekh, Z.; Li, X.; Zhang, H.; Baldridge, J.; and Wu, Y. 2022. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. TMLR
2022
-
[52]
Yu, T.; Yao, Y.; Zhang, H.; He, T.; Han, Y.; Cui, G.; Hu, J.; Liu, Z.; Zheng, H.-T.; Sun, M.; et al. 2024. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In CVPR, 13807--13816
2024
-
[53]
Yuksekgonul, M.; Bianchi, F.; Kalluri, P.; Jurafsky, D.; and Zou, J. 2023. When and why vision-language models behave like bags-of-words, and what to do about it? In ICLR
2023
-
[54]
D.; and Fergus, R
Zeiler, M. D.; and Fergus, R. 2014. Visualizing and understanding convolutional networks. In ECCV, 818--833. Springer
2014
-
[55]
Zhang, B.; Zhang, P.; Dong, X.; Zang, Y.; and Wang, J. 2025. Long-clip: Unlocking the long-text capability of clip. In European Conference on Computer Vision, 310--325. Springer
2025
-
[56]
Zhao, C.; Wang, K.; Zeng, X.; Zhao, R.; and Chan, A. B. 2024. Gradient-based visual explanation for transformer-based clip. In ICML, 61072--61091. PMLR
2024
-
[57]
C.; and Dai, B
Zhou, C.; Loy, C. C.; and Dai, B. 2022. Extract free dense labels from clip. In ECCV, 696--712. Springer
2022
-
[58]
Zhu, D.; Chen, J.; Haydarov, K.; Shen, X.; Zhang, W.; and Elhoseiny, M. 2024. Chat GPT Asks, BLIP -2 Answers: Automatic Questioning Towards Enriched Visual Descriptions. TMLR
2024
-
[59]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[60]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.