REVIEW 2 major objections 3 minor 1 cited by
The in-context inductive biases of vision-language models differ across modalities
T0 review · 2 major / 3 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper shows that the same ambiguous category examples, presented as images or text, shift how vision-language models generalize, with images amplifying a shape-over-color bias and text favoring whichever feature is named first.
desk verdict A useful, careful workshop study whose central modality claim is currently confounded with adjective order, but the fix is internal and cheap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The measuring instrument is a shape-versus-color bias score: the proportion of shape-based generalizations minus the proportion of color-based generalizations, computed across trials. It is applied to three paradigms borrowed from cognitive science: one-category generalization, where several same-color same-shape exemplars share a novel label and probes match on only one feature; two-category cue conflict, where two categories differ in both features and probe objects mix them; and odd-one-out, where a set contains one object unique in shape and one unique in color. The stimuli are photographs of geometric solids, deliberately new so they cannot be memorized from pretraining, and the text conditions describe the same objects with adjective order varied, such as 'color red and shape cube' versus 'shape cube and color red.' This turns an unobservable inductive bias into a measurable choice between features.
What would settle it
Run the same three paradigms with text descriptions that carry no pragmatic cues, such as a preamble stating that features are listed in random order and are equally diagnostic. If the first-mentioned-feature bias persists, word order is not being read pragmatically; if it disappears while the image-versus-text shape-bias gap remains, the modality claim survives but the word-order explanation needs revision.
Extended reading notes
Core claim
At the paper's core is a comparison of three tasks—one-category generalization, two-category cue conflict, and odd-one-out—in which the correct generalization is deliberately ambiguous because both shape and color are consistent with the examples. The paper's central finding is that vision-language models generalize more by shape than by color on average, that this shape bias is usually larger when the examples are presented as images, and that in text the adjective order shifts the bias toward the first-named feature, such as 'red cube' versus 'cube that is red.' The authors interpret this as evidence that the in-context representation of the examples differs across modalities and formats, while noting that the direction and size of the effects are idiosyncratic across models and paradigms, including an inversion on the odd-one-out task for two models.
Load-bearing premise
The image-versus-text comparison assumes the image and the text description are informationally equivalent except for modality; if the wording adds its own meaning, the observed differences may reflect language conventions rather than modality-specific representation.
Editorial extensions
If this is right
- Image presentations generally strengthen the shape-over-color bias relative to text presentations.
- When category cues are given in text, the feature named first is favored, so 'shape cube and color red' shifts generalization toward shape, while 'color red and shape cube' shifts it toward color.
- These bias shifts are not universal; they vary in magnitude and sometimes direction across models and task paradigms, so any practical use must be verified for the specific model and task.
- In-context inductive bias is not a fixed property of a model but a function of how the context is rendered.
- The novel image dataset built for the study avoids the concern that effects come from familiar pretraining images.
Reading between the lines
- A natural next test is whether the first-mentioned-feature bias disappears when the text explicitly marks both features as equally relevant; if it does, the pragmatic-reading explanation is supported, and if it does not, the effect is a purely syntactic order effect.
- A practical consequence the paper leaves implicit is that prompt engineers can use modality and adjective order to steer few-shot generalization without changing the underlying category information.
- Given the paper's observation that many multimodal models are text-first and vision-later, the modality differences may trace to training-data skew rather than to any intrinsic difference between vision and language; comparing models with different training mixtures would separate these accounts.
- The odd-one-out set-size effect, where the smallest set of three objects reverses the modality pattern, suggests that claims about modality biases are scoped to the task's configuration and should be revisited as tasks are scaled up.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies in-context inductive biases of three vision-language models (Gemini 1.5 Pro, Claude 3.5 Sonnet, GPT-4o Mini) using three cognitive-science paradigms: one-category generalization, two-category cue conflict, and odd-one-out. The authors compare performance when category examples are presented as images versus as text, and also vary the order of color and shape adjectives in the text descriptions. The main empirical claims are that (1) models generally show a shape-over-color bias when generalizing from ambiguous in-context examples, (2) this shape bias is amplified for image presentations relative to text in the 'standard' color-first phrasing, and (3) in text, models favor the first-mentioned descriptor. The paper introduces a new dataset of toy geometric objects intended to avoid pretraining contamination, reports results across three models and three paradigms, and includes bootstrapped CIs and regression-based significance tests. The authors are careful to note that the effects vary across models and task configurations, and they offer a pragmatic-language explanation for the text-order effects in the Discussion.
Significance. If the central modality claim holds, the paper provides a clean demonstration that presentation format changes in-context inductive generalization in VLMs, with practical implications for prompt design and for the cognitive-science alignment of these models. The work has several concrete strengths: a purpose-built image dataset that reduces pretraining-contamination concerns; three distinct paradigms adapted from the human category-learning literature; tests across three commercially relevant VLMs; bootstrapped confidence intervals; and an appropriately hedged writing style that flags model- and task-dependence. The 'first-mentioned adjective' bias in text is a crisp, falsifiable result. However, the significance of the modality claim is currently limited by a confound between presentation modality and adjective order in the primary image-vs-text comparison, as detailed below. The paper is a workshop-length preliminary investigation, and its contribution is best judged as a set of empirical observations that require the additional control analyses before the modality interpretation can be accepted.
major comments (2)
- [Section 3, Fig. 2; Section 2 (Methods); Fig. 3] The central image-vs-text comparison in Fig. 2 uses as its text condition the 'standard adjective ordering' descriptions described in Section 2, e.g. 'a red cube,' in which color precedes shape. Fig. 3 independently shows that the models generalize toward the first-mentioned descriptor in text, so the standard text condition is precisely the text condition least favorable to a shape bias. The paper does not report whether the image advantage over text survives when the text condition is the shape-first order ('an object that has shape cube and color red') or when the two text orders are averaged. Because the shape-first text data appear in Fig. 3 and the authors' own methods state that the different orderings were collected 'in order to assess whether adjective order affects the model biases,' this is an internal, testable confound rather than an external concern. If the image advantage disappears against the shape-first text condition, the effect attributed to modality would instead be a word-order (or word-order plus surface-form) effect. Please report the image condition against each text-order condition separately, and include a text-order dummy (or equivalent) in the regression that supports the main modality claim.
- [Section 3, regressions; Appendix B, Fig. 6] The statistical support for the overall modality difference is reported only as t-stats and p-values (Section 3: t(823) = 12.478 for Gemini, etc.), without the regression specification. It is therefore unclear whether the model controls for task type only, whether adjective order or text phrasing is a factor (relevant to the confound above), and whether task-configuration variables such as set size in the odd-one-out paradigm are included. The appendix (Fig. 6) shows that the modality effect inverts for set size 3 in the odd-one-out task, and the appendix text describes the one-category results (Fig. 5) as 'very noisy and inconsistent.' To assess robustness of the pooled claim, please provide the full specification (fixed and random effects), list which trial-level variables were covariates, and report a sensitivity analysis that excludes or interacts set size 3 and that separates the explicit category-learning tasks from the odd-one-out task.
minor comments (3)
- [Section 1; Fig. 2; Fig. 3] There are minor typos: 'idiosyncractic' in Section 1 should be 'idiosyncratic,' and 'boostrapped' in the Fig. 2 and Fig. 3 captions should be 'bootstrapped.'
- [Section 3, order-effect regressions] The t-stats for the order effect are reported as t(1681) for Gemini and t(839) for the other two models, but the differing degrees of freedom are not explained; please clarify whether this reflects exclusions such as the one-category refusals noted in the Fig. 3 caption.
- [Section 2, 'Shape vs. color bias plotting'] Please define the shape-vs-color bias metric explicitly in the main text as the difference between the proportion of shape generalizations and the proportion of color generalizations, and state its range, rather than leaving the reader to infer the definition from the figures.
Circularity Check
No circularity: the modality and adjective-order effects are direct measurements of model outputs with no fitted parameters or self-citation chain.
full rationale
All outcome measures are computed directly from raw model choices across the image and text conditions; no parameter is fitted to the data and then reported as a prediction. The image/text contrast (Fig. 2) and the adjective-order contrast (Fig. 3) are measured separately and neither result is defined in terms of the other. The only self-citations (Chan et al. 2022 for the two-category paradigm; Lampinen et al. 2024 for in-context learning) are methodological or background citations and are not load-bearing for the empirical claims. The possible confound that the standard text phrasing places color before shape (Sec. 2, Sec. 4) is a validity and interpretation concern about what the modality comparison isolates, not a circularity: the reported effects do not reduce to their inputs by construction.
Assumptions & free parameters
assumptions (2)
- domain assumption The three tasks (one-category generalization, two-category cue conflict, odd-one-out) measure a stable 'shape vs color' inductive bias, and aggregating them via a fixed bias metric is valid.
- domain assumption The textual descriptions are semantically equivalent to the images apart from modality (and deliberately varied adjective order).
Cite this review
Pith. "Pith review of The in-context inductive biases of vision-language models differ across modalities." pith.science (2026). https://pith.science/paper/ICHXSD2T
@misc{pith2026250201530,
author = {Pith},
title = {Pith review of: The in-context inductive biases of vision-language models differ across modalities},
year = {2026},
howpublished = {\url{https://pith.science/paper/ICHXSD2T}},
note = {Machine review of arXiv:2502.01530}
}
read the original abstract
Inductive biases are what allow learners to make guesses in the absence of conclusive evidence. These biases have often been studied in cognitive science using concepts or categories -- e.g. by testing how humans generalize a new category from a few examples that leave the category boundary ambiguous. We use these approaches to study generalization in foundation models during in-context learning. Modern foundation models can condition on both vision and text, and differences in how they interpret and learn from these different modalities is an emerging area of study. Here, we study how their generalizations vary by the modality in which stimuli are presented, and the way the stimuli are described in text. We study these biases with three different experimental paradigms, across three different vision-language models. We find that the models generally show some bias towards generalizing according to shape over color. This shape bias tends to be amplified when the examples are presented visually. By contrast, when examples are presented in text, the ordering of adjectives affects generalization. However, the extent of these effects vary across models and paradigms. These results help to reveal how vision-language models represent different types of inputs in context, and may have practical implications for the use of vision-language models.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
Multimodal LLMs reproduce gendered instrument stereotypes across text, image, and audio inputs, with text showing the strongest and audio the weakest alignment.
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35: 0 23716--23736, 2022
2022
-
[2]
Colour-name versus shape-name learning in young children
Marc H Bornstein. Colour-name versus shape-name learning in young children. Journal of Child Language, 12 0 (2): 0 387--393, 1985
work page 1985
-
[3]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...
arXiv 2005
-
[4]
Transformers generalize differently from information stored in context vs in weights
Stephanie CY Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K Lampinen, and Felix Hill. Transformers generalize differently from information stored in context vs in weights. MemARI Workshop, NeurIPS 2022, 2022
work page 2022
-
[5]
Sebastian J Crutch, Sarah Connell, and Elizabeth K Warrington. The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements. Quarterly Journal of Experimental Psychology, 62 0 (7): 0 1377--1390, 2009
work page 2009
-
[6]
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data. Advances in Neural Information Processing Systems, 36, 2024
work page 2024
-
[7]
Ordering adjectives in referential communication
Kumiko Fukumura. Ordering adjectives in referential communication. Journal of Memory and Language, 101: 0 37--50, 2018
work page 2018
-
[8]
Paul Gavrikov, Jovita Lukasik, Steffen Jung, Robert Geirhos, Bianca Lamm, Muhammad Jehanzeb Mirza, Margret Keuper, and Janis Keuper. Are vision language models texture or shape biased and can we steer them? arXiv preprint arXiv:2403.09193, 2024
arXiv 2024
Show all 27 references
-
[9]
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019
2019
-
[10]
Shortcut learning in deep neural networks
Robert Geirhos, J \"o rn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2 0 (11): 0 665--673, 2020
2020
-
[11]
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Partial success in closing the gap between human and machine vision. Advances in Neural Information Processing Systems, 34: 0 23885--23899, 2021
2021
-
[12]
Logic and conversation
HP Grice. Logic and conversation. Syntax and semantics, 3, 1975
1975
-
[13]
Revealing the multidimensional mental representations of natural objects underlying human similarity judgements
Martin N Hebart, Charles Y Zheng, Francisco Pereira, and Chris I Baker. Revealing the multidimensional mental representations of natural objects underlying human similarity judgements. Nature human behaviour, 4 0 (11): 0 1173--1185, 2020
2020
-
[14]
What shapes feature representations? exploring datasets, architectures, and training
Katherine Hermann and Andrew Lampinen. What shapes feature representations? exploring datasets, architectures, and training. Advances in Neural Information Processing Systems, 33: 0 9995--10006, 2020
2020
-
[15]
The broader spectrum of in-context learning
Andrew Kyle Lampinen, Stephanie CY Chan, Aaditya K Singh, and Murray Shanahan. The broader spectrum of in-context learning. arXiv preprint arXiv:2412.03782, 2024
2024 arXiv
-
[16]
The importance of shape in early lexical learning
Barbara Landau, Linda B Smith, and Susan S Jones. The importance of shape in early lexical learning. Cognitive development, 3 0 (3): 0 299--321, 1988
1988
-
[17]
Aligning machine and human visual representations across abstraction levels
Lukas Muttenthaler, Klaus Greff, Frieda Born, Bernhard Spitzer, Simon Kornblith, Michael C Mozer, Klaus-Robert M \"u ller, Thomas Unterthiner, and Andrew K Lampinen. Aligning machine and human visual representations across abstraction levels. arXiv preprint arXiv:2409.06509, 2024
2024 arXiv
-
[18]
Adversarial training for free! Advances in neural information processing systems, 32, 2019
Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! Advances in neural information processing systems, 32, 2019
2019
-
[19]
Intriguing properties of neural networks, 2014
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks, 2014. URL https://arxiv.org/abs/1312.6199
2014 arXiv
-
[20]
What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models
Tessa Verhoef, Kiana Shahrasbi, and Tom Kouwenhoven. What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pp.\ 199--213, 2024
2024
-
[21]
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023
2023 arXiv
-
[22]
Word learning as bayesian inference
Fei Xu and Joshua B Tenenbaum. Word learning as bayesian inference. Psychological review, 114 0 (2): 0 245, 2007
2007
-
[23]
What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023
2023
-
[24]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[25]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[26]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[27]
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.