Pith. sign in

REVIEW 2 major objections 3 minor 1 cited by

The in-context inductive biases of vision-language models differ across modalities

T0 review · 2 major / 3 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper shows that the same ambiguous category examples, presented as images or text, shift how vision-language models generalize, with images amplifying a shape-over-color bias and text favoring whichever feature is named first.

desk verdict A useful, careful workshop study whose central modality claim is currently confounded with adjective order, but the fix is internal and cheap. read the letter →

arxiv 2502.01530 v2 pith:ICHXSD2T submitted 2025-02-03 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords in-contextlearninginductivebiasvision-languagemodelsshape-vs-colorcategorygeneralizationmodalitydifferencesadjectiveorderingfew-shot
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Learners must guess when evidence is ambiguous, and this paper asks whether a vision-language model's guess about a new category depends on how the examples are given in the prompt. Across three cognitive-science paradigms, the authors find that when the same category examples are given as images, the models tend to generalize by shape over color, and more strongly so than when the examples are described in text. When the examples are in text, the order in which features are named changes the bias, with models favoring the feature mentioned first. The results are overall tendencies that vary across models and tasks, but they indicate that the format of in-context examples measurably changes how these models represent and generalize ambiguous evidence. This matters practically because prompting choices, such as using an image versus a sentence or putting one adjective before another, can steer a model's category guess.

What carries the argument

The measuring instrument is a shape-versus-color bias score: the proportion of shape-based generalizations minus the proportion of color-based generalizations, computed across trials. It is applied to three paradigms borrowed from cognitive science: one-category generalization, where several same-color same-shape exemplars share a novel label and probes match on only one feature; two-category cue conflict, where two categories differ in both features and probe objects mix them; and odd-one-out, where a set contains one object unique in shape and one unique in color. The stimuli are photographs of geometric solids, deliberately new so they cannot be memorized from pretraining, and the text conditions describe the same objects with adjective order varied, such as 'color red and shape cube' versus 'shape cube and color red.' This turns an unobservable inductive bias into a measurable choice between features.

What would settle it

Run the same three paradigms with text descriptions that carry no pragmatic cues, such as a preamble stating that features are listed in random order and are equally diagnostic. If the first-mentioned-feature bias persists, word order is not being read pragmatically; if it disappears while the image-versus-text shape-bias gap remains, the modality claim survives but the word-order explanation needs revision.

Watch

Extended reading notes

Core claim

At the paper's core is a comparison of three tasks—one-category generalization, two-category cue conflict, and odd-one-out—in which the correct generalization is deliberately ambiguous because both shape and color are consistent with the examples. The paper's central finding is that vision-language models generalize more by shape than by color on average, that this shape bias is usually larger when the examples are presented as images, and that in text the adjective order shifts the bias toward the first-named feature, such as 'red cube' versus 'cube that is red.' The authors interpret this as evidence that the in-context representation of the examples differs across modalities and formats, while noting that the direction and size of the effects are idiosyncratic across models and paradigms, including an inversion on the odd-one-out task for two models.

Load-bearing premise

The image-versus-text comparison assumes the image and the text description are informationally equivalent except for modality; if the wording adds its own meaning, the observed differences may reflect language conventions rather than modality-specific representation.

Editorial extensions

If this is right

  • Image presentations generally strengthen the shape-over-color bias relative to text presentations.
  • When category cues are given in text, the feature named first is favored, so 'shape cube and color red' shifts generalization toward shape, while 'color red and shape cube' shifts it toward color.
  • These bias shifts are not universal; they vary in magnitude and sometimes direction across models and task paradigms, so any practical use must be verified for the specific model and task.
  • In-context inductive bias is not a fixed property of a model but a function of how the context is rendered.
  • The novel image dataset built for the study avoids the concern that effects come from familiar pretraining images.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test is whether the first-mentioned-feature bias disappears when the text explicitly marks both features as equally relevant; if it does, the pragmatic-reading explanation is supported, and if it does not, the effect is a purely syntactic order effect.
  • A practical consequence the paper leaves implicit is that prompt engineers can use modality and adjective order to steer few-shot generalization without changing the underlying category information.
  • Given the paper's observation that many multimodal models are text-first and vision-later, the modality differences may trace to training-data skew rather than to any intrinsic difference between vision and language; comparing models with different training mixtures would separate these accounts.
  • The odd-one-out set-size effect, where the smallest set of three objects reverses the modality pattern, suggests that claims about modality biases are scoped to the task's configuration and should be revisited as tasks are scaled up.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper studies in-context inductive biases of three vision-language models (Gemini 1.5 Pro, Claude 3.5 Sonnet, GPT-4o Mini) using three cognitive-science paradigms: one-category generalization, two-category cue conflict, and odd-one-out. The authors compare performance when category examples are presented as images versus as text, and also vary the order of color and shape adjectives in the text descriptions. The main empirical claims are that (1) models generally show a shape-over-color bias when generalizing from ambiguous in-context examples, (2) this shape bias is amplified for image presentations relative to text in the 'standard' color-first phrasing, and (3) in text, models favor the first-mentioned descriptor. The paper introduces a new dataset of toy geometric objects intended to avoid pretraining contamination, reports results across three models and three paradigms, and includes bootstrapped CIs and regression-based significance tests. The authors are careful to note that the effects vary across models and task configurations, and they offer a pragmatic-language explanation for the text-order effects in the Discussion.

Significance. If the central modality claim holds, the paper provides a clean demonstration that presentation format changes in-context inductive generalization in VLMs, with practical implications for prompt design and for the cognitive-science alignment of these models. The work has several concrete strengths: a purpose-built image dataset that reduces pretraining-contamination concerns; three distinct paradigms adapted from the human category-learning literature; tests across three commercially relevant VLMs; bootstrapped confidence intervals; and an appropriately hedged writing style that flags model- and task-dependence. The 'first-mentioned adjective' bias in text is a crisp, falsifiable result. However, the significance of the modality claim is currently limited by a confound between presentation modality and adjective order in the primary image-vs-text comparison, as detailed below. The paper is a workshop-length preliminary investigation, and its contribution is best judged as a set of empirical observations that require the additional control analyses before the modality interpretation can be accepted.

major comments (2)
  1. [Section 3, Fig. 2; Section 2 (Methods); Fig. 3] The central image-vs-text comparison in Fig. 2 uses as its text condition the 'standard adjective ordering' descriptions described in Section 2, e.g. 'a red cube,' in which color precedes shape. Fig. 3 independently shows that the models generalize toward the first-mentioned descriptor in text, so the standard text condition is precisely the text condition least favorable to a shape bias. The paper does not report whether the image advantage over text survives when the text condition is the shape-first order ('an object that has shape cube and color red') or when the two text orders are averaged. Because the shape-first text data appear in Fig. 3 and the authors' own methods state that the different orderings were collected 'in order to assess whether adjective order affects the model biases,' this is an internal, testable confound rather than an external concern. If the image advantage disappears against the shape-first text condition, the effect attributed to modality would instead be a word-order (or word-order plus surface-form) effect. Please report the image condition against each text-order condition separately, and include a text-order dummy (or equivalent) in the regression that supports the main modality claim.
  2. [Section 3, regressions; Appendix B, Fig. 6] The statistical support for the overall modality difference is reported only as t-stats and p-values (Section 3: t(823) = 12.478 for Gemini, etc.), without the regression specification. It is therefore unclear whether the model controls for task type only, whether adjective order or text phrasing is a factor (relevant to the confound above), and whether task-configuration variables such as set size in the odd-one-out paradigm are included. The appendix (Fig. 6) shows that the modality effect inverts for set size 3 in the odd-one-out task, and the appendix text describes the one-category results (Fig. 5) as 'very noisy and inconsistent.' To assess robustness of the pooled claim, please provide the full specification (fixed and random effects), list which trial-level variables were covariates, and report a sensitivity analysis that excludes or interacts set size 3 and that separates the explicit category-learning tasks from the odd-one-out task.
minor comments (3)
  1. [Section 1; Fig. 2; Fig. 3] There are minor typos: 'idiosyncractic' in Section 1 should be 'idiosyncratic,' and 'boostrapped' in the Fig. 2 and Fig. 3 captions should be 'bootstrapped.'
  2. [Section 3, order-effect regressions] The t-stats for the order effect are reported as t(1681) for Gemini and t(839) for the other two models, but the differing degrees of freedom are not explained; please clarify whether this reflects exclusions such as the one-category refusals noted in the Fig. 3 caption.
  3. [Section 2, 'Shape vs. color bias plotting'] Please define the shape-vs-color bias metric explicitly in the main text as the difference between the proportion of shape generalizations and the proportion of color generalizations, and state its range, rather than leaving the reader to infer the definition from the figures.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the modality and adjective-order effects are direct measurements of model outputs with no fitted parameters or self-citation chain.

full rationale

All outcome measures are computed directly from raw model choices across the image and text conditions; no parameter is fitted to the data and then reported as a prediction. The image/text contrast (Fig. 2) and the adjective-order contrast (Fig. 3) are measured separately and neither result is defined in terms of the other. The only self-citations (Chan et al. 2022 for the two-category paradigm; Lampinen et al. 2024 for in-context learning) are methodological or background citations and are not load-bearing for the empirical claims. The possible confound that the standard text phrasing places color before shape (Sec. 2, Sec. 4) is a validity and interpretation concern about what the modality comparison isolates, not a circularity: the reported effects do not reduce to their inputs by construction.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper rests on domain assumptions rather than free parameters: that the three paradigms measure a common shape-vs-color inductive bias, and that textual descriptions are comparable to images apart from modality and adjective order. No invented entities are introduced.

assumptions (2)
  • domain assumption The three tasks (one-category generalization, two-category cue conflict, odd-one-out) measure a stable 'shape vs color' inductive bias, and aggregating them via a fixed bias metric is valid.
    Section 2.1 defines the three paradigms and Section 3 combines them into a single bias axis with task type controlled in regressions. The metric assumes the underlying bias is comparable across structurally different tasks.
  • domain assumption The textual descriptions are semantically equivalent to the images apart from modality (and deliberately varied adjective order).
    Section 2 and Appx. A.3 give the prompt formats. The comparison interprets image-vs-text differences as modality effects, but language adds pragmatic cues (Section 4), which is a potential confound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The in-context inductive biases of vision-language models differ across modalities." pith.science (2026). https://pith.science/paper/ICHXSD2T

@misc{pith2026250201530,
  author       = {Pith},
  title        = {Pith review of: The in-context inductive biases of vision-language models differ across modalities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ICHXSD2T}},
  note         = {Machine review of arXiv:2502.01530}
}
read the original abstract

Inductive biases are what allow learners to make guesses in the absence of conclusive evidence. These biases have often been studied in cognitive science using concepts or categories -- e.g. by testing how humans generalize a new category from a few examples that leave the category boundary ambiguous. We use these approaches to study generalization in foundation models during in-context learning. Modern foundation models can condition on both vision and text, and differences in how they interpret and learn from these different modalities is an emerging area of study. Here, we study how their generalizations vary by the modality in which stimuli are presented, and the way the stimuli are described in text. We study these biases with three different experimental paradigms, across three different vision-language models. We find that the models generally show some bias towards generalizing according to shape over color. This shape bias tends to be amplified when the examples are presented visually. By contrast, when examples are presented in text, the ordering of adjectives affects generalization. However, the extent of these effects vary across models and paradigms. These results help to reveal how vision-language models represent different types of inputs in context, and may have practical implications for the use of vision-language models.

Figures

Figures reproduced from arXiv: 2502.01530 by the authors.

Figure 1
Figure 1. Conceptual overview of our experimental paradigms. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Across three category-generalization paradigms, VLMs generalize more by shapes than [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The order in which the features are presented in text shifts the VLMs generalization biases; [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Example stimuli used in our experiment, showing some of the range of colors, shapes, and [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Patterns of generalization of a single category presented with varying stimulus sets. There [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Color and shape choices on the Odd-One-Out tasks across set sizes. With four or more [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Multimodal LLMs reproduce gendered instrument stereotypes across text, image, and audio inputs, with text showing the strongest and audio the weakest alignment.

Reference graph

Works this paper leans on

27 extracted references · 11 canonical work pages · cited by 1 Pith paper

  1. [1]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35: 0 23716--23736, 2022

  2. [2]

    Colour-name versus shape-name learning in young children

    Marc H Bornstein. Colour-name versus shape-name learning in young children. Journal of Child Language, 12 0 (2): 0 387--393, 1985

  3. [3]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...

  4. [4]

    Transformers generalize differently from information stored in context vs in weights

    Stephanie CY Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K Lampinen, and Felix Hill. Transformers generalize differently from information stored in context vs in weights. MemARI Workshop, NeurIPS 2022, 2022

  5. [5]

    The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements

    Sebastian J Crutch, Sarah Connell, and Elizabeth K Warrington. The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements. Quarterly Journal of Experimental Psychology, 62 0 (7): 0 1377--1390, 2009

  6. [6]

    Dreamsim: Learning new dimensions of human visual similarity using synthetic data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data. Advances in Neural Information Processing Systems, 36, 2024

  7. [7]

    Ordering adjectives in referential communication

    Kumiko Fukumura. Ordering adjectives in referential communication. Journal of Memory and Language, 101: 0 37--50, 2018

  8. [8]

    Are vision language models texture or shape biased and can we steer them? arXiv preprint arXiv:2403.09193, 2024

    Paul Gavrikov, Jovita Lukasik, Steffen Jung, Robert Geirhos, Bianca Lamm, Muhammad Jehanzeb Mirza, Margret Keuper, and Janis Keuper. Are vision language models texture or shape biased and can we steer them? arXiv preprint arXiv:2403.09193, 2024

Show all 27 references
  1. [9]

    Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019

  2. [10]

    Shortcut learning in deep neural networks

    Robert Geirhos, J \"o rn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2 0 (11): 0 665--673, 2020

  3. [11]

    Partial success in closing the gap between human and machine vision

    Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Partial success in closing the gap between human and machine vision. Advances in Neural Information Processing Systems, 34: 0 23885--23899, 2021

  4. [12]

    Logic and conversation

    HP Grice. Logic and conversation. Syntax and semantics, 3, 1975

  5. [13]

    Revealing the multidimensional mental representations of natural objects underlying human similarity judgements

    Martin N Hebart, Charles Y Zheng, Francisco Pereira, and Chris I Baker. Revealing the multidimensional mental representations of natural objects underlying human similarity judgements. Nature human behaviour, 4 0 (11): 0 1173--1185, 2020

  6. [14]

    What shapes feature representations? exploring datasets, architectures, and training

    Katherine Hermann and Andrew Lampinen. What shapes feature representations? exploring datasets, architectures, and training. Advances in Neural Information Processing Systems, 33: 0 9995--10006, 2020

  7. [15]

    The broader spectrum of in-context learning

    Andrew Kyle Lampinen, Stephanie CY Chan, Aaditya K Singh, and Murray Shanahan. The broader spectrum of in-context learning. arXiv preprint arXiv:2412.03782, 2024

  8. [16]

    The importance of shape in early lexical learning

    Barbara Landau, Linda B Smith, and Susan S Jones. The importance of shape in early lexical learning. Cognitive development, 3 0 (3): 0 299--321, 1988

  9. [17]

    Aligning machine and human visual representations across abstraction levels

    Lukas Muttenthaler, Klaus Greff, Frieda Born, Bernhard Spitzer, Simon Kornblith, Michael C Mozer, Klaus-Robert M \"u ller, Thomas Unterthiner, and Andrew K Lampinen. Aligning machine and human visual representations across abstraction levels. arXiv preprint arXiv:2409.06509, 2024

  10. [18]

    Adversarial training for free! Advances in neural information processing systems, 32, 2019

    Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! Advances in neural information processing systems, 32, 2019

  11. [19]

    Intriguing properties of neural networks, 2014

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks, 2014. URL https://arxiv.org/abs/1312.6199

  12. [20]

    What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models

    Tessa Verhoef, Kiana Shahrasbi, and Tom Kouwenhoven. What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pp.\ 199--213, 2024

  13. [21]

    Larger language models do in-context learning differently

    Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023

  14. [22]

    Word learning as bayesian inference

    Fei Xu and Joshua B Tenenbaum. Word learning as bayesian inference. Psychological review, 114 0 (2): 0 245, 2007

  15. [23]

    What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023

    Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023

  16. [24]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  17. [25]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  18. [26]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  19. [27]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.