Pith. sign in

Paper Citation Record · LEDGER

The in-context inductive biases of vision-language models differ across modalities

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01530 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:35.449560Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:38:00.558288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 632f3338-6350-4f67-8c87-82e17ddce842 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

The in-context inductive biases of vision-language models differ across modalities Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.327832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.327832Z digest=sha256:b6c26726864248e89265e72736b62b64b6fab26fcba7192317c19f5e303e30a3

Observation 29d0f044-b06e-44dd-8eef-d7a3706fb66f · outbound

This paper cites Colour-name versus shape-name learning in young children.

The in-context inductive biases of vision-language models differ across modalities Colour-name versus shape-name learning in young children

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.759882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.332776Z digest=sha256:b7c51449e5efae3ad9d16554fff78fd811a7243f237ebdeae1035882ca28f346

Observation 4ea83377-4818-4216-a14f-f0b6d31930d4 · outbound

This paper cites Language Models are Few-Shot Learners.

The in-context inductive biases of vision-language models differ across modalities Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.336573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.336573Z digest=sha256:5d87c19a8da348938111f91e5b0c4d4c1ccd75664c088424f7ec391f0e142377

Observation 1b099067-f2a2-4a7d-be50-53140eb69d77 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

The in-context inductive biases of vision-language models differ across modalities Transformers generalize differently from information stored in context vs in weights

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.745176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.340817Z digest=sha256:e4a4c5f6124fa0680e9b9fda07081c363e3743dc25178672ec98cf7e39cdeb3c

Observation 4e96c286-cfdd-4e2f-8f61-85c6f9002d54 · outbound

This paper cites The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements.

The in-context inductive biases of vision-language models differ across modalities The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.732719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.349461Z digest=sha256:c0cd82ce684f4efedcc5976d0d6fae99dc08b9152b5719a5cbc7b7c0ddfd6e18

Observation d5eadd47-6556-4f83-a50a-1ec1702b9579 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

The in-context inductive biases of vision-language models differ across modalities Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.719158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.355141Z digest=sha256:9aa24f9dcd5867290fd57a3b27324c398af8988ad728bf2949908be9f1557084

Observation aa95eb3f-1628-4291-87da-c49cba0aa8d3 · outbound

This paper cites Ordering adjectives in referential communication.

The in-context inductive biases of vision-language models differ across modalities Ordering adjectives in referential communication

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.706032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.359616Z digest=sha256:8930dde82cc1c5d84ed2a1f63bed7469d6dabe19b9355f60b56ef6321da85a3a

Observation 89178649-e182-40a8-8d17-59889033142a · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

The in-context inductive biases of vision-language models differ across modalities Can We Talk Models Into Seeing the World Differently?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.363001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.363001Z digest=sha256:d5451fbdbe0771640da04f114dde00cd817e4d53425b70604b3ab4f0b9fc7d8e

Observation 919f6bc9-a5dd-4998-ab2a-aa2ffe5cfe16 · outbound

This paper cites Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness.

The in-context inductive biases of vision-language models differ across modalities Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.366822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.366822Z digest=sha256:89087e4ec937c20a2bd52b2f8944c6267c23edee3555d658ef259091becaf2f3

Observation 9b6bdda7-6d07-4b76-a214-ac6f42ebfa02 · outbound

This paper cites Shortcut learning in deep neural networks.

The in-context inductive biases of vision-language models differ across modalities Shortcut learning in deep neural networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.370467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.370467Z digest=sha256:c5285a5361c1fc3f878fa227ed12e93d9ec8874aa531ca631e09c283ee08a389

Observation ed4d4125-e2b0-4a67-80f9-2b11c493c7a9 · outbound

This paper cites Partial success in closing the gap between human and machine vision.

The in-context inductive biases of vision-language models differ across modalities Partial success in closing the gap between human and machine vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.374251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.374251Z digest=sha256:6b005e1b81b195893c65d3b2f1ad77d7b3e97b5d9eb9203cf65620be53401ac8

Observation 5c12d530-980f-4d81-bdc0-1643d2ff597f · outbound

This paper cites Logic and conversation.

The in-context inductive biases of vision-language models differ across modalities Logic and conversation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.674507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.377893Z digest=sha256:f92766936cd60d5621172071f20b7d88a689e44eb7bf2507f94a32e334e495b6

Observation 3a864aec-8413-4d71-951d-e12c1d91c422 · outbound

This paper cites Revealing the multidimensional mental representations of natural objects underlying human similarity judgements.

The in-context inductive biases of vision-language models differ across modalities Revealing the multidimensional mental representations of natural objects underlying human similarity judgements

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.663386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.383206Z digest=sha256:70f7f8164fff031b579c1b1f1945e2dcb809bf37c8b71e19712d587b16c99618

Observation be61e304-e86a-44f1-b1c8-7701f53ed20f · outbound

This paper cites What shapes feature representations? exploring datasets, architectures, and training.

The in-context inductive biases of vision-language models differ across modalities What shapes feature representations? exploring datasets, architectures, and training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.650489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.389182Z digest=sha256:881fc306026160ee7038493bd51fe0b956c0f7b2e575f79917e7a29d461043e6

Observation a1404e10-6e47-4563-9899-1ca548c6b98d · outbound

This paper cites The broader spectrum of in-context learning.

The in-context inductive biases of vision-language models differ across modalities The broader spectrum of in-context learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.394368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.394368Z digest=sha256:c16c8647d612188497905a7e8efc776bc4dbfbf50e78c3794f4935ec9b0f6fa2

Observation 815f8d0e-199a-4082-a2fc-e0c1c85c004b · outbound

This paper cites The importance of shape in early lexical learning.

The in-context inductive biases of vision-language models differ across modalities The importance of shape in early lexical learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.638485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.398462Z digest=sha256:0e3c116f96e64db2e00153a0914874f660a56c03abe6d9f91eebac2135dff5d7

Observation 949b7e2e-0fcb-4415-bd04-dd239d8e855f · outbound

This paper cites Aligning Machine and Human Visual Representations across Abstraction Levels.

The in-context inductive biases of vision-language models differ across modalities Aligning Machine and Human Visual Representations across Abstraction Levels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.404895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.404895Z digest=sha256:5291977456ddef0a0fc9cc88cde0029532af24e65e16e11e87e29a28bf9b7d92

Observation 56766abc-6cfe-4bbc-9c7e-adae61da79d4 · outbound

This paper cites Adversarial training for free! Advances in neural information processing systems, 32, 2019.

The in-context inductive biases of vision-language models differ across modalities Adversarial training for free! Advances in neural information processing systems, 32, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.412082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.412082Z digest=sha256:25746f0fb9003471c3a64c0262719854758a0f5b9cbc5e1922f97b9d18f474db

Observation 02f32883-6932-4d63-916f-9d85c6e4a1b1 · outbound

This paper cites Intriguing properties of neural networks.

The in-context inductive biases of vision-language models differ across modalities Intriguing properties of neural networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.416921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.416921Z digest=sha256:1c51332ab9b068d6208376c5dfb8862b04fd5f624c1843ff3d7481e323a818d3

Observation 3dd74afd-215a-4f9c-aea2-66a248d522d4 · outbound

This paper cites What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models.

The in-context inductive biases of vision-language models differ across modalities What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.616811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.421381Z digest=sha256:894c877cae98cbaaa35b4bf75f83ecc8dd09c9b2817ff4c9395d74c896a7f88b

Observation 23296d28-0426-4f7e-b581-d67b34d7db99 · outbound

This paper cites Larger language models do in-context learning differently.

The in-context inductive biases of vision-language models differ across modalities Larger language models do in-context learning differently

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.425399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.425399Z digest=sha256:e131a3e4a746f9a6ec39c217725049b5eeb5a5b1158b982954537108c1189cba

Observation 2c06b120-441d-44ef-b34e-7e0eff0858d3 · outbound

This paper cites Word learning as bayesian inference.

The in-context inductive biases of vision-language models differ across modalities Word learning as bayesian inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.604365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.429873Z digest=sha256:b60ff424f1cef05bce2498a5d8dc384e20e01fe4c2fb77d43aa5aa7576067632

Observation 0fa84074-350b-4d98-b226-d267066c5e05 · outbound

This paper cites What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023.

The in-context inductive biases of vision-language models differ across modalities What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.433581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.433581Z digest=sha256:40e861fdc7fa8b3dfdc69fe1e04cd9216e2957d3c61dc717b27940f632e5e8fe

Observation b3d06aa3-6000-40c4-8135-710786633bab · outbound

This paper cites write newline.

The in-context inductive biases of vision-language models differ across modalities write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.437148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.437148Z digest=sha256:36bccb08d46a814859599e401cd1e0e208c27b596507d0e7fee994da6441161f

Observation 8ac483a6-3e46-4cfc-b126-80dbfd4a62d6 · outbound

This paper cites @esa (Ref.

The in-context inductive biases of vision-language models differ across modalities @esa (Ref

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.441637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.441637Z digest=sha256:9d678fa1c334c1357c6bc3a5fbae780854ee0b99acee83c0245c338b2561ba4b

Observation 8b0d6d83-99cb-4a88-a8e6-dbfcbc83db96 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.445663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.445663Z digest=sha256:c8ef033e19eec66e59dc61de7d18e846d12854ddf13d7b483c0dc205fe27ff4b

Observation 9337145e-ba16-43af-83e0-bb3fd8fcbe85 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.449560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.449560Z digest=sha256:7e5691f3eb8c94f21a89caff31470e989a481ca6f186285f622dd555d8099ef1

Pith citing papers

Observation d06f6240-a08d-4c04-9a8b-60082a742f1f · inbound

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs cites this paper.

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs The in-context inductive biases of vision-language models differ across modalities

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-01T17:38:00.558288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:38:00.558288Z digest=sha256:ee7a6cdd6489d329fab0a5b8ca3d153ce5bba5c120080ae0828af601b1ae0000