Pith. sign in

Paper Citation Record · LEDGER

How Visual Representations Map to Language Feature Space in Multimodal LLMs

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 4 inbound Pith citation observations for arXiv:2506.11976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11976 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:08.321921Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:56.147155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:26.125055Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3699c23-1c90-4763-8c70-44707168202f · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning, 2023.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards monosemanticity: Decomposing language models with dictionary learning, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:12.186433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.197245Z digest=sha256:6e00f8185e5970f4c1d0fafda458373cb01ebaf1e1bbb6e5ebd9e01ab54d4aaa

Observation 07ce8781-2025-4982-8acf-de1552b550a7 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.285352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.285352Z digest=sha256:6c1f0caa80aa7827aa72876791a985b07d15201094b0043af6ec9b8017cf3263

Observation 66ce3c5d-6bcf-4256-b409-918c0da41987 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.979935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.381515Z digest=sha256:d6b5d46b95a8461e32cab13de6e7c00f66be82049a138cd27d4c1968af118c90

Observation 7e7d4643-cb6b-4356-abae-16728f76041d · outbound

This paper cites The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.736793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.453896Z digest=sha256:e9bc94795caa8c006fcf0443a342a3f9161b77e0f703fd893c2418891bc7f5c2

Observation 7c89694a-e792-44b8-a725-35da8d656625 · outbound

This paper cites Arik, Tejas Nama, and Tomas Pfister.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Arik, Tejas Nama, and Tomas Pfister

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.548557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.522248Z digest=sha256:7c9f304799e3c02ead30475e03908de8f7d8bcd4c900fff00696ce4be8bb5509

Observation 5236b704-0e78-42c1-9656-03a507e1ed67 · outbound

This paper cites A mathemati- cal framework for transformer circuits.

How Visual Representations Map to Language Feature Space in Multimodal LLMs A mathemati- cal framework for transformer circuits

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.309715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.622671Z digest=sha256:ee171de55bbf6e226f81fde1f9c2e01b024489195673b4977e7b22d8128652ee

Observation b317eb25-217e-4507-9681-f25af730f1b5 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.715063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.715063Z digest=sha256:4bd1695fa6f2bcd8bcd465641f3ae4418691d50cc47063979baba819f54f1dc4

Observation 7a117c9c-7741-4049-a70a-582d506c319e · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.784761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.784761Z digest=sha256:ccea30030714601f9890b28c6350a28bab8db6fc52b9e0ed99004f66c2e6800d

Observation b1144db9-52c8-4570-8c84-dbfcb944205a · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Sparse autoencoders find highly interpretable features in language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.122586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.856612Z digest=sha256:a3dff8285a9e93bab0a235b6878715fcad6f1a9ffffe7249c31a1d1aa7317882

Observation 404ab531-414f-4122-a3e4-11bce77cf204 · outbound

This paper cites Hudson and Christopher D.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Hudson and Christopher D

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.917163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:06.898241Z digest=sha256:d8a5a8b8151bae842f397400e022c90427e6ca17315ddf47069ae55086b0e520

Observation 25f00dca-de3e-4cfd-883d-653b0b12c346 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.969709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.969709Z digest=sha256:a6e985afc56e855d976970555a06b5a7b7f79f30a8e0efc2ff29b6f6ab354225

Observation f8983eb4-1142-47d0-a99a-f09f64e827d3 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Evaluating object hallucination in large vision- language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.725401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.035217Z digest=sha256:75b09957245018a5fbfacb0e5be8c6b6686514a1bfc975b27de0bd2e07a37181

Observation bc6efdc3-a151-4fdc-8221-4521bf9c9819 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.108965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.108965Z digest=sha256:28d54e95341060c0b28f571f3ea575875f1c2c9e571e22171f97db2597c89ceb

Observation fd8c4433-db51-480c-9055-8664a4eeb97a · outbound

This paper cites Vila: On pre-training for vi- sual language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Vila: On pre-training for vi- sual language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.496925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.179332Z digest=sha256:1913599acb113ce0fd917d3081a552546ff70bc0c2393a6330c15588bf16088e

Observation effa39d3-0c87-4fb8-82dc-ef0a4cb5c2e6 · outbound

This paper cites Visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.260602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.260602Z digest=sha256:be9e2dbbf1f512c7c1eba765ecc2f607deaf7a8a217a9cedbdcbd5a9490aded6

Observation 84d95b6d-8393-4dd3-8645-5dce02161448 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.270709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.352145Z digest=sha256:fa68064979824aa2b1e165eec5587762c1cf5c5b26fb68ca744769ef8c96896b

Observation f6af390d-ded8-46ac-8748-6fbbe7850e37 · outbound

This paper cites Linearly mapping from image to text space.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Linearly mapping from image to text space

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.054001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.431580Z digest=sha256:983d3cf379914f294e083ce9b868452a3c3b755654c7a2ec705ceed05c73f3d9

Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.506978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.506978Z digest=sha256:2b1e17931d4460fd9a4c8538c803fa5ca54071dc23dcf688d882b418d105623c

Observation 847a9d7a-2e48-49a5-be4d-14db007b01ea · outbound

This paper cites Interpreting GPT: The Logit Lens.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting GPT: The Logit Lens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.841272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.567052Z digest=sha256:11d83e7f30d376cf017099c76d0b06d2886bbb77995c6af24acfe125c12194fe

Observation 5022e038-1a89-4d2d-97d9-b8f83fd8bd5a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.644453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.644453Z digest=sha256:f3578d3281ef1dc23098f24f6cff12c8422a83cd4b94f5f92cb72f4f7d63f67b

Observation d98f162a-bd02-4070-869c-7b659162ab52 · outbound

This paper cites Bridg- ing vision and language spaces with assignment prediction.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Bridg- ing vision and language spaces with assignment prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.654978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.743327Z digest=sha256:62a8f7dadc5eccff91bc93ce285173dcbcf9268d4bfe0580096d079c7befc4d0

Observation ca3e0edd-32ae-4375-bd68-ed81739091d9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.493256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.811626Z digest=sha256:8dbbb6539916f60b784320c48860ab48d4bc7f4c61f6ed1b4d49950a9d5aa8ce

Observation 5f3e4212-a992-4b9f-8e5d-e364113e2139 · outbound

This paper cites Multimodal neurons in pre- trained text-only transformers.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Multimodal neurons in pre- trained text-only transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.338873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:07.913318Z digest=sha256:b91f6253f3d9d80bb3b67fbf5db745b29b3da885c1f3bf892165bbee8379db23

Observation 54ecaa4c-e899-43b5-925b-20aa394b7bd5 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.977614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.977614Z digest=sha256:7a5b12197a55decdd557ac083bda6ff7db5e1bd932f1b41e263aa5a2545ef6f2

Observation 1937f5de-3784-4034-89c3-3e0a890480da · outbound

This paper cites Paligemma 2: A family of versatile vlms for transfer, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Paligemma 2: A family of versatile vlms for transfer, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.161200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:08.055607Z digest=sha256:8768bef45ebfda5dbfc2defed736480cc719b45776c67e7aac10beb3ed64ff64

Observation 2083da55-8b62-4185-94f9-0d851b8bdc27 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:08.123176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:08.123176Z digest=sha256:8975b681e87fd704f7a1ec4aef6175d3482397d7da16a9e104429a18fc3f287c

Observation b9431d42-e8cf-47df-9a14-6e6a98d5e4d6 · outbound

This paper cites an unresolved cited work.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:06:08.998912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:08.199039Z digest=sha256:d8f40b789e2c4262a4209c972cd50c1b6e4f8a713e95a0017ebf02e14d98be56

Observation 35c4f53e-89b3-4c81-81b7-9e89d9844c1c · outbound

This paper cites Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.844753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:08.244823Z digest=sha256:ae5dde567cbce424ac9aabeff5a83110ac724de1a4f49e1081cb4d0e0224a279

Observation 85d98845-988b-458e-8ad4-bda971176a71 · outbound

This paper cites Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.657295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T01:06:08.321921Z digest=sha256:c6e6b32b343ef6fec7798e5085bbeae3170e322f0d556165c52f53e7d979224e

Pith citing papers

Observation 7f81683f-e654-4363-ad77-03ebc2dffce9 · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.449816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.449816Z digest=sha256:f135b7e68aa6b485346a80950c2da3dc1ae4760508ca4a37c875cb218fa84216

Observation ab9ce2f3-d682-4f24-bbce-0a77ab76cb60 · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.126763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:4817476867bbd98b3fa4060646489258824541cc5b7da89fb0d199ef770197fd

Observation 4f28d0ec-44f2-4d3f-927d-2075dbb823b4 · inbound

Pathways of Visual Information Flow in Vision-Language Models cites this paper.

Pathways of Visual Information Flow in Vision-Language Models How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T03:03:11.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:03:11.110105Z digest=sha256:f82e396fbbfaf4eaa9b7a1110505aa1d47b531ad55fe258a924fc0f6e9b2c8e8

Observation e97f989c-c33a-4426-8979-5210adbd71a5 · inbound

Multimodal Model Diffing for Feature Discovery and Control cites this paper.

Multimodal Model Diffing for Feature Discovery and Control How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.147155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.147155Z digest=sha256:000937d5e2146d61fb2e28fd5a501be2bf2687d01a483e27d77f6bfed0fdab55