Pith. sign in

Paper Citation Record · LEDGER

How Visual Representations Map to Language Feature Space in Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2506.11976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11976 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:08.321921Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.449816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:26.125055Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3699c23-1c90-4763-8c70-44707168202f · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning, 2023.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards monosemanticity: Decomposing language models with dictionary learning, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:12.186433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.197245Z digest=sha256:eb30ae780cacb2146ad7457eac3c3c65075e5190123d36d15a902db2979f6ff2

Observation 07ce8781-2025-4982-8acf-de1552b550a7 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.285352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.285352Z digest=sha256:0084b9ae65141d31bae1c64fff6f7368d873cef411170bc9019c6dd38a468318

Observation 66ce3c5d-6bcf-4256-b409-918c0da41987 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.979935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.381515Z digest=sha256:e2e566d0ce044dce9ad399517c45ee7204c90afac4a49e3c10316039646541b0

Observation 7e7d4643-cb6b-4356-abae-16728f76041d · outbound

This paper cites The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.736793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.453896Z digest=sha256:6380060357673607d07da2d803a38042495ba4e0da947361b3f2ded7e28e34a8

Observation 7c89694a-e792-44b8-a725-35da8d656625 · outbound

This paper cites Arik, Tejas Nama, and Tomas Pfister.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Arik, Tejas Nama, and Tomas Pfister

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.548557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.522248Z digest=sha256:a6e8c580f5a0ab46fd8c33a400c9bb693176a738476548ce69ac0900c989f99f

Observation 5236b704-0e78-42c1-9656-03a507e1ed67 · outbound

This paper cites A mathemati- cal framework for transformer circuits.

How Visual Representations Map to Language Feature Space in Multimodal LLMs A mathemati- cal framework for transformer circuits

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.309715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.622671Z digest=sha256:de0d724919c03040c1932dbfef898bebd4059a6ec763ca5e0eacedb8bbeb1429

Observation b317eb25-217e-4507-9681-f25af730f1b5 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.715063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.715063Z digest=sha256:cf12b9407b7d6261d51ad2a1af62a25c6f5eb4e059d8b6d2276a0698859860a5

Observation 7a117c9c-7741-4049-a70a-582d506c319e · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.784761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.784761Z digest=sha256:b5f08236ce5d9deb52b8727e70ea9065bb69c43507c9459f4d6b861cff74eb04

Observation b1144db9-52c8-4570-8c84-dbfcb944205a · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Sparse autoencoders find highly interpretable features in language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.122586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.856612Z digest=sha256:0401de18bac39ea89f68b838bf058fc3884670dbe6848f3bfb063ce55446ad4c

Observation 404ab531-414f-4122-a3e4-11bce77cf204 · outbound

This paper cites Hudson and Christopher D.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Hudson and Christopher D

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.917163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.898241Z digest=sha256:03f419e183b58b7a4da1d14bc90b98476ca8d2e2882a5eff2bcbbebc2a3bd688

Observation 25f00dca-de3e-4cfd-883d-653b0b12c346 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.969709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.969709Z digest=sha256:1a80045e96af1bc8a00c54e315ccd9b240a97d397643d978fa797bd24f521e7c

Observation f8983eb4-1142-47d0-a99a-f09f64e827d3 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Evaluating object hallucination in large vision- language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.725401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.035217Z digest=sha256:13a2d46c72196f606698200ff9e3ce31e6483dcf25ce035503dba56720c078cc

Observation bc6efdc3-a151-4fdc-8221-4521bf9c9819 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.108965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.108965Z digest=sha256:746bed6a57b454d96061e9c7867d0c26ff3e57dafd310918035415a39e90028b

Observation fd8c4433-db51-480c-9055-8664a4eeb97a · outbound

This paper cites Vila: On pre-training for vi- sual language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Vila: On pre-training for vi- sual language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.496925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.179332Z digest=sha256:2f9bbf249093230361ceda59d1b2b2eae733cdae93f0c694f94d6564fad64fa6

Observation effa39d3-0c87-4fb8-82dc-ef0a4cb5c2e6 · outbound

This paper cites Visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.260602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.260602Z digest=sha256:1aae24544579f1c0c5374de8ac22512170afdd8ab1bd592ac3a5b92f2d461270

Observation 84d95b6d-8393-4dd3-8645-5dce02161448 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.270709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.352145Z digest=sha256:85018a191d169ba203a7fa5eedd904dc5cf9ebdd93e2584e1cd24619bca3c9ad

Observation f6af390d-ded8-46ac-8748-6fbbe7850e37 · outbound

This paper cites Linearly mapping from image to text space.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Linearly mapping from image to text space

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.054001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.431580Z digest=sha256:801b7eba22411e6cf400aaa86704006f32856c38eb9358536e02931781cbd9b0

Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.506978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.506978Z digest=sha256:034439d712d17860e6868d91c021c0e08d6c368be547299278f7c3fecc65dd69

Observation 847a9d7a-2e48-49a5-be4d-14db007b01ea · outbound

This paper cites Interpreting GPT: The Logit Lens.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting GPT: The Logit Lens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.841272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.567052Z digest=sha256:6c6b0f251ded044b0fdd8bb57b2b7f5a2ec0cc18dd21070c38cd5e796ef0a1a1

Observation 5022e038-1a89-4d2d-97d9-b8f83fd8bd5a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.644453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.644453Z digest=sha256:b65864d0912f3fa043f0ff6d5a67fa8c4a09672985254512d86a050bbc1d7529

Observation d98f162a-bd02-4070-869c-7b659162ab52 · outbound

This paper cites Bridg- ing vision and language spaces with assignment prediction.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Bridg- ing vision and language spaces with assignment prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.654978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.743327Z digest=sha256:fe2f258d3a42701141f26caaea95c7c9cd95b7d9f204d80cd26884004ffd1ac8

Observation ca3e0edd-32ae-4375-bd68-ed81739091d9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.493256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.811626Z digest=sha256:1600c33c00dcb7be8cb2ff09a80c2307ab929281493c30bcea4413b4de45b44c

Observation 5f3e4212-a992-4b9f-8e5d-e364113e2139 · outbound

This paper cites Multimodal neurons in pre- trained text-only transformers.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Multimodal neurons in pre- trained text-only transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.338873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.913318Z digest=sha256:9106147d2d27cca838ba2d8868a0ce7d8e760587d4c04684d80e668aaa25e385

Observation 54ecaa4c-e899-43b5-925b-20aa394b7bd5 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.977614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.977614Z digest=sha256:5e67fc1af94fe11722d3c1eaf6e220577965c306fb0a7046cbc805f3307dd547

Observation 1937f5de-3784-4034-89c3-3e0a890480da · outbound

This paper cites Paligemma 2: A family of versatile vlms for transfer, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Paligemma 2: A family of versatile vlms for transfer, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.161200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.055607Z digest=sha256:1a3faa1524c199bc880ce62584a443bae8933b5fcf00984791b4b18fce3dc1f2

Observation 2083da55-8b62-4185-94f9-0d851b8bdc27 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:08.123176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:08.123176Z digest=sha256:669a8334e56c33adf6c31dca0b6197207aa60e1d2e6bea949d17c8a44343785a

Observation b9431d42-e8cf-47df-9a14-6e6a98d5e4d6 · outbound

This paper cites an unresolved cited work.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:06:08.998912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.199039Z digest=sha256:c5517c7e5d0b2242cc7e0f52d8b4b44e80f29da3c5795959f45a81e52b52a211

Observation 35c4f53e-89b3-4c81-81b7-9e89d9844c1c · outbound

This paper cites Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.844753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.244823Z digest=sha256:cb2e7551ca26599eeaff1e6e279b17513ef355f6c17565541da6d8a502862a23

Observation 85d98845-988b-458e-8ad4-bda971176a71 · outbound

This paper cites Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.657295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.321921Z digest=sha256:47090f04e629946eae49ed4e064c814fae3e5c23a8da29a4699c24ee9355db72

Pith citing papers

Observation 7f81683f-e654-4363-ad77-03ebc2dffce9 · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.449816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.449816Z digest=sha256:4a38c355456e170a11d2b5ce886fcae1fbc200740988cd8d2514e5e9cdc5bab5

Observation ab9ce2f3-d682-4f24-bbce-0a77ab76cb60 · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.126763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:7c7f36215d4f62c0e2e0b9bbd2e8d9e61f5bfe9b476ac7d7764cf315007c5b55

Observation 4f28d0ec-44f2-4d3f-927d-2075dbb823b4 · inbound

Pathways of Visual Information Flow in Vision-Language Models cites this paper.

Pathways of Visual Information Flow in Vision-Language Models How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T03:03:11.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:03:11.110105Z digest=sha256:d22e3383dd08e19010285f8e7d0ec593856928cab8dea33c17ee95d1cd7d58d2