Pith. sign in

Paper Citation Record · LEDGER

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

As of 23 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2508.20279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20279 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:52:36.195441Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T13:49:17.343893Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98858cfd-a862-4147-82de-a1a33a542fb1 · outbound

This paper cites Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:52:36.480094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:52:36.120038Z digest=sha256:e3b46be99f44c7f7db5c8db443ca979f3ef14418e7750b1ab206b06503769181

Observation 806ad925-4db3-45d3-8eaa-704092268396 · outbound

This paper cites The Llama 3 Herd of Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.124101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.124101Z digest=sha256:e8d10d984df1b899a31b002380f4fc1b6b028b91cb96d6ee47e14cd8730da665

Observation 7dfd39f7-2da9-4227-b077-c9105219996a · outbound

This paper cites Transcoders Find Interpretable LLM Feature Circuits.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Transcoders Find Interpretable LLM Feature Circuits

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.127974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.127974Z digest=sha256:5fa5a6dee4c644578ce1472e17bb0aaf86389a412697d05bead30009f09c8161

Observation e033c343-9192-40e3-ad5e-00400c0a5741 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Scaling and evaluating sparse autoencoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.132100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.132100Z digest=sha256:1a35e020354ba2a5f6f9d40ee05aa0f3a9c27e80cd26be6991113762df2e1edc

Observation b2656b8f-2356-46a8-8b2d-f17385bd7e63 · outbound

This paper cites What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.136411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.136411Z digest=sha256:cbf72be02ecadaefeb83e6616ec3da839425d2582b9a518724d7761c90b89a18

Observation 2f1d0792-cf1a-4ede-bc7a-fb26795b264d · outbound

This paper cites How to use and interpret activation patching.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding How to use and interpret activation patching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.144344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.144344Z digest=sha256:f69e2771e745ebe9788011af1dd41cd24baa055727a1a781d7e0bc9103737bd9

Observation aa626305-c01b-4707-b1a9-4c551602a5c9 · outbound

This paper cites MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.152984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.152984Z digest=sha256:fa7d1dde363b9064d4feb5273ec1b59d432ddd348e7d70ee86264980fdc3cfb9

Observation 315330f3-a536-4e49-b825-6dab7b2bf359 · outbound

This paper cites Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.157175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.157175Z digest=sha256:16361eac66fe7cbf8fe9e536a993abfa4555e156c8e04609cdbbaa7632e6398b

Observation de0f6f12-f9dd-4b7b-ad90-f5d2b918a6d3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:52:36.453104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:52:36.162890Z digest=sha256:06137f2e8ccd470db7e3af9ba7f6b272b44162a1fe8cd303e232e77ff09dcd5f

Observation d7f5f1c2-47ec-4ea0-acd2-ece3ba25cc58 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.166625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.166625Z digest=sha256:007fa61f338eeaeeb3e7932ca64cc634be15c185d504415cacaa56cf2d1da552

Observation 21e32dec-0b15-41cc-b523-251b9567f762 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.171048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.171048Z digest=sha256:ea286152afea5a17d418d4e339603aef282668acfab20ac9f106ec29578a34c6

Observation 8bf32ff1-16ac-4f5e-80d4-9b39ba6113a3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.174971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.174971Z digest=sha256:0d7ec8e61cc9b81d97c04399599b1cd073f46ad67388d6a96fccaa52e1d4bb64

Observation 18fb96d7-d4d9-42c1-b8ab-86104187e19a · outbound

This paper cites Probing Large Language Models from A Human Behavioral Perspective.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Probing Large Language Models from A Human Behavioral Perspective

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:52:36.283848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:52:36.179290Z digest=sha256:a370936748b7910f4a2b5bc4dee58a0f727050ef28ba79ec53d952b1eb8f7576

Observation 943f7d77-626a-4680-ac2d-994b9cd7e972 · outbound

This paper cites How Interpretable are Reasoning Explanations from Prompting Large Language Models?.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding How Interpretable are Reasoning Explanations from Prompting Large Language Models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.183560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.183560Z digest=sha256:ae12b87a4fac46433db558826bb8b1ca50013a552cf458c97a45e7eb9eb21060

Observation 28f9fed7-a112-4e81-b7ec-e4fa5bb813b0 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding OPT: Open Pre-trained Transformer Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.187710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.187710Z digest=sha256:f0082c25274e58b5ffecd6f455d9a252086f1c3c60f83cf3c73d91681409becf

Observation 49d9567f-a6ef-4923-90d6-94a88b39e795 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.195441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.195441Z digest=sha256:2fcda357de38aafd9afc679897aeb27829e01bb2d46a331fb1582b60494c0b05

Observation 7039e817-5b59-4788-9ce7-54c877c062e8 · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.140415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.140415Z digest=sha256:ca82db10b7d5b3392d330bc579ea58012930bacf4042496974e73123d36ec69a

Observation 6399d9f2-939c-4928-a073-41faf76a7ee8 · outbound

This paper cites Qwen Technical Report.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Qwen Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.102274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.102274Z digest=sha256:3d18eaf2389cc7e4130661017b12c9e2b91d4310c6f6e9487e20d8baf10ace8c

Observation ec95d0d9-3213-41fb-bdc4-79c81d0e13ce · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.115943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.115943Z digest=sha256:3fd2e89ab606d6616a773e4a961a259dd6b0c77e15c9c3c083914411fb61f544

Observation 97284a95-1243-46a2-99f3-ea0cc0e78305 · outbound

This paper cites A structural probe for finding syntax in word representations.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding A structural probe for finding syntax in word representations

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:52:36.467111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T16:52:36.148732Z digest=sha256:a93c6c7e839c4051763f2704d411f468ba8efcab23c73c85bb55827798fda70f

Observation 19697ec8-1d39-4cfa-9c0f-2a9f8313185c · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.191653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.191653Z digest=sha256:9288b4e96e8c300131ca87589208d82751226d60d9e311c249a28015d4fe1ba8

Observation 7a797298-bb6d-4bbe-a6d1-86b7ea7bc952 · outbound

This paper cites Understanding Information Storage and Transfer in Multi-modal Large Language Models.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Understanding Information Storage and Transfer in Multi-modal Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.107721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.107721Z digest=sha256:8b87e96af54b08fcb6d6ac99f86f8d8e1e1178265bfe233178ba78ee96a2216d

Observation 084f4cdc-dce6-41c9-98e9-189511aebde8 · outbound

This paper cites Language models are few-shot learners.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Language models are few-shot learners

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.111933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.111933Z digest=sha256:1a017cc76d61b1a2c02296f849dcedfd67e81dc80eabdd4d2d936790d9d5495e

Pith citing papers

Observation 325c9e4d-786b-4d8f-8499-f721c7ce50e8 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:09.429677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:114bc958bd88e848c736a933be5e8a930a3b0e2eaeea37f11baa12d9dbc990a7

Observation 6dbcd8c3-ad05-4cb3-9a1b-e44e89f2ed65 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.847233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:75c904d602833ce125eb1f83a27fcd30020f34ac29aae06e4323524e0dda6f1a

Observation c976d48f-b72b-45db-a734-c8aebcdc54a5 · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding

Reference 198

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:57:06.697512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:41b94b31d89c2350257fa28842731b0c9ea4579e83dfc272398c63874916b0bf