Pith. sign in

Paper Citation Record · LEDGER

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 4 inbound Pith citation observations for arXiv:2506.05344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05344 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:08.640131Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:58:24.192980Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:19:00.613121Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1a487d4-c014-41c4-b463-c685878d72ae · outbound

This paper cites Pixtral 12B.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.449132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.449132Z digest=sha256:4b3b42ae19140dd1f8e5f5b15463cee1f1e710c2eccd37dfb10f7ee64c24b0f1

Observation 30b24c70-51c7-4512-96b9-8039577fb7a0 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.454196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.454196Z digest=sha256:b05e2dd455062d9996e02fdd2043cbaa9431e91f9b7bc5b503846b9a922f69a3

Observation 00d6d9a0-2177-4bb9-8c8c-639e6a5d81af · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.455862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.458359Z digest=sha256:d6aaac933b022f8bc8bc1a1ab1efc16e795ec7c2fca36a1d630a9c9d5fb74af0

Observation d713e9f5-3123-4627-b5f8-88b5dfccacdf · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.462063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.462063Z digest=sha256:e054ee49bde3d7cc190a3cf974ec5c9bedaa1fda66a96e37bc69342195cc6a4b

Observation 09c4e02b-ce26-43cd-960b-ae0e4de99916 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.445708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.465830Z digest=sha256:ee1bca9d535b38ff94cc4f2f49a6472c6debee5243a1b959da91cfc718b6dfb3

Observation 41b0d03f-1c55-4e2e-9529-cea3d74a40f8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.469397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.469397Z digest=sha256:b3eaa5aff20fdecb46e56953da01785f2584367a2e34932f8bf47715bf0844af

Observation 121a992f-c3df-4580-a1ec-5b88c8971bfa · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.435577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.472691Z digest=sha256:8465f6806e25a3c21434d1db337ee2cdb7aa597e87aa89f675b230112d44b2e8

Observation 8d896789-5e63-4b86-bccd-3b1e19ff9551 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.476233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.476233Z digest=sha256:a686bc308f4566696e443f67b4671fcdeff58349f69b5a1a713611316d8a1329

Observation 7d288c32-e4eb-41c6-8ebd-666ffc17a68e · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.418655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.480207Z digest=sha256:e24fd59963aa1fbe3d7e4d835c02bc78febea41984863360453ec0c8b00e8df9

Observation 67242505-2a02-4a27-9a10-cf79db3fc8f0 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.484085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.484085Z digest=sha256:a72a7acc78cd0449cd825ef306e514fe182ba8163af5d6a233ba8ca66f1b72ca

Observation e5496a26-e016-4da9-a8fc-e4dda3744dea · outbound

This paper cites The Llama 3 Herd of Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.488311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.488311Z digest=sha256:affd5d3c0050c4bccd640e3f08e029904dfb8cb055d4d14121ceeca681337c9b

Observation 7952014f-ff33-40b3-bd24-7d4a7bfe2962 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.491781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.491781Z digest=sha256:b7db9f966181cd88e972cfc4b6f78af742e9619815356f5bf2f131c28ea2aa8c

Observation ac0ab6c1-82d3-4b2d-aca3-764827709fda · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.495177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.495177Z digest=sha256:f30d1252df6399607a14b663a8dd1e9d59f15b8c7329f1dbb11dfce5014c23ed

Observation 6c2de33d-1360-49a5-acef-c9dfee767dcd · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.498592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.498592Z digest=sha256:c3a0967288d88197bdb22f0653a61c081d7033f04aade3f1bf663af25b02a13a

Observation 63365273-1d7b-4e50-ba2d-491ff0d3623c · outbound

This paper cites Making the v in vqa matter: Elevating 10 Table 5.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Making the v in vqa matter: Elevating 10 Table 5

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.408293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.502558Z digest=sha256:3524447e0f85bcf0561aa9e5aaa625c14eaf84dfc5b2e4404fb96950fb94116b

Observation c0c7cb53-1ab4-4faa-885a-f2ed5bf3b6e3 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 3d-llm: Injecting the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.398366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.506621Z digest=sha256:74a05d63c6b3a79f3cd42116a732edc41cb957095062fbd1de74de587487c317

Observation 09be408f-3fbb-42dc-8bc7-ec2a5a59dbff · outbound

This paper cites Mixtral of Experts.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.510563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.510563Z digest=sha256:ad72e89fc8eb324c9dc7fa599c2c84ec088f3e9ffa129624ef7777ef7b305253

Observation f7abf4ff-69a0-4b69-bd96-f571d0398808 · outbound

This paper cites OCR-free Document Understanding Transformer.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs OCR-free Document Understanding Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.514999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.514999Z digest=sha256:7fcb448435d512d0b8cb3458d397a511b48709444faa345f5c1a043f9777e85f

Observation 640f0432-a767-47d9-a292-4374171e99cf · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.518799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.518799Z digest=sha256:e8956d20a068426601779083820dcf5fa151be06b874b3a5cdeeed7ed9dd86c4

Observation 396755af-003a-4a95-979d-1e876c601d0a · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.522356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.522356Z digest=sha256:e393a6abd4885eab6f2d0cb04dbd744ccbd9c2b66ac3f168ee454cce5270b416

Observation 566c391c-e7fb-4133-96e9-fc9bb910babf · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.385069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.526082Z digest=sha256:f2363aa2059c5d6ada1a936b0c226a3710b212e8a950b70856a1fb146515dcee

Observation 0d849c97-1c3e-48d8-8227-a99196537bef · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Snapkv: Llm knows what you are looking for before generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.374282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.529743Z digest=sha256:08ec2a0f2620a62b43460d9cb6622caacaa66d25ad96c4daa5fcfd0cb905cb80

Observation dfdf24db-2228-4f65-a7fc-ac019be7592f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.533143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.533143Z digest=sha256:8af4b00ede66725340a07a14e1da83df28990d1f7a5b19bc9cb9ba6a8bceb0c7

Observation 7567f1c9-e52e-4032-85df-1142b59b0f03 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Vila: On pre-training for visual language models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.362742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.536916Z digest=sha256:f15d89b8ae6af32a20dd6db7759752c09c73dbf56f083d0c1cf91658d6d220ac

Observation 3d4782ca-92f8-4721-be6a-e396167148e9 · outbound

This paper cites Microsoft coco: Common objects in context.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Microsoft coco: Common objects in context

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.351838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.540172Z digest=sha256:a15f1c79fd65e7280357d68f3ea76a60ebceedc6786ac23cc514a0341399f020

Observation 3cd9a652-3352-4233-829d-792e61a63944 · outbound

This paper cites Improved baselines with visual instruction tuning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Improved baselines with visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.340899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.544030Z digest=sha256:9d32a693f9ac21ac1553c69d45cb1b26f0c98b658ec9e6f57ebad26c302479dd

Observation f7c7c1e2-b227-435a-b947-0953c3abe4eb · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.330141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.547308Z digest=sha256:4c539300c6db02c251f2ae783c09ea591019e367803c0ce20c0bf054df5f75cd

Observation a83c741a-2e47-4567-b496-19a971700bbd · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.550404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.550404Z digest=sha256:d5b59c5ce94cd68155d41168989585b73a72d396c49fb579667d20794bb887d7

Observation 9578d87c-21c6-4210-923b-26d4dc0c536c · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.553683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.553683Z digest=sha256:4f32ad49209848ee80ca564bf75db4afb57b4db9526193670968d1751886ea43

Observation 899732fe-3962-44d9-9f9a-12075b47872f · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.556727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.556727Z digest=sha256:5fbd64c7c426d315149cfddf9e50375490f4ef15cffa5e9404b0570cdf7680a7

Observation 880c1dff-0c4b-48f9-9bf6-96a0e4d84d16 · outbound

This paper cites Effi- cient inference of vision instruction-following models with elastic cache.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Effi- cient inference of vision instruction-following models with elastic cache

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.318901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.559674Z digest=sha256:e38767a62963d4ea70c23f137a36f446092ab8d9125ba7bf3e82cf55f64863b5

Observation bf4ec20d-c45d-4fb2-90a6-085f09707196 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.562512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.562512Z digest=sha256:8d1183a100bd1b5701f497b0d3d531c402bb916fd1b8e0ca2f993895bfa56433

Observation 9e4548a8-5541-4013-8f38-477bf4129879 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.565388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.565388Z digest=sha256:eee288ed252c4541ddd836ffd4d12bf968f1a8493ba679cb95a088bec05610c1

Observation e633d34b-836b-43a8-adb4-eb8ea93ab707 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.568861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.568861Z digest=sha256:f3501959879584539c6e67ec91e2459e2c507ab35cb23627329bcca91bd285c2

Observation 1c093e96-8433-46a2-ae12-eaae3306a507 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Docvqa: A dataset for vqa on document images

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.306818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.572674Z digest=sha256:c7994470e56f2752aec36a9bf6620f7b0e2b5874631da40ec40e95ecd0c7c9db

Observation d9cf2fbf-4904-4792-a442-a735996fbe93 · outbound

This paper cites Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.296422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.575877Z digest=sha256:7e6cc770393278289e332c40cd458e581a8a5cc65c35b9275eae34ea71b8af3a

Observation a6100129-a842-4b65-b044-da096e0eb340 · outbound

This paper cites Openai gpt-3.5 api.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Openai gpt-3.5 api

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.285401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.578986Z digest=sha256:fcfc0beee891c2907528813d1096ab26761e7afe5614cd6c3556d597b396eafe

Observation 6e54d431-a56a-42c6-a734-97427afd117e · outbound

This paper cites Gpt-4v(ision) system card.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Gpt-4v(ision) system card

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.269553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.581851Z digest=sha256:dcf529194931f1a7cea8d72909c5654eb4aa21b14f8e9161c9474fdf79780171

Observation 59f5f77e-6697-48bf-8c7f-a9aa61506fd2 · outbound

This paper cites Hello gpt-4o — openai.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Hello gpt-4o — openai

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.256952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.584744Z digest=sha256:8effcb0ea4ef22f0c44c74a9edb1bdd33de7429d31775b822795ba827fbead3e

Observation db801acc-99f8-4e8d-8389-511d707c6541 · outbound

This paper cites Qwen2 Technical Report.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.587561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.587561Z digest=sha256:11743e5840e31c3a7d0a12ec3d0b19528738360df80a0e9acf106b4299178e44

Observation 26bcdf97-5519-4b34-a50a-92ea65acc44d · outbound

This paper cites Qwen2-vl: To see the world more clearly.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2-vl: To see the world more clearly

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.244338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.590765Z digest=sha256:3d1c08d64743e74a7b81442f0d7ade8074498f1c476cb41fe32636d65895d703

Observation 939c2c07-1104-419f-970f-1bad73e39310 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Learning transferable visual models from natural language supervision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.233310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.594024Z digest=sha256:9554004534701ae9cb71db7caeddbc514ffd8efe981a41101f0c81c471905e65

Observation 0489a506-c797-44c8-af9c-9ff845903f67 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Textcaps: a dataset for image captioning with reading comprehension

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.222002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.597175Z digest=sha256:88852f90b6e10ef0bb8995fb476b76c7c852a92e1f2a1f17c6c578c843e5f4f2

Observation 6f2a7073-6f3c-4251-b5db-017bc3ff8172 · outbound

This paper cites Towards vqa models that can read.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Towards vqa models that can read

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.210343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.599910Z digest=sha256:cf0711b28f120aa3127a93df62f678b1d7ad28e0119d14f0342348f65e402dd4

Observation 48163766-a825-40a6-b97c-042baceaf8a4 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2.5: A party of foundation models, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.603979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.603979Z digest=sha256:b834be412fbaefa9eb7f279aad452be6e7d6a35913d9c515e4f6cb86cf848d38

Observation 0bb75663-bad7-414c-be79-be02a3ec5560 · outbound

This paper cites Qwen2.5-vl, 2025.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2.5-vl, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.607076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.607076Z digest=sha256:367f8b2ca84d874f92562eaa75bc30883c3435bcac207a3d29b7ef34a78e2386

Observation 0b70eac0-5ef7-4e51-b076-1f80beee6f17 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.610568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.610568Z digest=sha256:cd3d8c2a7b0c114faadeff3ea33efc6cfa1edcebeb5405a19d24f6c08dabeeff

Observation 55c11c4a-8f90-4ef1-9412-4ffe2239f111 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.614050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.614050Z digest=sha256:c9fd187395db56201b1081305f907fe079063acf3d0cd7433060adc7318f8661

Observation 4389f7f5-64db-46bb-aade-3da369c6420b · outbound

This paper cites Attention is all you need.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.182854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.617520Z digest=sha256:d015cb6f54ec8815b833aedeccd72e4b35ef6aee479cae39a9fb94300d531fd3

Observation 2a8c20f4-b3a0-4d56-8234-0aa88e43a211 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Efficient Streaming Language Models with Attention Sinks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.621107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.621107Z digest=sha256:d7ddcf52c3f274f59b1e9a4759f48d668177728fbb65e647982b943872da8e81

Observation a948cb9a-ed4a-49d7-a9c0-9c532f81ba55 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.624340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.624340Z digest=sha256:b9ca3dc80fe2cc86fe2625f9bec6422fb1154e336c4222ca9003d1c62777f98d

Observation 8a695192-8751-4bbe-a402-e1d30c00b363 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.627677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.627677Z digest=sha256:4652984f066b12e765cd6ac089bbedd7a6cd7d37ebe3feed25d09213063ecfdb

Observation 84ec8c9f-a0de-44ff-85f4-0e26332e6a45 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.171594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.630668Z digest=sha256:4332d9182cf4723e8144f57a56442161e11ee1d0267129be794965070b8398c8

Observation ddf2ac94-c5c9-4c65-8cd1-2f8063209dec · outbound

This paper cites A large chinese text dataset in the wild.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs A large chinese text dataset in the wild

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.159665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.633733Z digest=sha256:01be2654fe4ae2e9752104cf0461df617a35db804be304d6810ba3c0ff401502

Observation 38f2f6b7-7182-4e57-8b4a-32b8fc3e6976 · outbound

This paper cites Sigmoid loss for language image pre-training.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Sigmoid loss for language image pre-training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.148847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.637152Z digest=sha256:6a08e83d8815e832d75c5f82cce7f49fd558728529f9148cf47bbbf26f5377fc

Observation 6ac57d46-2557-46db-91bc-d5bb35590948 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.136099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:27:08.640131Z digest=sha256:b10c2737d9987b300db681b1ed00ea6ae3186a29f8f8245f4c9a0de355102632

Pith citing papers

Observation 0b1cdba5-adfa-49f9-afcc-4a467b4f0d1d · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:19:00.615863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:f5994d2db2cc7e299ea5e1a06c274a00a38134c08ad43e66b8b662efae024fce

Observation a261b1fc-71ba-4d0d-a62a-b64b28187b29 · inbound

HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference cites this paper.

HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:48.359879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:58:45.374139Z digest=sha256:bea2805d2ec7203534f5a2ee9b626c2d9758938bee722a5a39692d5458105408

Observation 212ef584-9dba-4f4d-b6bc-97a47fb067a8 · inbound

Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke cites this paper.

Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:49:44.490639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T04:46:46.504349Z digest=sha256:152779d3df593781cdf98a998720426f49bb1d64e7d13b50ec7bd8e6649d3cda

Observation 20c69d2f-748e-4b1f-8a59-577e5f5e0335 · inbound

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models cites this paper.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.192980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.192980Z digest=sha256:0bdbb0e69f0edae6b4daa8fcd8d473f5c0941ecbb526d7c34590ceaf5732b6c9