Pith. sign in

Paper Citation Record · LEDGER

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 4 inbound Pith citation observations for arXiv:2506.05344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05344 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:08.640131Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:58:24.192980Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:19:00.613121Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1a487d4-c014-41c4-b463-c685878d72ae · outbound

This paper cites Pixtral 12B.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.449132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.449132Z digest=sha256:3ba3d3f18dd97cacdfab8024f6d82c8a9968519b2f04e9986354cc7f52b79b9b

Observation 30b24c70-51c7-4512-96b9-8039577fb7a0 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.454196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.454196Z digest=sha256:6b1b9527473d0c9284995a4d6cc1ec04cd8e08835f8f1f61e70155fcf070482e

Observation 00d6d9a0-2177-4bb9-8c8c-639e6a5d81af · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.455862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.458359Z digest=sha256:927a8f662f878e0ca885d7be7853cc5554d71c236b9127710aa3060dee834f7a

Observation d713e9f5-3123-4627-b5f8-88b5dfccacdf · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.462063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.462063Z digest=sha256:22c223963fb935f15ae5acd01364cc67ed0691edbdc216be3616138a3532f973

Observation 09c4e02b-ce26-43cd-960b-ae0e4de99916 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.445708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.465830Z digest=sha256:50799c5e3deca14eaa001942e33844d7828729c44324544b78c23126a79e2a1e

Observation 41b0d03f-1c55-4e2e-9529-cea3d74a40f8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.469397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.469397Z digest=sha256:3cc98c449e896e31485602ed4cdac3cd9b3a297c72484fd0bafc69eb92c7ed30

Observation 121a992f-c3df-4580-a1ec-5b88c8971bfa · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.435577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.472691Z digest=sha256:a0cf288d55684c42bfa110a553382ec9ecdb6b87cb62a42c96cd023e4ff9c807

Observation 8d896789-5e63-4b86-bccd-3b1e19ff9551 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.476233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.476233Z digest=sha256:5e4e5a7b55ce91e372277fb22eaa476f32a7e656a1e89f25c2dbf3a4165ce1a1

Observation 7d288c32-e4eb-41c6-8ebd-666ffc17a68e · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.418655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.480207Z digest=sha256:fb638addd3d98e8a3c2f711998f978aacb4351bbf741f81c2d8ccc4456d542d7

Observation 67242505-2a02-4a27-9a10-cf79db3fc8f0 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.484085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.484085Z digest=sha256:3388d06a4066b60e525c86f9e880f28bfc4a7a69111813d50d6cb16d653f47ce

Observation e5496a26-e016-4da9-a8fc-e4dda3744dea · outbound

This paper cites The Llama 3 Herd of Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.488311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.488311Z digest=sha256:e3c38fbbef055754f24745e9233ae6fbbb6f42e0a1ecc9a079a5183cd2fecc7c

Observation 7952014f-ff33-40b3-bd24-7d4a7bfe2962 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.491781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.491781Z digest=sha256:d0d64de069b9cd22f7c8facf9acf96d679f1b7aeef2424db8d2df697aa987162

Observation ac0ab6c1-82d3-4b2d-aca3-764827709fda · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.495177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.495177Z digest=sha256:f465ab4c80ca3d20b550c599f3505384f5416fba64d52a01113b790fbecacef2

Observation 6c2de33d-1360-49a5-acef-c9dfee767dcd · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.498592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.498592Z digest=sha256:d412482d4d8e5d34a2a19e5fcf4809adaf9b970f78c2f7927e87378ba87e4cb5

Observation 63365273-1d7b-4e50-ba2d-491ff0d3623c · outbound

This paper cites Making the v in vqa matter: Elevating 10 Table 5.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Making the v in vqa matter: Elevating 10 Table 5

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.408293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.502558Z digest=sha256:b528860e694cb8fbc24aad858ccec51e193d420a53d803bdebe1b93192d58c49

Observation c0c7cb53-1ab4-4faa-885a-f2ed5bf3b6e3 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs 3d-llm: Injecting the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.398366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.506621Z digest=sha256:98bbf355c6c8b55af524367ee97146c28d6d86d57d2282d503651c7b18e60ca1

Observation 09be408f-3fbb-42dc-8bc7-ec2a5a59dbff · outbound

This paper cites Mixtral of Experts.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.510563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.510563Z digest=sha256:1c091098818b6ad7160c3a4fbde98968eb1515d1c7e63b1565c04e15e126b539

Observation f7abf4ff-69a0-4b69-bd96-f571d0398808 · outbound

This paper cites OCR-free Document Understanding Transformer.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs OCR-free Document Understanding Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.514999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.514999Z digest=sha256:675bacc01bfe741c4c4ce9e1e0d3e12a2347a9b5cef61d1ef569acea095edaca

Observation 640f0432-a767-47d9-a292-4374171e99cf · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.518799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.518799Z digest=sha256:a4949e6b3a2c00ec23373164faf4b9952367c67653ac82141a1999e6d4f461e9

Observation 396755af-003a-4a95-979d-1e876c601d0a · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.522356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.522356Z digest=sha256:a8dc572dec895e33ad9e1dd83de6e753b59c24ccb5ff98d40f6db8535f544a04

Observation 566c391c-e7fb-4133-96e9-fc9bb910babf · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.385069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.526082Z digest=sha256:c6200a8d12110611634c9a813fe7b4387ca2492772b865f942bbc43aee5e771d

Observation 0d849c97-1c3e-48d8-8227-a99196537bef · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Snapkv: Llm knows what you are looking for before generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.374282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.529743Z digest=sha256:4094c41655d984296c7344c6f9420b0f234c76c53ce8852a2586fdd0a8c882b2

Observation dfdf24db-2228-4f65-a7fc-ac019be7592f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.533143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.533143Z digest=sha256:7caf0a5e32c1388c3a245c6774f3a7ef781ffc93aab0f530aa13c6b81b3312b7

Observation 7567f1c9-e52e-4032-85df-1142b59b0f03 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Vila: On pre-training for visual language models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.362742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.536916Z digest=sha256:933cce697a1de99d126160eb2899f10d5672a094653c7bc6b5d3aa6ff9187dfc

Observation 3d4782ca-92f8-4721-be6a-e396167148e9 · outbound

This paper cites Microsoft coco: Common objects in context.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Microsoft coco: Common objects in context

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.351838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.540172Z digest=sha256:595acd96ab70a0ce4a4256b1879e74e3739c46089dcf763e6c3272091f0a2615

Observation 3cd9a652-3352-4233-829d-792e61a63944 · outbound

This paper cites Improved baselines with visual instruction tuning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Improved baselines with visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.340899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.544030Z digest=sha256:7d57a00e0404232afd811d4f8f25d81b90e4a84c3b05e7c8903118904f67be39

Observation f7c7c1e2-b227-435a-b947-0953c3abe4eb · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.330141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.547308Z digest=sha256:09c193791c9a768718a383010b10dafeed9259286a2788a12e5b578ae43d8ac9

Observation a83c741a-2e47-4567-b496-19a971700bbd · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.550404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.550404Z digest=sha256:b8bf3cb3004e008833597b50903c033ec7c849f672b05f0cd8e7bd59bffa79ce

Observation 9578d87c-21c6-4210-923b-26d4dc0c536c · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.553683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.553683Z digest=sha256:1ac5bf7f7854b28450a8972211cd4a4b518d5b843a2c4e8046ae50a66745b175

Observation 899732fe-3962-44d9-9f9a-12075b47872f · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.556727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.556727Z digest=sha256:2ed1f5103a16007499e907753603d226be51fec7f9b20cddae421b84ab4baca4

Observation 880c1dff-0c4b-48f9-9bf6-96a0e4d84d16 · outbound

This paper cites Effi- cient inference of vision instruction-following models with elastic cache.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Effi- cient inference of vision instruction-following models with elastic cache

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.318901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.559674Z digest=sha256:26cfab6f789b5765cb0a3144a346cf4f8cadde6c0366fd000384f77f5453e464

Observation bf4ec20d-c45d-4fb2-90a6-085f09707196 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.562512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.562512Z digest=sha256:8b59344e803365ce47dd797323f75f61e6ede4df79921a93df9f166e7b561781

Observation 9e4548a8-5541-4013-8f38-477bf4129879 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.565388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.565388Z digest=sha256:4b4c541ad279a6e5b9714b42ef0bb63657c770c46732876c040834bb046a3d2e

Observation e633d34b-836b-43a8-adb4-eb8ea93ab707 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.568861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.568861Z digest=sha256:2fa0c934552c247c4ba18ff20975142b614c44175f130b06edd1c20f69ebb6da

Observation 1c093e96-8433-46a2-ae12-eaae3306a507 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Docvqa: A dataset for vqa on document images

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.306818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.572674Z digest=sha256:c8723b10e18e1aa0546fde52933baf1c55abbbe64bfaad4aac2f83641059a748

Observation d9cf2fbf-4904-4792-a442-a735996fbe93 · outbound

This paper cites Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.296422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.575877Z digest=sha256:06cbd001a76e276cb14f21c702bb6cb2b18fd4b3c91ba526df8fa621bc52f797

Observation a6100129-a842-4b65-b044-da096e0eb340 · outbound

This paper cites Openai gpt-3.5 api.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Openai gpt-3.5 api

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.285401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.578986Z digest=sha256:4e1e23fe17c0d2b06b58d6433f5f9c86ff269d3d8b99fe825daf5cd3d4b61167

Observation 6e54d431-a56a-42c6-a734-97427afd117e · outbound

This paper cites Gpt-4v(ision) system card.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Gpt-4v(ision) system card

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.269553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.581851Z digest=sha256:8a50d06b389018eb28121b85f3af2f901321c85a76425826644477d424e0e37b

Observation 59f5f77e-6697-48bf-8c7f-a9aa61506fd2 · outbound

This paper cites Hello gpt-4o — openai.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Hello gpt-4o — openai

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.256952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.584744Z digest=sha256:b8a28afaf1e651964c0135c3811920e1b8ddeb3ef81609ace41fd3cd8ebe3f30

Observation db801acc-99f8-4e8d-8389-511d707c6541 · outbound

This paper cites Qwen2 Technical Report.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.587561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.587561Z digest=sha256:bdb0cbf6e1c505ed1cd5b71b36c8bd5cae37cefce9e2d484dc6d18e874fa5cbb

Observation 26bcdf97-5519-4b34-a50a-92ea65acc44d · outbound

This paper cites Qwen2-vl: To see the world more clearly.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2-vl: To see the world more clearly

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.244338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.590765Z digest=sha256:a89d31ffb9632fc8c8c19342667500925c4e510f0999cfc4cb304898092a4ff1

Observation 939c2c07-1104-419f-970f-1bad73e39310 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Learning transferable visual models from natural language supervision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.233310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.594024Z digest=sha256:d9e21bf8a92b97515b4c04e413cd5bd075f3a5185c8b1f134512005cdaa304c1

Observation 0489a506-c797-44c8-af9c-9ff845903f67 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Textcaps: a dataset for image captioning with reading comprehension

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.222002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.597175Z digest=sha256:c15bde8fdeddd8adfd9c2cefd7b4a2ccf087b08652b10ac3b3dc6509a91bbd2a

Observation 6f2a7073-6f3c-4251-b5db-017bc3ff8172 · outbound

This paper cites Towards vqa models that can read.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Towards vqa models that can read

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.210343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.599910Z digest=sha256:480affdfcedb3277504e4e66fa84cd1781e7cf790f0a76899a794ecb9ddab183

Observation 48163766-a825-40a6-b97c-042baceaf8a4 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2.5: A party of foundation models, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.603979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.603979Z digest=sha256:82f2a6c744b50b169ac010d8ccd03c4a0190b7a8050cabade9f2fd4eaaac920b

Observation 0bb75663-bad7-414c-be79-be02a3ec5560 · outbound

This paper cites Qwen2.5-vl, 2025.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Qwen2.5-vl, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.607076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.607076Z digest=sha256:2decaf9523f0c0363538e1ece9262e99a6896cd1853262d00cb663e5db94219f

Observation 0b70eac0-5ef7-4e51-b076-1f80beee6f17 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.610568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.610568Z digest=sha256:9195ae89d14523a6aa517e8ce350477d69abb1a4b010c1b9df9243c88664fe62

Observation 55c11c4a-8f90-4ef1-9412-4ffe2239f111 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.614050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.614050Z digest=sha256:cce7730d67c1c733495084e047b20d68a042414bc01a1db480074904ad8a7c79

Observation 4389f7f5-64db-46bb-aade-3da369c6420b · outbound

This paper cites Attention is all you need.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.182854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.617520Z digest=sha256:2e51c6b767bf95367ddcba504b740ef787daf62f7b3d7f2e08587a346658738f

Observation 2a8c20f4-b3a0-4d56-8234-0aa88e43a211 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Efficient Streaming Language Models with Attention Sinks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.621107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.621107Z digest=sha256:3c49a1905b82e04d2faae5a7f05191a851ea3182ec2fbbd49bb6a57884cb2c53

Observation a948cb9a-ed4a-49d7-a9c0-9c532f81ba55 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.624340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.624340Z digest=sha256:f8c0fd15619f315c9c05635df9ba90ae550057037dd0f603b0d306f7a038a5a0

Observation 8a695192-8751-4bbe-a402-e1d30c00b363 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.627677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.627677Z digest=sha256:9706c759d19e6ba16eda251b9448a2413dd2f4768761b6fdfcfdb83753fba17d

Observation 84ec8c9f-a0de-44ff-85f4-0e26332e6a45 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.171594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.630668Z digest=sha256:3f1816344a458c9de2ff187f9cffdfb9dae3ef8ff6f6c075ea002c44b7a86681

Observation ddf2ac94-c5c9-4c65-8cd1-2f8063209dec · outbound

This paper cites A large chinese text dataset in the wild.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs A large chinese text dataset in the wild

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.159665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.633733Z digest=sha256:41a4942696a5addded9f1a9b9d7556dc547ff837b20089a717911012e812b8fe

Observation 38f2f6b7-7182-4e57-8b4a-32b8fc3e6976 · outbound

This paper cites Sigmoid loss for language image pre-training.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Sigmoid loss for language image pre-training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.148847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.637152Z digest=sha256:a5b9137f5ede710fb09f8b69d817a620ef642fb730bdb68518cf779f4342423e

Observation 6ac57d46-2557-46db-91bc-d5bb35590948 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:09.136099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:27:08.640131Z digest=sha256:70fa9f98a21b4888ce49cbd8caabbf9478c66a208f8bdf7c6a85f1eaec1f0b1d

Pith citing papers

Observation 0b1cdba5-adfa-49f9-afcc-4a467b4f0d1d · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:19:00.615863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:886a6d8906e1038857651bf13762daebe95f79130983f932c3f27af38c579931

Observation a261b1fc-71ba-4d0d-a62a-b64b28187b29 · inbound

HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference cites this paper.

HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:48.359879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:58:45.374139Z digest=sha256:5734e84cfaf597c4fa0ff1e6317d3481f3ccac0cf73a771e4abad30b374ecffa

Observation 212ef584-9dba-4f4d-b6bc-97a47fb067a8 · inbound

Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke cites this paper.

Vision-Core Guided Contrastive Learning for Balanced Multi-modal Prognosis Prediction of Stroke SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:49:44.490639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T04:46:46.504349Z digest=sha256:faecaf3863650d934362e490abb0695c3fc05d1eb57aa809c022399bffbc342c

Observation 20c69d2f-748e-4b1f-8a59-577e5f5e0335 · inbound

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models cites this paper.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.192980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.192980Z digest=sha256:c363e6cd9b035c60bc078576f7dfe18b19abeaa4690105648b02fc71e79e8621