Pith. sign in

Paper Citation Record · LEDGER

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

As of 16 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.15816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15816 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:12.577103Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T06:33:36.846345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:34:40.990815Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8fc6bbe7-7958-4e92-bd2d-706296eb87c7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.474248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.474248Z digest=sha256:dfe49e3b03c84d9ec86f0b3498e1cd70e17522b5018a071ed3ab56295d8b74fb

Observation 8d8a034f-3e83-4553-b13a-e6c2bf8d0a3d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.862879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.862879Z digest=sha256:5864a8f1527b036d8d0e952fb1600c3e46739ecbb623d0adb227c67968e99f1b

Observation a33f1fad-b67d-4fe2-9011-183d95200ed5 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.933843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.933843Z digest=sha256:9655b20bd90cb689c8f9c278864f43de33b27d880216ef594ef505838d088b63

Observation f6891c77-ab35-42fa-b24e-ef3c1e9f02e6 · outbound

This paper cites The Llama 3 Herd of Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.001417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.001417Z digest=sha256:9574732a97494a09232bab5a9a0a4d8c9888d37e17a036e59b939dd59df0a24c

Observation 94df1105-1a71-4884-8784-4d314371a3d8 · outbound

This paper cites Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.049280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.049280Z digest=sha256:9ef54a357503de341f47cf0c6406e734b3e7df2607aae055e298c30ec0b4329d

Observation d2b73828-db75-412c-864a-12046d6b3fda · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.166214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.166214Z digest=sha256:4228f4eece4cda5533602c7a5c7a7a7bc54b76d0c2642b401629976d574dd614

Observation 1ec9bd7f-f068-49cd-898b-b626a761f3dc · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.230877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.230877Z digest=sha256:29c3c8b6b4420b16240aa2e077ecbeb3ea2c30713aca0ea4457edfeda256316a

Observation 94b725be-4ed1-45d0-8108-8c76e00512fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.315553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.315553Z digest=sha256:bc8084d7fae4f1ef4447ead229ff1afcbe1807cab50e996ad6233b0a9bda8142

Observation 196ae3f1-9380-4d33-b451-9513513ea221 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.437146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.437146Z digest=sha256:3face7a86b7c4d13d7f03c86820abf72b42bef4172efe1ba964542e24011e1b5

Observation 4a86901b-42a9-4789-a6b3-f2f91fcd9c1f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.683559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.683559Z digest=sha256:dc53e6fedc9e9e5c2de170e13316502f0becb2d917825345fe8f9d6e82a8eb41

Observation 0ab86f51-7054-440a-9b1d-2c94dd4b4332 · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.721014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.721014Z digest=sha256:26b51a30e37c3d5ffc6ba1a3e9712d0b83ab781827234f58368fcf2b98829bff

Observation e0eeda1f-9f52-4ae6-9519-2ecb36638b25 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.922497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.922497Z digest=sha256:64e18f56f0211e9081414e9a25fda349f14711388568a18aa1fff72191a3deab

Observation b2fbaa53-77e2-40d8-81f3-e9f1d92dd550 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.997946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.997946Z digest=sha256:085f5e8d4c1469f2d51161ae6030781d3d740dac32491671c3e64d7dd368ad3a

Observation 176cf2c9-49f7-4d31-8b89-614b50ad2a88 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.085010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.085010Z digest=sha256:05fdbf9471bea60525d141c13686cc1bb72cab7900822bb82440f7fc1957f983

Observation cccd9eeb-d8cc-44b5-a363-03d7c376bcdd · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.159249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.159249Z digest=sha256:b9a2b542ad14f3f824ea2dca5442f3d382c69f9ba15e03ef85291bed9b4e10be

Observation 4a2e253e-b2ae-4f1c-89e2-10eceb1764c8 · outbound

This paper cites Qwen2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.191635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.191635Z digest=sha256:b0a2fc593e167b090a3e3ffd8c0105b291c40b371af522ab9a2876060d90b16f

Observation c45523ae-bfe6-4356-b6b4-3f7ce6e92d7d · outbound

This paper cites Long Context Transfer from Language to Vision.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Long Context Transfer from Language to Vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.258703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.258703Z digest=sha256:a6e6056b1764ccc4096e15eff7e637e8895193956a82c08260de70da986b233f

Observation 709f18f3-12eb-4ec6-8bff-aff98e463c23 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.354668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.354668Z digest=sha256:01e4a560c99908526d48f10090c53516669678f02c2c9314dbb02095cb8b9797

Observation f61d0b3a-33cc-4d09-9490-858863c934ec · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.434364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.434364Z digest=sha256:fbd7a560c971ce14f1d9dc16928322c230100039561199df9be6ca7fc616b991

Observation bfd9234e-3dbf-4d97-8a2c-b8b439605a8a · outbound

This paper cites an unresolved cited work.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:13.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:15:12.501855Z digest=sha256:2f5177b539bf7aa77d3b69fa54774b3a93cebe4ffa4981438afd2521d5cb5547

Observation 5871d1eb-3091-4db2-ba61-b4acaf0125b4 · outbound

This paper cites Table 7.Evaluation on a comprehensive set of general multimodal benchmarks.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Table 7.Evaluation on a comprehensive set of general multimodal benchmarks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:13.371875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:15:12.577103Z digest=sha256:780e26444fbcf04f6d38bc977421e7638ddd179ff94e27bb35c1faab9f305589

Observation 4251acff-7482-4f44-af19-847fd10a068c · outbound

This paper cites Cross-Self KV Cache Pruning for Efficient Vision-Language Inference.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Cross-Self KV Cache Pruning for Efficient Vision-Language Inference

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.783480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.783480Z digest=sha256:e9e75c48477d762969551e477e2acbbeb955a78a6ac3ba022869ce3c5aa1cc42

Observation c1cb8a37-0ca7-474e-885f-539cee6ae2e8 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.612694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.612694Z digest=sha256:391bc31a3803e884a5c33e5ca8400f6285bebf77b68e7f11fb5f3f1806783c00

Observation 4daf7dd9-aab9-4b1e-9937-3369bc9fb644 · outbound

This paper cites J., and Yan, Y.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM J., and Yan, Y

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.852764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.852764Z digest=sha256:4b5d07fe53d6a23730f5f46f3d36ed87c8dc7b501085af142374a06b1addbe2c

Observation 75c46191-e6b4-40e8-a942-e0c02d96293f · outbound

This paper cites InternLM2 Technical Report.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM InternLM2 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.553356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.553356Z digest=sha256:14b707aaffd38e181a90b27bdd77f55e125132f919d795300c2c9d89a57a5e3a

Observation 8824895c-1718-43fe-88e5-984ba981523e · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM NVLM: Open Frontier-Class Multimodal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.736485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.736485Z digest=sha256:965560d1a2238f2b024d57f150059116d2a83e2ae72430f41eb218c66459cecb

Observation 519b9643-f2d2-4432-b1a8-638dcc103ef5 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.643408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.643408Z digest=sha256:e32c6f0a11e3b86f5f5b23ed9611ff2a5bb636585ce7a22a0bc3420720c647f2

Observation d6970a0a-68eb-4b8b-8841-49d7ff231925 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.527741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.527741Z digest=sha256:6e6ec1f1fb006f456328983ec043e35d072eb5e5fcbca37febea430cf439aac8

Pith citing papers

Observation f2b7f28a-61ff-4f35-b0ed-ca024f3c5c1b · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.992784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:c4e0a5cda877774aad815b85b4e74a8467bf5ca6c925c776cc6f9c3a3c6d67c9