Pith. sign in

Paper Citation Record · LEDGER

Learning Compact Vision Tokens for Efficient Large Multimodal Models

As of 15 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2506.07138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07138 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:46:07.815623Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:18:00.869557Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T22:18:04.010679Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51da7a16-600f-4cdb-aa5e-df1301accd5a · outbound

This paper cites Gpt-4 technical report.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Gpt-4 technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.224549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.693148Z digest=sha256:299861657d186400f32bb3e1f5b6e92de5051304ee0848f76875baa59055aad5

Observation 242480fd-5e69-4532-937f-d1e28933de3d · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.216143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.696975Z digest=sha256:8df78dfad8d11e2b8b0dd8159275edd651d4547f8fe4de5052eecc89060531b5

Observation 47b7ccb5-ecec-4232-aa66-fdd0e4a612cf · outbound

This paper cites Llavolta: Efficient multi-modal models via stage-wise visual context compression.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llavolta: Efficient multi-modal models via stage-wise visual context compression

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.207041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.700136Z digest=sha256:ec3fe89c7da37763d62dd8ef72910d0f545249a09371752fff80d36112243746

Observation 801ecf1d-d2b5-4ccb-8796-7d3024829c01 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.197947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.703600Z digest=sha256:8c9adc5b57eb25b5ae66c1b6cf75f87ddc7d57bf5523e6cf7bfdaa39ccebe49e

Observation a5725ba3-bb26-4992-8729-7bca59097c77 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.188962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.706893Z digest=sha256:6d48e3457f7c024c058897a0a184f84473c03f7d61ebd8ca6657c325ffa22a07

Observation 2daf0eb1-0da0-4355-a70c-aa749f051295 · outbound

This paper cites Funnel-transformer: Filtering out sequential redundancy for efficient language processing.NeurIPS, 33:4271–4282, 2020.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Funnel-transformer: Filtering out sequential redundancy for efficient language processing.NeurIPS, 33:4271–4282, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.179814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.710429Z digest=sha256:f015da2be8586dd83205d848af479dbb088e828fd432f8eef98c0eb058287a31

Observation 7bd3348e-dfab-4b74-9a06-0e3d3a1b3da2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Learning Compact Vision Tokens for Efficient Large Multimodal Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.170478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.713930Z digest=sha256:39422e146fc7fa003d509ab6e8ceb9c98fa2327684768353a81fbadbaf9c4e8a

Observation daba7e17-70db-4bb9-bec1-7ae646e68133 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Eva: Exploring the limits of masked visual representation learning at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.160697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.717719Z digest=sha256:259b7f2c94de6fc67b6a5f996c20d3978ca22aa2dc203b9c1790f491f8154366

Observation d92b89d5-1e62-425e-b63f-2aa2053392e7 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.151659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.720813Z digest=sha256:0c58aa1847e925f6ce88345c787d586c19ee4cf396e27fac48158d17e353eb19

Observation aca5367f-5d25-4ca1-a0fe-0fd6e65d4311 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.142867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.726782Z digest=sha256:4fcbbe88cfa5ba01b06f1c7601796fd65be4770dc704b9d93fddc820c43155a9

Observation d5bca0a3-8e1d-4418-af6a-33bd85c37fc3 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.133673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.730088Z digest=sha256:18d0b013987d25d02144db2aa48dd3ab2ccd086bb8656a7f452bd7d179c63028

Observation 8e06c92b-03c3-444c-a377-47214adbdf06 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.124924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.732964Z digest=sha256:ca78e905ada074a7054059309b9702696200fbf6c57d368028d7838ac1715512

Observation 4769ae5f-cd76-453f-bbed-d41799705769 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.115256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.735580Z digest=sha256:feb2cb41df43e6c7211697bf7755349b540b75d7f96e7cc6a11aebcbec30744d

Observation bfaac42d-eeec-4e09-8532-e67ba4abd52b · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Learning Compact Vision Tokens for Efficient Large Multimodal Models Gaussian Error Linear Units (GELUs)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.738326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.738326Z digest=sha256:ace675705d314e90e79c911075ff0f290ad2935d432dad259ecc3a50a143bcf5

Observation e66b3781-7f0e-48ba-ac60-57e40cc8b359 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.104245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.741473Z digest=sha256:334becccbf7c5e5736b0f4eb86d62439cf6ababa752a2224359d9a203027a028

Observation 950b0f9c-fc84-421f-8560-abb520e05b58 · outbound

This paper cites Phi-2: The surprising power of small language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Phi-2: The surprising power of small language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.094345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.744103Z digest=sha256:073de03300d4043ccd612eeee0a82de789bb614f4c203ef62e6783451cd9b685

Observation e73ff6e0-9ac9-4781-8b96-d51dee059a30 · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Tokenpacker: Efficient visual projector for multimodal llm

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.083761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.746877Z digest=sha256:44d88651e8effc3ec6699e88aadc52bb959de0f06fc2eebb75b1bede125ff92b

Observation e7917c06-20eb-4107-8422-23d6d2c4e03a · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Mini-gemini: Mining the potential of multi-modality vision language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.073771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.749749Z digest=sha256:9a4a2d5ba4d78f7f8dbfb54b5d3e57a8f2b497c4321e8d4e453dc6368ed895f8

Observation b47089e5-752d-45e9-b401-9ac7e88e70e1 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Evaluating object hallucination in large vision-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.063681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.752764Z digest=sha256:491c53f8471484804fb7f74e4f879296100b8c2fc6172071ad3ef14b24e7cba1

Observation 09630c67-f9b0-4bd9-bca4-aae99b6e66cf · outbound

This paper cites Improved baselines with visual instruction tuning.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Improved baselines with visual instruction tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.755519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.755519Z digest=sha256:db351437e73b1024aac21e682aa6e4ac82de5490581a6ed96cbd75f68bdf3601

Observation d5b7854e-9639-456f-b092-1c9da409e187 · outbound

This paper cites Visual instruction tuning.NeurIPS, 36:34892–34916, 2023.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Visual instruction tuning.NeurIPS, 36:34892–34916, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.758411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.758411Z digest=sha256:31d7e8e81a2e6f7b58b633d0fd5546dbfe8e361f1cd6195c69818847f91fe51e

Observation 3230107c-e0a4-43c6-a16d-116a70d0056b · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.039655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.761374Z digest=sha256:f1ff79d9616390f3ee77ac356e8f4b312b3045a86228295fe296f78365632afc

Observation f77dbf56-8bcc-49c9-b43f-de9eea84a03a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35:2507–2521, 2022.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering.NeurIPS, 35:2507–2521, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.029686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.763943Z digest=sha256:3fd25f6e0370370926bd86614c5c5e2b1471f54c8d3e677af614b80b7d566995

Observation c35a9f06-369b-4659-9e0d-b75570a45183 · outbound

This paper cites Are sixteen heads really better than one? 32, 2019.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Are sixteen heads really better than one? 32, 2019

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.019633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.766796Z digest=sha256:82279d56d0b68155c6f3c71de368207dc737f3c82d83c84626f43271f157ee91

Observation a5937e6c-aeca-4516-acf9-e22563477cc9 · outbound

This paper cites Efficient transformers with dynamic token pooling.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Efficient transformers with dynamic token pooling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:08.009085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.769515Z digest=sha256:cb04bd4a886d704668b59fe9860d42aca67db53c6dd83e4e4fc31a4c94dd0d9e

Observation 1974f68b-547a-4176-ac2b-2e6379512855 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.772408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.772408Z digest=sha256:1bf04ae8009af513bb55f65239a285220edf90cbe0de663e4bbdbb9dd6c937a6

Observation 65373317-e8cd-4802-be78-97e2c620473b · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.775586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.775586Z digest=sha256:1812220acd14235f79b985e29d631e8564dfe781a3da979f3f653923daf5e005

Observation 6c900fbc-8a38-45fb-8e84-9bc63e91248f · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.985357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.778748Z digest=sha256:9e6d9717368b2d3afe698ca4b1262e3af6f54e2c03a0ea50bfd33e968ffa97e9

Observation 34caa49b-be5c-4abe-ac0b-bb8a93e3d73a · outbound

This paper cites Inter-Instance Similarity Modeling for Contrastive Learning.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Inter-Instance Similarity Modeling for Contrastive Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.781803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.781803Z digest=sha256:64204713f011fbf45953d757d4028e488514b166485f5abf51dbe0d9cc4e684b

Observation 9472e061-3934-41ec-b802-547a485a9f57 · outbound

This paper cites Towards vqa models that can read.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Towards vqa models that can read

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.785130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.785130Z digest=sha256:7708466a5a4a4a05273d8d1ad7cf8fb8d5924dd2e413e366645bc3e642e7e07f

Observation bcd25046-4cd4-4547-bf4a-a9afc2d8984d · outbound

This paper cites Data-efficient multi-scale fusion vision transformer.Pattern Recognition, 161:111305, 2025.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Data-efficient multi-scale fusion vision transformer.Pattern Recognition, 161:111305, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.969118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.787948Z digest=sha256:47b0fb03acf5ebf127092a7925511b5eeee22aaa83b389c3abcebef4abe5c0e9

Observation 1e4993ce-e0cb-48f2-82c4-8bb6a410038d · outbound

This paper cites Gemini: a family of highly capable multimodal models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Gemini: a family of highly capable multimodal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.959421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.790822Z digest=sha256:1bf1a0992c23ee39a3749022338687f4b05ea5c39ef2a52be92fd4d70ab14fba

Observation 39e2a535-6473-4740-b768-40ed8d323bdc · outbound

This paper cites Llama: Open and efficient foundation language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llama: Open and efficient foundation language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.949990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.793639Z digest=sha256:15db3bd1b11b977f7173ce9858c05a729c98fe86968fa04b781a3f0a94df0d43

Observation e5e1a986-84f4-4c31-9c98-24dd65ce52d2 · outbound

This paper cites Attention is all you need.NeurIPS, 30, 2017.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Attention is all you need.NeurIPS, 30, 2017

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:46:07.796342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:46:07.796342Z digest=sha256:8daac53973d805526d19211078ca2a2f35247e2f076ec67031fa400d8922f53e

Observation 1aec081c-b095-4e6d-82fb-aba3f29c9209 · outbound

This paper cites Calflops: A flops and params calculate tool for neural networks in pytorch framework, 2023.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Calflops: A flops and params calculate tool for neural networks in pytorch framework, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.932441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.799344Z digest=sha256:dce60da81dd8a2e13dafaa86cfc81c3a810cd0d14d3d3bc31d94d1c5332f8754

Observation 20d354e2-bcc9-4296-84ce-84320e6d704f · outbound

This paper cites Texthawk: Exploring efficient fine-grained perception of multimodal large language models.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Texthawk: Exploring efficient fine-grained perception of multimodal large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.921814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.802200Z digest=sha256:ffc1ad7ddd1f9839b49a01baca9fbb01d93d0f4597cc1abe9a3707a4181e9161

Observation 9f0b0680-df8c-4f28-beba-fce3e75b9e52 · outbound

This paper cites Vcc: scaling transformers to 128k tokens or more by prioritizing important tokens.NeurIPS, 36:20260–20286, 2023.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Vcc: scaling transformers to 128k tokens or more by prioritizing important tokens.NeurIPS, 36:20260–20286, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.911031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.804847Z digest=sha256:4ef5d2260e53d1b3b699b27dd582865809f0e4ae665668a4b20b9b02cc805e61

Observation e8199289-00c0-46d1-8199-3dfd759be749 · outbound

This paper cites Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.901187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.807634Z digest=sha256:5c644e7ed2cbc6409e3cd94c41afcbecbb86c6a261450bb95249996d69592655

Observation 45577e7b-f0ee-4765-8fb5-8d49e649f5e2 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.889734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.810330Z digest=sha256:18c33741f6cc462726c35731c503e89e36449ebaf78936ea552923ceb30d2f26

Observation a8f7b169-0b02-4f72-a8ee-6909f489d59f · outbound

This paper cites Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.879090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.812915Z digest=sha256:49906b23bcadb216df5b9516d15a1fb91c3807605a240c88a85e48914fb29a1e

Observation a30af405-ec88-473f-a549-af8c0623ceb9 · outbound

This paper cites Treat visual tokens as text? but your mllm only needs fewer efforts to see.

Learning Compact Vision Tokens for Efficient Large Multimodal Models Treat visual tokens as text? but your mllm only needs fewer efforts to see

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:46:07.867788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:46:07.815623Z digest=sha256:7dd31d8adbfe07841640d4a382fa8fccfdfbd117ecc24cf3960be5ea66272bb1

Pith citing papers

Observation c71ed3b0-259e-4ab7-9df1-52cbaf754290 · inbound

SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models cites this paper.

SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models Learning Compact Vision Tokens for Efficient Large Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:00.869557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:00.869557Z digest=sha256:60f81dcd5c3225343ab5459095d0d1b8750ef4eef2fc4b4372fef5e678b872fa

Observation cb49eef2-ff18-401d-a49c-9f8578789311 · inbound

Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression cites this paper.

Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression Learning Compact Vision Tokens for Efficient Large Multimodal Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:18:04.014300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:17:47.203589Z digest=sha256:dc225b30472220ad36abbe1058efdfddcf842877d0b05ca42bd671dfac7b8923