Pith. sign in

Paper Citation Record · LEDGER

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.11945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11945 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:14:08.479610Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.980388Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f62878cd-7277-4c8d-a45e-65f364e40591 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.822537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.621753Z digest=sha256:09738a612c77816d5634a91e4b1fbdc4044f7ea0164a9636dbc4a4f0f6b8485f

Observation 2f9599b4-91f2-4f73-a3f4-4b4f9829b513 · outbound

This paper cites Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.804808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.632476Z digest=sha256:28733556eac3ffe97c96640b9430869039560536869ff74c2dea8761f9db9253

Observation 0861b875-b5a0-4ac6-8897-3dfa2c7dd57c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.639409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.639409Z digest=sha256:650cc65de76e61b4601e5ef4a7e38e2d97f12c669a118a16afd9c03e0e3bf990

Observation 9b1e6f6c-b87d-48a1-8d1b-2ff46a275d0c · outbound

This paper cites Honeybee: Locality- enhanced projector for multimodal llm.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Honeybee: Locality- enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.788084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.652965Z digest=sha256:571bf7d3d8508c21961a0b74015d93bb0dc5fcc183826400b03b4534dc817929

Observation 5e5dd523-61d0-4079-9fc0-db871464aa34 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.771427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.662831Z digest=sha256:a6555a28f440f5a13e52a7399fe24cbaae2ea430f8c45ec174c73adb84ed54f3

Observation f9dee378-b90f-4403-a288-3f01828d9a88 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.674264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.674264Z digest=sha256:190c451809addfc8dc9ae6c184943aac2bddb0979decfb3e02db11c16a9713c0

Observation 6d2fa2d2-90f0-4c95-bf16-5145408b2b86 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.681307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.681307Z digest=sha256:c4305fc08728353158c0e4c940dccba081f8389224fba51ea906284f4a96c53e

Observation 3d3af01d-f0fb-4eab-bff0-c6108479f7a8 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.687604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.687604Z digest=sha256:afdc7954810a2e03328dfa29e66d7c39ee436174d16e8d94bff7a22d77c2038b

Observation 77f1ecc6-94c2-4125-a27b-15255b2174bf · outbound

This paper cites Don’t look twice: Faster video transformers with run-length tokenization.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Don’t look twice: Faster video transformers with run-length tokenization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.717428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.694497Z digest=sha256:c5c7e24298eddd2365c66ab89b93eda13ff32e9429b0d7131c7de740715e82ce

Observation d8f9a9cd-0d68-487f-b0a9-7502e8ef8f87 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.699140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.699140Z digest=sha256:c6a92e7a9f5d147b6126c7e2b9fcfc75b55d312c3397809a7d34d8f34a95699a

Observation 31b2899d-2898-49e0-aba3-f4e768f28024 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.701206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.705318Z digest=sha256:04734da88bac9aec96f8ba0d481f867a7790608af2a20bcf23fb00a45e809e96

Observation 0df83c2a-7098-47dc-8e16-ab180eeb1650 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.713164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.713164Z digest=sha256:2fac12b7315b3afc4bc145cb2b2fd1794a36178aaf5f355a97c7f2edd3c667f2

Observation 12e9f6d2-3058-4da3-91a6-df0fd278f016 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.683415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.718760Z digest=sha256:0ff15b3c0322ae819c70278e6244c405e6bc7eeb8073302e00938a887115f853

Observation 16cdcd33-faf0-496f-96e7-ee146fc3d012 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.724986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.724986Z digest=sha256:8c51ec25a6ca1810b39637b896f86a8f293908e8c0fb99fb5d54072005e01c20

Observation 40832420-5bc7-414a-96e0-626514df2af2 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Efficiently modeling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.666969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.732650Z digest=sha256:8fae823f2d69a4c8f49a6e62cd130c99478a632818f2a42601c13a3c3e6a68f9

Observation 98a84017-9be8-4346-9658-a016bab80d0a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.647982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.739852Z digest=sha256:9b9621f887bc6fc6df95802391bc03ed8446bd4e0d74301f93726c9106d6deeb

Observation 58d24812-16da-4f2f-8774-ff66ed7f39a4 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.631083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.746571Z digest=sha256:dbef10aa1846bedc5eb7310611a4f40a594a8ad61d1735ab68f0682062766ab0

Observation 0cf33c95-b05d-45aa-a327-2103b752530d · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.608997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.752539Z digest=sha256:a622bf8c7dd54e7fb461cc6b7d2ec504d4c1eea11c8e94a8d0eaf158469312de

Observation 055fb796-142a-4db1-9318-de1e4de9a968 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.759340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.759340Z digest=sha256:b08b764bc9921bfd0b682acd899d9c2f8905867a2eb2be8c1e5dd5283d731a36

Observation d026d179-89cc-4830-9ce7-8dd29369ae1e · outbound

This paper cites Token compensator: Altering inference cost of vision transformer without re-tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Token compensator: Altering inference cost of vision transformer without re-tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.580654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.767540Z digest=sha256:bb7975e127d5b34d124490eba904bcfd573da03e749490b3cf4cd7d8f6d053b4

Observation cee4edba-c37a-4e81-92ca-93dbacb28d72 · outbound

This paper cites Logicad: Explainable anomaly detection via vlm-based text feature extraction.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Logicad: Explainable anomaly detection via vlm-based text feature extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.561972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.773640Z digest=sha256:3833416723876b95790c2da418c801aac33345bc6746c183a444d614f139cd52

Observation 56fef084-41e6-4492-88a2-c0fbfcbcb10a · outbound

This paper cites Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.545350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.780005Z digest=sha256:13eb8e35bf22f8c4155ec8865cf43c12f70eb13e14d8be8421edcdbfb3d849ad

Observation c870c746-a0ec-401c-8f2e-ef8df16d13c7 · outbound

This paper cites Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.528772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.786893Z digest=sha256:b914c32897f3f4ac5d0d933253340116246c4a1588509e5b8e82b507e920fae1

Observation 40462433-8ad0-4bf1-a6ce-6d052dd65f80 · outbound

This paper cites Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.510182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.795802Z digest=sha256:c6b2018457ba7e1ff08faf6ced798dc27848d16c8e39ae9eccfac7291da29e68

Observation 48131032-88e0-498d-b4d1-dba3fbf890d3 · outbound

This paper cites Lookupvit: Compressing visual information to a limited number of tokens.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Lookupvit: Compressing visual information to a limited number of tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.492474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.801917Z digest=sha256:fc7bd8620a86ca3b793c1ed07db2eeb32338fdb9c86439e20b72b5d3e5f9b370

Observation 16f1ee93-5146-4761-b480-e14b0c4e315a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.808159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.808159Z digest=sha256:6180328ef3ff2a905cd30a457793d4cc9cbcc47383e92d2dc929b9ee7e0a8262

Observation 04c32cae-143e-4ffc-9aee-35d91ae85dbb · outbound

This paper cites Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.463194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.814482Z digest=sha256:b26804f9d06ec5e6278a2e842b4ce27b638a243cfa619ba427d45b79bde622ed

Observation 3601f0f1-856f-4326-8837-e95d082e1208 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.820689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.820689Z digest=sha256:e06ad64b121ee31491092124a5da0fa8397ef8a643d4db0a843601c005a7154d

Observation d080d537-2f39-4dff-b5bb-e43875d8d4e6 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.829103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.829103Z digest=sha256:07fbf225fb54b13593cc9800bb9f36f8334b744b93a35b422f0e27c33309d05f

Observation 832cf5ed-d8b9-492e-921a-8ff379890722 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.444494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.836127Z digest=sha256:fe57aad58966de3197e75dff94e97d8e8d17bee123c2bad1b9975307ab6f59ff

Observation f1754f53-f7dd-44a1-b20f-3ddb0d75739e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.407845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.842487Z digest=sha256:be5d5645a4f74d01a45a6e11746b6db5cbdb1440c406623f75346f09f00fb156

Observation 82cf2356-dfab-4ab6-ac15-a58e88dd345e · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.850545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.850545Z digest=sha256:124a91ec5c4ed71de7249fafcf5013fbfa6025ff825f3731817a0abfedc8092c

Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.856988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.856988Z digest=sha256:c6e9ab83b195a1744e92000537d26be4ef3bb6d75517ff3f9ae12c32a6b6a309

Observation bb0d0afe-d458-4f1f-8384-fc77cc1ea76f · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.322034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.863990Z digest=sha256:78b19875e31cefde3d91475dd6be7dda3e8dcb9ef33876ef0a8c2789812b4736

Observation 204187f5-f6e1-456a-8ead-79c436535559 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.872063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.872063Z digest=sha256:41d4e779f64882424565e1314921a2ec40ab60f516e7110d9a9059ba9d3b3049

Observation e10ef480-4df8-47ca-afc9-4380f5a04b49 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Evaluating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.304591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.879546Z digest=sha256:cbe08966ff4d50e99d6ad18209751c80016beef7cde0d0c31e12dacbd348fb7e

Observation f4aab4f0-8be9-4c5d-aee7-460d243a8e03 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.288374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.885567Z digest=sha256:6134977e5a19272f8bc83cf74bc11a582adaecd1c5fd6151d91084763b962d65

Observation c32f603e-a5f6-491e-9037-dc2f2dfe2e2a · outbound

This paper cites Vila: On pre-training for visual language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.272213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.892662Z digest=sha256:d8772d947abfe7bef61eb410f6c3284cd6b1c4ffb953fcfafe154c95e7bc502d

Observation 12e92281-0c0f-4df5-b138-a885811ccef5 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.899314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.899314Z digest=sha256:4fb023d91c9d04a00cabf7ebb708e6493294e7cb4854d93cc643e935fcce85ed

Observation 2e34a6ef-d5ad-49b4-9c7e-9fbcb4e6a523 · outbound

This paper cites Improved baselines with visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Improved baselines with visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.907057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.907057Z digest=sha256:0aa3a7fbc21a0212a92f836a33a6bc0767a20dcb9914490889fa52fc0e6635cf

Observation b195c883-cf1e-4be7-8117-1de4ed7fd9f6 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.912723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.912723Z digest=sha256:4e65d4d51ca741fa23d0ae879c906ae9e5e20d8522dcbf1bc339dd36ee849667

Observation b99e44de-561b-47f8-bb64-419825bc6dc4 · outbound

This paper cites Visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.917778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.917778Z digest=sha256:99dbcb7f29d2c456930365285a2bea3023402fe9287dfef217831c5ae7958bc4

Observation 2af17887-9523-4eef-99ba-52009acf4f1a · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.924724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.924724Z digest=sha256:39edd896cfdedc9ae4706a57c25c9533c880d0efb175c77b5b7a2be78e2c7bd8

Observation 53b9be14-69f3-4160-b451-211b1d53e57a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.932776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.932776Z digest=sha256:a0571c788bb72d4bd927dfd011b6d99a6ec7c2891a01d0669873af9e6b7ab18e

Observation f529d4dc-f324-43c3-9a9b-4adefaa2e53f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.938326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.938326Z digest=sha256:3bdbe5e36630b48f74eeb0e9cf36a56a27fc2acb4f500bdd73ec894ce690d2dd

Observation 7d473a4a-7a10-416e-9a8a-d5879019c480 · outbound

This paper cites Questioning, answering, and captioning for zero-shot detailed image caption.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Questioning, answering, and captioning for zero-shot detailed image caption

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.203605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.944171Z digest=sha256:1167dbd46ad62086681722b5f1ffaa1c1748ea7f879e7345cca9d9176da78ac0

Observation 7ed2d803-7c77-41b0-a3a6-b31c97e84256 · outbound

This paper cites Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.183406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.950924Z digest=sha256:6beea951b668e483b0ac5f7355080c4fd045296edbabd925d38dc5c7c5183579

Observation 6b057137-b19c-434d-9e52-c74a4c018f8b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.956976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.956976Z digest=sha256:fde810a3832f7fcf8c95848d010c0229ef97bff801cb4789629d83141f1f57af

Observation 33740278-dc71-4c6b-a779-fae282c58dc3 · outbound

This paper cites Infographicvqa.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Infographicvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.162471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.962777Z digest=sha256:3ad8e65f40b89cd00d8f9533a03b4b8c399b925aae421851e1d869d880b4638a

Observation 44c17a7c-f79c-437d-89f4-eb8960cce2f0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Docvqa: A dataset for vqa on document images

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.142650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.968333Z digest=sha256:dcded7ecf166aabdcbd6e134c9f0c8cf84119a7bd3f8dc9d9f61b29b8118138c

Observation 404e1c35-cf10-4e80-b998-2b129e6b152d · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ocr-vqa: Visual question answering by reading text in images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.106857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:15.974801Z digest=sha256:705639a1adffeafcc7a5ec45e25dd0e27d413a1ba893b86ab9e88fc0c62a74c7

Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.981135Z digest=sha256:77a7e766611970b492d4fdef6c6915b622ec330092b3b1327a1adc8b358ec083

Observation 4c172126-ab99-41b1-96c4-0fb76ffaa218 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Learning transferable visual models from natural language supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.988182Z digest=sha256:cf14e8a593bab4e8276aee4c7328fd70923f690d5c2c0194698d9baeb45e3a86

Observation 524c67da-e1ee-419d-8dca-9bef68068679 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.996657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.996657Z digest=sha256:ba4070fcd3b8338503fb28e4bd248db2c59464a47111c399bdfcd6b7f14e2503

Observation 2f748f79-0d26-435b-8d19-8eed9c027152 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Textcaps: a dataset for image captioning with reading comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.005761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.005761Z digest=sha256:d8a34d6decf57d54c010fa97b8ed18fc7b317c2cc62dddf6277cad77b1548149

Observation d8c46163-dada-4f41-a7ff-6e278ad53d2a · outbound

This paper cites Towards vqa models that can read.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.014992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.014992Z digest=sha256:ac7a22aa0cdd074e843da7bff1f1999baa195be7b130c9126c6770805b537b3a

Observation cbc2e32c-88a0-4a0e-9659-ace1bc2b6246 · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.053996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.022616Z digest=sha256:22526530f9490fa15764826f34c95c5c14bc0eb740a68375b0d2d4dd064327fc

Observation 77bb4be7-8de6-4b4c-9f86-11a44302791e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.030070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.030070Z digest=sha256:1f22431ae8e9bc829ae5572787f56cc7bd7f36fc6f43e0d1470b2dbd7ac80a12

Observation e2c41836-a8eb-4564-bef9-77486e1a450e · outbound

This paper cites FastVLM: Efficient Vision Encoding for Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FastVLM: Efficient Vision Encoding for Vision Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.039140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.039140Z digest=sha256:6166defa7aad3b73c6b7e77970afe46cb86fd12bdcf5864706a600aba523fba6

Observation 1d4f5561-1963-4097-8a5c-13ca50a36610 · outbound

This paper cites Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.023069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.050031Z digest=sha256:588b254496fd3edb0106826e131ebcd57bb06d4c0a86d2f95b3c4248a1967ff6

Observation 2a1795cc-d234-44ba-9259-e91f96f51404 · outbound

This paper cites Fashionvqa: A domain-specific visual question answering system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Fashionvqa: A domain-specific visual question answering system

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.002072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.059553Z digest=sha256:a6b7f72600515f50ad606f99d8c1acebe90337595c0db443eb426e7ad0f936d1

Observation ae314ba2-7d11-4713-9bc4-f20418aa3413 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.069768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.069768Z digest=sha256:4a173c5a75a0968dab0b9f237120e50938c4d3bc202defedd7cdcd7a528ea6f6

Observation e9ba78d9-d7b9-4fdc-9969-9a291cf2c333 · outbound

This paper cites Rl-vlm-f: reinforcement learning from vision language foundation model feedback.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Rl-vlm-f: reinforcement learning from vision language foundation model feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.974689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.075698Z digest=sha256:7a0a62a767f8c2cbc20a61314aa1007b956b192a9f543a92192883950b20434f

Observation f8610864-5872-44be-8ce8-29b904cb20d5 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision- language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vary: Scaling up the vision vocabulary for large vision- language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.956137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.081522Z digest=sha256:ebc5a6d32654dc86734c5d42c7b3dc5224a1039b702b53627fe6e3189eac0ca9

Observation 31167906-34d2-4242-b1e7-f9d258976af0 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.088460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.088460Z digest=sha256:81066f9c36e31c89934bd0f761c9b655ca333d9af1102a78a89e06b58a9dc530

Observation 19fe73d8-5170-463b-81d9-29179927abc9 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.094655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.094655Z digest=sha256:e87529313b4bd91391e00c5c67bf0b8c22699fa1cdb4297ee4ed59705c104a05

Observation 3dce44ee-3c80-4486-8fa8-017cc21a1f7b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.099677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.099677Z digest=sha256:f378e967d7a9e2752405e0d2c07042493f7d7ab1319af730419fec33844d937b

Observation 39af3fc9-e525-42ce-b7e7-76b20d873fbd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.930470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.104787Z digest=sha256:74be81cdcd9871ace522220df3ac9b299e6a9e46e2de044d596bac82b0e7d985

Observation c5d5a9a8-76c3-467a-ba48-ed8a6b1fbd4f · outbound

This paper cites Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.911023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.110320Z digest=sha256:d212ee7c54cd825c2a36d2e52d11d5c88e2d7a0a6d8d9bab00283f5a2ea37394

Observation 5b262117-94e1-4ebc-a208-b54d27729888 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.116057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.116057Z digest=sha256:3314c63c291783a22100a9d3f638e35ff0ecf046b600ba512599ba461948a2fe

Observation 0e4b296a-43fb-4a93-82f4-ad1d4a8dabd1 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.889049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:48:16.122911Z digest=sha256:531ad5422aa711cce92d6a432f8045fa5b5b99c58ed13458393ddaad81e4b16c

Observation ec64c60e-2bcd-4618-af58-83360ab3a170 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.128579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.128579Z digest=sha256:e7d62e03b3a96c73c3c6c351377690c31228fe0957241adc91245e0414d90661

Observation bc6ac840-ab06-4b33-b0cd-7d4836c0c254 · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.136813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.136813Z digest=sha256:c0bc1afb849bc66c1849b8e314610b3c181ef4868989ed7f92bff838de2c31dd

Pith citing papers

Observation bbac5bfd-94e9-46e3-a2ec-9a76cecb0132 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.131202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:2bf200b1808033ebc7a72c99f5240e01f0f8664f0d96ae366c80b1f9ee599be8

Observation 0403b79d-6273-4590-8527-cf80380382d0 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.981648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:e0e18315fdb12b6186e07e7ef4bcc1b3e0fa156b75bf1ea4e16040c670d3e44c