Pith. sign in

Paper Citation Record · LEDGER

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.11945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11945 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:14:08.479610Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.980388Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f62878cd-7277-4c8d-a45e-65f364e40591 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.822537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.621753Z digest=sha256:5e0e340430b0e920d875609af3c4c9383384c1d6b4ba5c58beb7cc290204208b

Observation 2f9599b4-91f2-4f73-a3f4-4b4f9829b513 · outbound

This paper cites Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.804808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.632476Z digest=sha256:a939841758f7a96d6b4c6a65d952a9f224c14d132da24ae2e63936f103f51fde

Observation 0861b875-b5a0-4ac6-8897-3dfa2c7dd57c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.639409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.639409Z digest=sha256:4432cfc113392fe499eeb93c320aff679eb22b46146bda4a9b070cbc7c11c968

Observation 9b1e6f6c-b87d-48a1-8d1b-2ff46a275d0c · outbound

This paper cites Honeybee: Locality- enhanced projector for multimodal llm.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Honeybee: Locality- enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.788084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.652965Z digest=sha256:d32c56c6034e19081ebb940c22cefa5ab2ca786c7faa394ab9671645ff5dd8d8

Observation 5e5dd523-61d0-4079-9fc0-db871464aa34 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.771427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.662831Z digest=sha256:7b018da371a90dd1416f9087b134940cd27ea93b82ddb17a9dab72de0ab21ad0

Observation f9dee378-b90f-4403-a288-3f01828d9a88 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.674264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.674264Z digest=sha256:bba1015289a09e0c68a891b5e1b171fa520d0cb24f77a25b728a27b707fd00db

Observation 6d2fa2d2-90f0-4c95-bf16-5145408b2b86 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.681307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.681307Z digest=sha256:611be525c596a9457c913b918a41c5390a1bf7b4ccebc52a8c20ecd6d0c2cd76

Observation 3d3af01d-f0fb-4eab-bff0-c6108479f7a8 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.687604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.687604Z digest=sha256:634582f18808d674ee9f943e0b51d2ad10a638f81922eb978bfdcc6271ef7de5

Observation 77f1ecc6-94c2-4125-a27b-15255b2174bf · outbound

This paper cites Don’t look twice: Faster video transformers with run-length tokenization.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Don’t look twice: Faster video transformers with run-length tokenization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.717428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.694497Z digest=sha256:ae40cfca9d92f8fd8b5834c027e109def79948403a7d17e275dad1f29d3465a2

Observation d8f9a9cd-0d68-487f-b0a9-7502e8ef8f87 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.699140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.699140Z digest=sha256:70394c2fa14ae4931062000636ec1f4aa41a310b015ed5a27d0131716063beb7

Observation 31b2899d-2898-49e0-aba3-f4e768f28024 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.701206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.705318Z digest=sha256:c23f3e89851d9df393d52a8228778191950da222f611261bcbb8ee4086590031

Observation 0df83c2a-7098-47dc-8e16-ab180eeb1650 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.713164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.713164Z digest=sha256:59e1e2650ce43c5d28811d06c7f7117dbff12e6b745757145aaa970ce5d33310

Observation 12e9f6d2-3058-4da3-91a6-df0fd278f016 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.683415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.718760Z digest=sha256:8cd0eb41e3c8d1ad040208c8807e75807337555f0a48fccd4552421847c67f05

Observation 16cdcd33-faf0-496f-96e7-ee146fc3d012 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.724986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.724986Z digest=sha256:4a9062eb5a1bfc2e439baee7f394543d05d9c9fbab5c1e491a19c8ce97a81f57

Observation 40832420-5bc7-414a-96e0-626514df2af2 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Efficiently modeling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.666969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.732650Z digest=sha256:327d6dc8023ba33f88b80a1313329535fd049277f70f99a8cb7e06ee840477f3

Observation 98a84017-9be8-4346-9658-a016bab80d0a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.647982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.739852Z digest=sha256:75cdddf21c73b02aeaf5844974764ec8421d187aa5b2916a0dc7bf985edc3163

Observation 58d24812-16da-4f2f-8774-ff66ed7f39a4 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.631083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.746571Z digest=sha256:e21d78ba1510c8a67d00aefd4ba5e6aee4c4eb3dcfbbd49cb43da41614c8e104

Observation 0cf33c95-b05d-45aa-a327-2103b752530d · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.608997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.752539Z digest=sha256:4d4e3d9429202d5e28650967499a33b74d57ec5f8684ee23c2e04976c19e2ac2

Observation 055fb796-142a-4db1-9318-de1e4de9a968 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.759340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.759340Z digest=sha256:d557060dcf1723da2df89c35abf40f9cf7c07e0478d08e6a07253f7ece1d6c1d

Observation d026d179-89cc-4830-9ce7-8dd29369ae1e · outbound

This paper cites Token compensator: Altering inference cost of vision transformer without re-tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Token compensator: Altering inference cost of vision transformer without re-tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.580654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.767540Z digest=sha256:119b04b30abbba015368e32a8136532638dee5cff21cde5008b9a174214b2367

Observation cee4edba-c37a-4e81-92ca-93dbacb28d72 · outbound

This paper cites Logicad: Explainable anomaly detection via vlm-based text feature extraction.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Logicad: Explainable anomaly detection via vlm-based text feature extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.561972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.773640Z digest=sha256:e22bc554e14ad1cf6c63af2d374c8a6e38bc8fbe3e4df919cf601d4ae2508de9

Observation 56fef084-41e6-4492-88a2-c0fbfcbcb10a · outbound

This paper cites Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.545350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.780005Z digest=sha256:25462092f8abcd84d3d0ad6d2f8116257ce3476f3c57be44ab7c938bcec30338

Observation c870c746-a0ec-401c-8f2e-ef8df16d13c7 · outbound

This paper cites Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.528772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.786893Z digest=sha256:86b9eba921e1eaf7c3d7e8a2613debccd1a1bf00532f51c47b20fffedced6622

Observation 40462433-8ad0-4bf1-a6ce-6d052dd65f80 · outbound

This paper cites Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.510182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.795802Z digest=sha256:9d22d2835de881493a7a6d1d2b7e043f746b73a7c6a9e44f53f2f97c216856b7

Observation 48131032-88e0-498d-b4d1-dba3fbf890d3 · outbound

This paper cites Lookupvit: Compressing visual information to a limited number of tokens.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Lookupvit: Compressing visual information to a limited number of tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.492474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.801917Z digest=sha256:60a8612231e7963bff5e08d771465afdfa1736ff6c5e75d6efede1ce2fb01df8

Observation 16f1ee93-5146-4761-b480-e14b0c4e315a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.808159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.808159Z digest=sha256:221e2b1a1ed178241ffa851c1584750bca9413813195bb6e3031e56e4393de28

Observation 04c32cae-143e-4ffc-9aee-35d91ae85dbb · outbound

This paper cites Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.463194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.814482Z digest=sha256:1fc7e55c5e30c2bbf8d38a88f11d66c11a44a71d55fee78af54e0a5fb48f6e2c

Observation 3601f0f1-856f-4326-8837-e95d082e1208 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.820689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.820689Z digest=sha256:f08bfc0aa48216c4f1425c0aa541c05b11f13a71faab4bedb83700f65c305025

Observation d080d537-2f39-4dff-b5bb-e43875d8d4e6 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.829103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.829103Z digest=sha256:6a7ad4929f2134e5c1be55a57b30cbf1f495d053dccc4505bae5362ef0f13409

Observation 832cf5ed-d8b9-492e-921a-8ff379890722 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.444494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.836127Z digest=sha256:c39cc8fcee323b373ae1eeda101970fd1affcf538885d1bfaa7b4fe6f663e30c

Observation f1754f53-f7dd-44a1-b20f-3ddb0d75739e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.407845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.842487Z digest=sha256:eb528c36ff5b1511f585ec9435562961676aeec3fa7bf877653760d621381851

Observation 82cf2356-dfab-4ab6-ac15-a58e88dd345e · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.850545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.850545Z digest=sha256:d515088473ebcbd3a03025a4f2494ae29f0c53a6f94968932bdb5d8c91c05d04

Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.856988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.856988Z digest=sha256:2a9ea77920c59b89674d81d5cd49b89ec55f7fb2ee681a5e70cd74a11b30b6da

Observation bb0d0afe-d458-4f1f-8384-fc77cc1ea76f · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.322034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.863990Z digest=sha256:3a39da8dd4cc3588ffeda72c2c5841224cbe5005bbf81d57223f19bad0eda695

Observation 204187f5-f6e1-456a-8ead-79c436535559 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.872063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.872063Z digest=sha256:c86b576f0a6d915bd54a12dfd05b3ab003f58b4a1a4eda1ed375b5436e7c6975

Observation e10ef480-4df8-47ca-afc9-4380f5a04b49 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Evaluating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.304591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.879546Z digest=sha256:4119adf38e45686214d43db0b67b92af20be20907cfd68327a94dab5f03f8018

Observation f4aab4f0-8be9-4c5d-aee7-460d243a8e03 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.288374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.885567Z digest=sha256:cd6881239e5457ef8a527c8bb887212bc94c70604311e2c2f7d72784df5e9397

Observation c32f603e-a5f6-491e-9037-dc2f2dfe2e2a · outbound

This paper cites Vila: On pre-training for visual language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.272213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.892662Z digest=sha256:d0faf8c886e6be0aa17faadb0cf382702b843428aa88123e43adb362d194e508

Observation 12e92281-0c0f-4df5-b138-a885811ccef5 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.899314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.899314Z digest=sha256:0b89b57e89870cc1ef1dc8f1765f3246951666f4b5ab6857096835849d729645

Observation 2e34a6ef-d5ad-49b4-9c7e-9fbcb4e6a523 · outbound

This paper cites Improved baselines with visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Improved baselines with visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.907057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.907057Z digest=sha256:62f07831f4453f559ae5c6c88536edeb1fb4ba1a6cf159b3da1f684d3a0f33f5

Observation b195c883-cf1e-4be7-8117-1de4ed7fd9f6 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.912723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.912723Z digest=sha256:cfbdca5614c4669cd4dfda6aab73178068a8f5cb2cebc13412d887a9c434a5e1

Observation b99e44de-561b-47f8-bb64-419825bc6dc4 · outbound

This paper cites Visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.917778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.917778Z digest=sha256:24dadc32b31a1680f0a2d4497919b5701855dea294b5938e894147469138e76e

Observation 2af17887-9523-4eef-99ba-52009acf4f1a · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.924724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.924724Z digest=sha256:fc01023da9a9abe7b8c4413f9276e093d4def5d2048c31886d1d32e2ba7dc001

Observation 53b9be14-69f3-4160-b451-211b1d53e57a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.932776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.932776Z digest=sha256:1019794451a2e0f04bd4ab021e209079c5fd5a7ea0182de89cbd2dadcb8855f6

Observation f529d4dc-f324-43c3-9a9b-4adefaa2e53f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.938326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.938326Z digest=sha256:654254a9e64c9f56ffd89ad2b8ece7f73954c71ca77aa95e89002c6aadb827a4

Observation 7d473a4a-7a10-416e-9a8a-d5879019c480 · outbound

This paper cites Questioning, answering, and captioning for zero-shot detailed image caption.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Questioning, answering, and captioning for zero-shot detailed image caption

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.203605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.944171Z digest=sha256:8f09ff208952d490bc5b8d420c6efb04e0866fc9f4e1527dfab5479ef7765694

Observation 7ed2d803-7c77-41b0-a3a6-b31c97e84256 · outbound

This paper cites Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.183406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.950924Z digest=sha256:bed2ffc159d3503971311200cd5856f43315e926c028aea976b848c9f480d2ba

Observation 6b057137-b19c-434d-9e52-c74a4c018f8b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.956976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.956976Z digest=sha256:f0d091d82096dc5bdb0b764aaffdf818b1747a578845eedc68f983714032844e

Observation 33740278-dc71-4c6b-a779-fae282c58dc3 · outbound

This paper cites Infographicvqa.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Infographicvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.162471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.962777Z digest=sha256:ff3ef0ed13c550a2862b48dd70fc21a296d9426f0aed16e90135372dbfb40482

Observation 44c17a7c-f79c-437d-89f4-eb8960cce2f0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Docvqa: A dataset for vqa on document images

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.142650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.968333Z digest=sha256:9fb4a5adcf1d9fe7b8a7a34e09a655d510f4fa8982a7bba0c42486be5002e0ae

Observation 404e1c35-cf10-4e80-b998-2b129e6b152d · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ocr-vqa: Visual question answering by reading text in images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.106857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.974801Z digest=sha256:837b304c2b34dd1b3204f56c52ae97492e37b6b1d1f9bc4cb3c31ac81382c799

Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.981135Z digest=sha256:73801e3d683d971ebe7f6aa2ae070f7c2a5ef8f4bbbb5ba420fdf4ee907f62c5

Observation 4c172126-ab99-41b1-96c4-0fb76ffaa218 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Learning transferable visual models from natural language supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.988182Z digest=sha256:1fc8f45c3b51dd0e266f553a736158b8293ccf36500065a95b5bf20d724c5ca6

Observation 524c67da-e1ee-419d-8dca-9bef68068679 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.996657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.996657Z digest=sha256:82cbd5f2286e2bd24de436f096ee16b19d5078c8ac70acb17e372ec7effb02da

Observation 2f748f79-0d26-435b-8d19-8eed9c027152 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Textcaps: a dataset for image captioning with reading comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.005761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.005761Z digest=sha256:132d8cd8094e295b1d17478ff0e4191e08eab80d407796707b3e14b5a1631051

Observation d8c46163-dada-4f41-a7ff-6e278ad53d2a · outbound

This paper cites Towards vqa models that can read.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.014992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.014992Z digest=sha256:b52edc0aef72d5d69b6a17836c3d7d7b8aa9540264f9afd55cacffc6fed59575

Observation cbc2e32c-88a0-4a0e-9659-ace1bc2b6246 · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.053996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.022616Z digest=sha256:bef0d000b6deff291cdabf2319850ff08296a496222e7169d024ab2f677a08f1

Observation 77bb4be7-8de6-4b4c-9f86-11a44302791e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.030070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.030070Z digest=sha256:44fc190c2f007cc6d266952ce6c471761d121e9b8b367b34bcdf309e479fec2c

Observation e2c41836-a8eb-4564-bef9-77486e1a450e · outbound

This paper cites FastVLM: Efficient Vision Encoding for Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FastVLM: Efficient Vision Encoding for Vision Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.039140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.039140Z digest=sha256:e2e7cf8754184f075de51a6bb871e06dc0e98652f0a4e3ad6f4bffe1b98f8b01

Observation 1d4f5561-1963-4097-8a5c-13ca50a36610 · outbound

This paper cites Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.023069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.050031Z digest=sha256:8499fa8225bd00d486029b48d353c725675f8abcf6aaf49b4e846d5f27e013e7

Observation 2a1795cc-d234-44ba-9259-e91f96f51404 · outbound

This paper cites Fashionvqa: A domain-specific visual question answering system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Fashionvqa: A domain-specific visual question answering system

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.002072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.059553Z digest=sha256:5834a69a55b139c171c5e25947c57031c3dbdfa392e6ae94e5a71ede327ed0fc

Observation ae314ba2-7d11-4713-9bc4-f20418aa3413 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.069768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.069768Z digest=sha256:8e3c11527832b8ee0b56c7661648c2ee2d43097359de03f7bdce466915cc69e6

Observation e9ba78d9-d7b9-4fdc-9969-9a291cf2c333 · outbound

This paper cites Rl-vlm-f: reinforcement learning from vision language foundation model feedback.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Rl-vlm-f: reinforcement learning from vision language foundation model feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.974689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.075698Z digest=sha256:6810eb314aacc1b281711f1704a544af967e10e671372870bf823a54e92dc991

Observation f8610864-5872-44be-8ce8-29b904cb20d5 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision- language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vary: Scaling up the vision vocabulary for large vision- language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.956137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.081522Z digest=sha256:a1b8ca3749dd03e53adb3ff50c0158945c2af79fef71fdfaf24bf940e8929031

Observation 31167906-34d2-4242-b1e7-f9d258976af0 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.088460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.088460Z digest=sha256:d11087fbbac307fb2c1a5dc1d963ed2d26735495505f4cb0255d6ec4c77b0f39

Observation 19fe73d8-5170-463b-81d9-29179927abc9 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.094655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.094655Z digest=sha256:bc4758f9f05ea05730cd3a5a14fe277178d37ad4ee8182073c712b529d3ab798

Observation 3dce44ee-3c80-4486-8fa8-017cc21a1f7b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.099677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.099677Z digest=sha256:1c76dea70796d170a6d6dcf327ec7c816888e3c126ba3de2adbe729641e956db

Observation 39af3fc9-e525-42ce-b7e7-76b20d873fbd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.930470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.104787Z digest=sha256:7fc68441ebc9f61abba379ee404b1d64a4526f1d6fe72fd78d4c5b372e18bfac

Observation c5d5a9a8-76c3-467a-ba48-ed8a6b1fbd4f · outbound

This paper cites Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.911023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.110320Z digest=sha256:0e9952c0524c41b53a1cff6be52585541b41d764f17b7b568db5ccc53f936383

Observation 5b262117-94e1-4ebc-a208-b54d27729888 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.116057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.116057Z digest=sha256:aca45439870aa58df86087c04701e529913669ddb06434e5fb84d545ee49db68

Observation 0e4b296a-43fb-4a93-82f4-ad1d4a8dabd1 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.889049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.122911Z digest=sha256:8c2faf6f397a2ffaa34427b25dfbca6cc0157e3480850b64a37f7a8005c0447e

Observation ec64c60e-2bcd-4618-af58-83360ab3a170 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.128579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.128579Z digest=sha256:eb25e4c7e8eb7948fd1f99b284b59ddef7d7eb417afd4fcc62b60bcb09f63c61

Observation bc6ac840-ab06-4b33-b0cd-7d4836c0c254 · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.136813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.136813Z digest=sha256:85e891f93abcaf107ce0c87903dde76b3353528b1079ae02e694b834480fd6fe

Pith citing papers

Observation bbac5bfd-94e9-46e3-a2ec-9a76cecb0132 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.131202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:6ce3aaad4051e0298347c8ae52827cdd091f478e1c09b67c250dc76b520b6967

Observation 0403b79d-6273-4590-8527-cf80380382d0 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.981648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:0de79ba04a263fc3e59c67d3ab2505dfbf85e6a6311dd9d62c765bdac008f625