Pith. sign in

Paper Citation Record · LEDGER

Efficient Multi-modal Large Language Models via Visual Token Grouping

As of 14 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.17773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17773 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:25:00.399099Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T22:40:39.892802Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.026280Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5230cde1-3152-4f7c-b770-1f9f6c3b7551 · outbound

This paper cites GPT-4 Technical Report.

Efficient Multi-modal Large Language Models via Visual Token Grouping GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.306947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.306947Z digest=sha256:817989001c91167221e400254d75f113a0ec33823ccf53bd2b794feedd6e7bfc

Observation ef091b08-659e-444b-a3ad-8becc7b819e9 · outbound

This paper cites Qwen Technical Report.

Efficient Multi-modal Large Language Models via Visual Token Grouping Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.310224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.310224Z digest=sha256:6ad1e64b2fea41b58b64e9ebbf7cf23b1b0e03a61786ac07cd94d56144b4c2b8

Observation f2e6de39-00dc-4dfb-8d64-9cea8fe19938 · outbound

This paper cites Improving image generation with better captions.

Efficient Multi-modal Large Language Models via Visual Token Grouping Improving image generation with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.313096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.313096Z digest=sha256:4abae768755d3a9d3daa2db364651e3681dedc6734e9efd2d8e281f6d9f4802d

Observation afd818f9-a137-4b25-88c2-cf1002a6f4b2 · outbound

This paper cites Token Merging: Your ViT But Faster.

Efficient Multi-modal Large Language Models via Visual Token Grouping Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.315976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.315976Z digest=sha256:fe5588d232c9050cd88aaf438cebf45593fc190e3e049b12b6d941b57c0f4265

Observation 64644470-221f-407f-89bd-3744c2d2c8ee · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Efficient Multi-modal Large Language Models via Visual Token Grouping Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.318802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.318802Z digest=sha256:86f699ab9fd507b3048d2c106b315fd68bb3d3f474cd9671ae1afe78d62ac898

Observation cd12837a-9450-408b-9fd3-ad712256895b · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Efficient Multi-modal Large Language Models via Visual Token Grouping AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.321557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.321557Z digest=sha256:8a106182e405c17b5795e3bd33dee9e1bcd0cbb1822af5ba03503236ee927c06

Observation d47b8885-67fb-41c9-b767-57ae6ed2ef53 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.324543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.324543Z digest=sha256:14e5e7cf8081d37d40dbb462a3c035d69992380921dd358ddefd93a9c91e532c

Observation 4b1a71e4-03c2-48aa-9ae9-0176bd548584 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Efficient Multi-modal Large Language Models via Visual Token Grouping Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.327432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.327432Z digest=sha256:9cedf4474ed78582eb08a009634b13c68a274f9577fb3d513192a1c92046893e

Observation 09893b2c-e01d-47ac-9d3c-3cc20284bdd4 · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Large Language Models via Visual Token Grouping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:25:00.591947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:25:00.330264Z digest=sha256:06c07c82e4096c89e2122da87e98dde086f0cd279b0637fd69d40c77b438c2b1

Observation 214b2d25-4901-43fd-bf37-467ef56475b7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.332531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.332531Z digest=sha256:784ee70b3b184eefe2bac4a3db8b4ab4a3998ae90ed8d0c9b8eee74433e80858

Observation 9a977841-9611-4c50-928e-2d65fd479c4a · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Efficient Multi-modal Large Language Models via Visual Token Grouping Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.335430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.335430Z digest=sha256:8ab302ac08c65c4f8847bb193241c11599642eadbea7e8bb33933dff3757aeec

Observation 45b90320-8df4-49bd-b9f6-45ecfd133974 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Efficient Multi-modal Large Language Models via Visual Token Grouping Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.338172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.338172Z digest=sha256:fc7b1c3478b869d78962edbc77c593d54eeb778d89d93c77475c0cb648b1492a

Observation f60001f2-cd20-4437-804b-4f2d2c99ec73 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.340410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.340410Z digest=sha256:9cfb7141ff0e4d1e5b30043efb825706de4164cf02a8191f17a86445966dca42

Observation 1b4b2d2c-4d3e-4d45-874d-7c293341aebc · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Evaluating Object Hallucination in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.342758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.342758Z digest=sha256:acd520636d160cfb466f0904c88331c72c421085f21dad75a357e735a0fadd0e

Observation 0358a844-7806-4255-9a7c-cc027c1d2bf3 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Efficient Multi-modal Large Language Models via Visual Token Grouping Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.345323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.345323Z digest=sha256:c736e9ccd0762368de16597aedbd827031d35c6e23f8aecb912b8ce148e5553f

Observation 0442edb4-31e6-4af3-84aa-3c588239285b · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.348083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.348083Z digest=sha256:3331d2db36a5ff75bf3fed7d025e90a92ce5cd21bc77a9afbb1f5522fe7c887e

Observation 3c04a221-ead8-4f7c-9d06-9c06c01835e9 · outbound

This paper cites Visual instruction tuning.

Efficient Multi-modal Large Language Models via Visual Token Grouping Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.577189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:25:00.350609Z digest=sha256:71a3727c4a139cdd0a80d42199cc12d51f2b18fc0ba182b031dc34986c90aeaf

Observation c3ade4e1-85ea-4eab-89ca-8399625fe3f9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Efficient Multi-modal Large Language Models via Visual Token Grouping MMBench: Is Your Multi-modal Model an All-around Player?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.352950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.352950Z digest=sha256:0a93a19ce8902b4a6dd668b0f9eacb0fd26ba96b2cb03a1adda5bf81d52f82dc

Observation 57c8ee26-7b4e-4e2b-a9b3-98e9b32fed22 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Efficient Multi-modal Large Language Models via Visual Token Grouping Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.355823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.355823Z digest=sha256:286ab0a5ff9eb86caf443df458efb2763631105c6bb632448047971957e590e5

Observation 405752d7-7327-4b50-a1a7-d897cf44dcd3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Efficient Multi-modal Large Language Models via Visual Token Grouping Learning transferable visual models from natural language supervi- sion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.358347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.358347Z digest=sha256:351ae43497861dd25cfaf728a85730f41889f5f15f1be9c0a3d7b0af9d800ded

Observation 5a07fea6-86ee-449d-ae0e-ece8976773c9 · outbound

This paper cites Towards vqa models that can read.

Efficient Multi-modal Large Language Models via Visual Token Grouping Towards vqa models that can read

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.360675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.360675Z digest=sha256:08992ea539b4065bb2bad246c66f69f9e1e6f58373c25a7c72fb27868f17dab6

Observation 13052108-b106-43fe-8134-5a5b9cf59b3f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Efficient Multi-modal Large Language Models via Visual Token Grouping Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.363060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.363060Z digest=sha256:17f40b2e53c6f29fd44e2cb95b6bfea90b6ad7f5b61253d9e27b03224aa69e10

Observation 6956228f-437a-49b9-924a-2f03b47292b4 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Efficient Multi-modal Large Language Models via Visual Token Grouping Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.559249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:25:00.365716Z digest=sha256:8cca62063f417fd44922ad8a765d0a3b255e2946f5cbcd2482fcdd7ed3240fb7

Observation d9fce147-c50d-4345-a667-700ba991a541 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.368115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.368115Z digest=sha256:5e3e1925e463874fcc2601f855df0eca972e78d94be2b7782e3a6cdff9c8e85d

Observation 0dd63667-5baa-4b2c-a5ee-0bf1ad13ba94 · outbound

This paper cites Neural discrete representation learning.

Efficient Multi-modal Large Language Models via Visual Token Grouping Neural discrete representation learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.370673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.370673Z digest=sha256:5c92d48fe3b8f5240ddf9fc0db213c019c0466f5408497dcab0d057abcfe4b4f

Observation 5357ac85-5fa8-4cb4-9c01-a3e34be52516 · outbound

This paper cites Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data.

Efficient Multi-modal Large Language Models via Visual Token Grouping Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.372991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.372991Z digest=sha256:4dc3a98435e918aace825e66f1dae431078a340839a7fe73aa09ae6ec10a8dc5

Observation 0b66a686-60c9-43f5-8695-8dfcd767f616 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Efficient Multi-modal Large Language Models via Visual Token Grouping Efficient Streaming Language Models with Attention Sinks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.375740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.375740Z digest=sha256:6aedb7c8fb5df535b2a4d309c407768584889411c03edec35cdfe6149de41211

Observation 4df0f32f-6fbc-4d4e-b8ea-51feb0162fd8 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Efficient Multi-modal Large Language Models via Visual Token Grouping Groupvit: Semantic segmentation emerges from text supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.548185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:25:00.378156Z digest=sha256:0cd471b180f61f3bc5292a49155b6fdc58a032a4396331e0d9206979c19601a6

Observation 5ecc5ec3-ec4b-4125-a631-81b5b5c457dd · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Efficient Multi-modal Large Language Models via Visual Token Grouping The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.380558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.380558Z digest=sha256:94abbb0d6f857224745128e4c4ebc3cbf0c3cf3bc29be66fb221f67aa5dcff1c

Observation ae0cb8b0-25e8-4e56-a9ff-a24c72415139 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.383749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.383749Z digest=sha256:e8023c4189c678d6999326bb0701ae0ca784ead10aff2ff7529f4acffe09a9ec

Observation 49af9acd-90f2-4ab2-8674-a7dcf84265aa · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.386323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.386323Z digest=sha256:083075f5a5ddfa840d55f5bbb7ee315175e53f27b36aa3cf8e2b21d08f1b310b

Observation 4e842cc9-a5c7-4467-b472-628f30265967 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Efficient Multi-modal Large Language Models via Visual Token Grouping TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.388877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.388877Z digest=sha256:5655e674f2b64d067c9acdb1803f82bde3a5d4a3f33c249d248133e60b561830

Observation 1a9a7bdd-8385-4453-a25e-79d65249bb9c · outbound

This paper cites Sigmoid loss for language image pre-training.

Efficient Multi-modal Large Language Models via Visual Token Grouping Sigmoid loss for language image pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.391401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.391401Z digest=sha256:0a61442e343c6e477758c78bec468285b9fcbdee8f6cde9726e86e4271f3c3ff

Observation b6806a50-cea7-4973-82db-3383cef06e19 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Efficient Multi-modal Large Language Models via Visual Token Grouping InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.393682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.393682Z digest=sha256:8297fbb94131b3f6e879f248d84d3e5bc5ae255b64a7232b83d0c3e0550ce863

Observation 0471f26e-8674-48b7-bf4f-07ea6732e36c · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.396491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.396491Z digest=sha256:56f45f133d80b6086e7eba86dca8e0099046cf07dac714349b9a5388127a44e8

Observation c17ded9e-09bf-4628-a2fd-3688b23fad5b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.399099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.399099Z digest=sha256:c42c9ba79768f84f00261508ceb509122c3cc0166725b575ddbeac03aaa1f5af

Pith citing papers

Observation 87aad0bf-2f97-415f-9013-4e57a68f7c86 · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models Efficient Multi-modal Large Language Models via Visual Token Grouping

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.028908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:7fa0a5f925d12d82c4661e66824e1ce9e0bb1fe8b6fe6ff48555f941cbfe3b55