Pith. sign in

Paper Citation Record · LEDGER

Efficient Multi-modal Large Language Models via Visual Token Grouping

As of 15 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2411.17773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17773 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:25:00.399099Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T22:40:39.892802Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.026280Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5230cde1-3152-4f7c-b770-1f9f6c3b7551 · outbound

This paper cites GPT-4 Technical Report.

Efficient Multi-modal Large Language Models via Visual Token Grouping GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.306947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.306947Z digest=sha256:32f7f96c40d47fc1e2141acba269b1366b2bf460bd1ea0a04c1b28027e8b1b32

Observation ef091b08-659e-444b-a3ad-8becc7b819e9 · outbound

This paper cites Qwen Technical Report.

Efficient Multi-modal Large Language Models via Visual Token Grouping Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.310224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.310224Z digest=sha256:e50c319c2be973b564324c59637071086168f4e55b956cddcb72987ae8a327c2

Observation f2e6de39-00dc-4dfb-8d64-9cea8fe19938 · outbound

This paper cites Improving image generation with better captions.

Efficient Multi-modal Large Language Models via Visual Token Grouping Improving image generation with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.313096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.313096Z digest=sha256:8e81f02baa3102665badd84c665706bf06f3fb63b91880ce71de0212f6842dd5

Observation afd818f9-a137-4b25-88c2-cf1002a6f4b2 · outbound

This paper cites Token Merging: Your ViT But Faster.

Efficient Multi-modal Large Language Models via Visual Token Grouping Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.315976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.315976Z digest=sha256:34df756070c91ec22fae313fcece9f102567fab420fef2a22d7de7f414dae971

Observation 64644470-221f-407f-89bd-3744c2d2c8ee · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Efficient Multi-modal Large Language Models via Visual Token Grouping Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.318802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.318802Z digest=sha256:ecd6d1bd7f26283dd57e998404252e0b65e1d10b54a833a1ec1942dd228d4621

Observation cd12837a-9450-408b-9fd3-ad712256895b · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Efficient Multi-modal Large Language Models via Visual Token Grouping AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.321557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.321557Z digest=sha256:0ce4cda6dc19cf550acc9ca481926fca6f4f3f14b3b16e375126955fa4fb8bb1

Observation d47b8885-67fb-41c9-b767-57ae6ed2ef53 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.324543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.324543Z digest=sha256:a87c2dadf0dbe186417e27842cd35db5198a406799f8f3b5b1b9268a84d44637

Observation 4b1a71e4-03c2-48aa-9ae9-0176bd548584 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Efficient Multi-modal Large Language Models via Visual Token Grouping Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.327432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.327432Z digest=sha256:8e43580598bffec77cfd16290ff7fc3e3a8706281c1c8863ed1ddffce4ef913f

Observation 09893b2c-e01d-47ac-9d3c-3cc20284bdd4 · outbound

This paper cites an unresolved cited work.

Efficient Multi-modal Large Language Models via Visual Token Grouping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:25:00.591947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:25:00.330264Z digest=sha256:3ae5dc9124a575acdc8c22a738ce0c84d8578d5ca879cfd7f62a2668effafdb9

Observation 214b2d25-4901-43fd-bf37-467ef56475b7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.332531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.332531Z digest=sha256:a3b66adc33fbd663e271db174e15f81d33482fd6af40409596ebade278a592ce

Observation 9a977841-9611-4c50-928e-2d65fd479c4a · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Efficient Multi-modal Large Language Models via Visual Token Grouping Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.335430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.335430Z digest=sha256:da126079d8026e99f1aa62f7f098530ee22b7ceffd32c687689fc3fe66ab4bab

Observation 45b90320-8df4-49bd-b9f6-45ecfd133974 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Efficient Multi-modal Large Language Models via Visual Token Grouping Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.338172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.338172Z digest=sha256:572b9e4807cedf5fdc0af53423dbb4cca8b90eb6b695b0d6637a54a1c28656c9

Observation f60001f2-cd20-4437-804b-4f2d2c99ec73 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.340410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.340410Z digest=sha256:82e346f8a80790289826fdf5afae3ffdb4ba119797c218daed43e3d3a57e26e2

Observation 1b4b2d2c-4d3e-4d45-874d-7c293341aebc · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Evaluating Object Hallucination in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.342758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.342758Z digest=sha256:4a51aa5c37e5b21f4686ea9be8ed47f9b3bbe4eb9a749ce40dd4d6300004f963

Observation 0358a844-7806-4255-9a7c-cc027c1d2bf3 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Efficient Multi-modal Large Language Models via Visual Token Grouping Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.345323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.345323Z digest=sha256:9a0a558200a4d5da0178156c8c0be0d1dde86493a9e7b84ba817536e45357598

Observation 0442edb4-31e6-4af3-84aa-3c588239285b · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.348083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.348083Z digest=sha256:dcfeebd9dffa255bcbd1ec3c9416b7c09357f504b6ba48f92da9260d82edd989

Observation 3c04a221-ead8-4f7c-9d06-9c06c01835e9 · outbound

This paper cites Visual instruction tuning.

Efficient Multi-modal Large Language Models via Visual Token Grouping Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.577189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:25:00.350609Z digest=sha256:d10fbf87160732e065dfc6405b5ece7d9ff7f48ad22c2752655d6fdb64d0c302

Observation c3ade4e1-85ea-4eab-89ca-8399625fe3f9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Efficient Multi-modal Large Language Models via Visual Token Grouping MMBench: Is Your Multi-modal Model an All-around Player?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.352950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.352950Z digest=sha256:467363d276d51be8d828a013d607e1c0e562d408995faf371778337ad0e5c17f

Observation 57c8ee26-7b4e-4e2b-a9b3-98e9b32fed22 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Efficient Multi-modal Large Language Models via Visual Token Grouping Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.355823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.355823Z digest=sha256:815755c2916348afa4a0b073c99e1e3f53289b626ea93cb586ee93dccdeda6ba

Observation 405752d7-7327-4b50-a1a7-d897cf44dcd3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Efficient Multi-modal Large Language Models via Visual Token Grouping Learning transferable visual models from natural language supervi- sion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.358347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.358347Z digest=sha256:0ba04970cf820c43966922e3a1469659794fa10bdd71c5c9e29b75c7ab578ef9

Observation 5a07fea6-86ee-449d-ae0e-ece8976773c9 · outbound

This paper cites Towards vqa models that can read.

Efficient Multi-modal Large Language Models via Visual Token Grouping Towards vqa models that can read

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.360675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.360675Z digest=sha256:044eacde615d8136ff650250a2c1b6059cd2ac48f752d4a0fd98d0b94562c35d

Observation 13052108-b106-43fe-8134-5a5b9cf59b3f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Efficient Multi-modal Large Language Models via Visual Token Grouping Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.363060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.363060Z digest=sha256:53ee5a997bb5f62e4cbf0ce61c670bdf482341fbe91385cc53552f51dc3548f1

Observation 6956228f-437a-49b9-924a-2f03b47292b4 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Efficient Multi-modal Large Language Models via Visual Token Grouping Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.559249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:25:00.365716Z digest=sha256:8b7b0ace573aaae89197e377a4c090b0507ffb2945f22f65aa0d801a19ad54f8

Observation d9fce147-c50d-4345-a667-700ba991a541 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.368115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.368115Z digest=sha256:0395fad5f03b3722ec259bc52406b9ef046aa6d1ae8fdc7793dce098fd177c85

Observation 0dd63667-5baa-4b2c-a5ee-0bf1ad13ba94 · outbound

This paper cites Neural discrete representation learning.

Efficient Multi-modal Large Language Models via Visual Token Grouping Neural discrete representation learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.370673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.370673Z digest=sha256:bc7ea445ea76245b71b3c009d2358a1ae8c7e8b043a75707889891e3387edb0e

Observation 5357ac85-5fa8-4cb4-9c01-a3e34be52516 · outbound

This paper cites Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data.

Efficient Multi-modal Large Language Models via Visual Token Grouping Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.372991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.372991Z digest=sha256:df353d412d87481f5e0aa6c770475bb1c30239c795f3cbb1374b7d462271d91f

Observation 0b66a686-60c9-43f5-8695-8dfcd767f616 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Efficient Multi-modal Large Language Models via Visual Token Grouping Efficient Streaming Language Models with Attention Sinks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.375740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.375740Z digest=sha256:264c062e0063b2825bfdac3993c5e8440688bbdeb3fa3e3d8a33f40f6afbbab0

Observation 4df0f32f-6fbc-4d4e-b8ea-51feb0162fd8 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Efficient Multi-modal Large Language Models via Visual Token Grouping Groupvit: Semantic segmentation emerges from text supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:25:00.548185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:25:00.378156Z digest=sha256:a195c6390383e92e708a9ded463dcfc3561e65e4acca8afd72a050e397805914

Observation 5ecc5ec3-ec4b-4125-a631-81b5b5c457dd · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Efficient Multi-modal Large Language Models via Visual Token Grouping The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.380558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.380558Z digest=sha256:bb0cdc77272e8cca238d4ab3bc3696d6b789c61be74add20beddf4c98248af07

Observation ae0cb8b0-25e8-4e56-a9ff-a24c72415139 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.383749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.383749Z digest=sha256:ef055552deb46e384ceceeacc28577dfc40387a52b780e9b89491a4206c7f0b2

Observation 49af9acd-90f2-4ab2-8674-a7dcf84265aa · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.386323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.386323Z digest=sha256:c34b74fa20586ed467c48fef8ded989565b714950ef26f0cecf29a74f53e7cca

Observation 4e842cc9-a5c7-4467-b472-628f30265967 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Efficient Multi-modal Large Language Models via Visual Token Grouping TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.388877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.388877Z digest=sha256:d34c9a7f335df3687b10a0dbb59e3ccc1cbc63c8bddacb36743490dc512b946c

Observation 1a9a7bdd-8385-4453-a25e-79d65249bb9c · outbound

This paper cites Sigmoid loss for language image pre-training.

Efficient Multi-modal Large Language Models via Visual Token Grouping Sigmoid loss for language image pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.391401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.391401Z digest=sha256:2f22396bd2d85c84900960d4418ed9a8a120297874953f4e9e2dcfb7a1c38320

Observation b6806a50-cea7-4973-82db-3383cef06e19 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Efficient Multi-modal Large Language Models via Visual Token Grouping InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.393682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.393682Z digest=sha256:7ed8b022e283df98a9814e0fbe5e710359d58c3a2cbc74afab80401a38430c61

Observation 0471f26e-8674-48b7-bf4f-07ea6732e36c · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.396491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.396491Z digest=sha256:1fcb4e804a3cd9f5c5564fa9fb81add374e06ca5cbc2e081af071fdc85860692

Observation c17ded9e-09bf-4628-a2fd-3688b23fad5b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Multi-modal Large Language Models via Visual Token Grouping MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.399099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.399099Z digest=sha256:bee229b7ff37c1440f7ba8061594aa879f9b50266e7ff5972fab10745219af66

Pith citing papers

Observation 87aad0bf-2f97-415f-9013-4e57a68f7c86 · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models Efficient Multi-modal Large Language Models via Visual Token Grouping

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.028908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:5a42970d2784ce404434643238546479e1c29d3dee2400e16a71fbd86e5968ae