Pith. sign in

Paper Citation Record · LEDGER

On the Limits of Token Reduction for Efficient Unified Vision Language Training

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2606.01503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01503 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:57:56.333595Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbfd7d18-a573-4d64-a687-e89dd65c2022 · outbound

This paper cites Token merging: Your vit but faster, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Token merging: Your vit but faster, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:c8f02f88123ce6b6e0f6858cbb529726c459f5cdbb143106e358d3f337756126

Observation 5182b6d0-8fae-42f7-8347-08ec0a408930 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Sharegpt4v: Improving large multi-modal models with better captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:56a9647e260e9baec143dd4ed159755498e3fa28a90889c4608d189e7fe3258a

Observation faa7894c-0022-48b4-a455-4aed482fb906 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:5cb4261fdf0daa195b73c45c429aa882641c7b8788af230f28838f9bc4e53fce

Observation 770372eb-6b43-4e78-8d0e-72b8116ef617 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

On the Limits of Token Reduction for Efficient Unified Vision Language Training InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:a0167ba0312d76cf758ea2110cb18dde627bcf78a98e0fdaeff7db444a5dda48

Observation ea0ef4b0-f621-4a27-99ac-9cade3a0fa49 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

On the Limits of Token Reduction for Efficient Unified Vision Language Training The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:d043b18caf5a47b51ae439945357d9da2f81f366de2e5b9e0fe3a8eabb902cf7

Observation 58521017-b486-4c5d-9016-065457d67803 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Taming transformers for high-resolution image synthesis, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:48171ad5974604570475c974b16fd8602fdda4c4eed2d5dc65caa9f132fcc8cd

Observation 44ff07d6-1c22-4cf0-9af8-e9d4234b7a7a · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ea27a41cf616eb9d0b19dd98d57e6e7a132855f92e0a68f42de16e5d0d96e7fe

Observation 38cf4303-dc7f-441a-8193-7b90a452f987 · outbound

This paper cites MME: A comprehensive evaluation benchmark for multimodal large language models.

On the Limits of Token Reduction for Efficient Unified Vision Language Training MME: A comprehensive evaluation benchmark for multimodal large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:cc06c2a55e018f9c043755f5137f702e59bdfd4fcbf78eb42ef082d9d23ebe2c

Observation 4db44839-f96d-4531-8039-1693ce4b100d · outbound

This paper cites Making llama see and draw with seed tokenizer, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Making llama see and draw with seed tokenizer, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:061f8f36a11ece272967fce3eba6bf122c3cbc30ef994ceed82cde580672fbac

Observation a8bcef66-4e17-4cc4-a754-1eb5699ce92a · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

On the Limits of Token Reduction for Efficient Unified Vision Language Training When Attention Sink Emerges in Language Models: An Empirical View

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.305996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:871758bd09aa492e54214ce970a25d94fd6a86ade82d77b0701606eab5d457ad

Observation f4be08c8-aa4a-428b-8a45-1e8138c42d1a · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Denoising diffu- sion probabilistic models, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:7e878faad7acb7a67913823b4a31aa7e24efda94f124adcd58dafeae6b6c39e3

Observation 157e9bb7-b29d-4280-af17-4e1a4c143bc3 · outbound

This paper cites Matryoshka query trans- former for large vision-language models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Matryoshka query trans- former for large vision-language models, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:4de386d73630f7faea976be8c65f134d74be1f697514307851c797de7549c1f1

Observation ed59ce8f-b34b-40a4-8569-5e2f1db43333 · outbound

This paper cites Hudson and Christopher D.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Hudson and Christopher D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:00778251d50a880c734227e6ca8ac6fda75b408da20c5592e755f870d89ba502

Observation afdfbab7-b92a-409a-9f15-ace553f9f1a4 · outbound

This paper cites Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:46830abafc3c2cb45f8c77aadfa8cd7a03cb238d3e76323bee63a67b2eee14c6

Observation 42b01c53-b1aa-48ea-879d-453a84b3f40c · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:4f1bf63632577d7871fde5e4a2087334dd7c79b51da0a130e3554bf1c7f3ba43

Observation 89ad5c31-298a-4fe0-8615-788fbc0fa72e · outbound

This paper cites Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ee77e0e32ab6e6d77d08d7177c6b0858eded34f09959d9cf890c157649c47088

Observation 98055ec9-bb71-4cd4-b80f-211c61d6739f · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Tokenpacker: Efficient visual projector for multimodal llm, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ee92b1306c2e7cf3f1548c307412f6a9271074fe9be5421398acfd764c90f97f

Observation ff3921b5-542c-46d0-9329-ce3915b223b9 · outbound

This paper cites Evaluating object hallucination in large vision-language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Evaluating object hallucination in large vision-language models, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:82cfd290350f05ba0ae87f6f9eb92e8a4a2d0b24850bf92e0a8e27a338f326f9

Observation 57a49fa9-51e9-4bd3-beb3-457280b55024 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama-vid: An image is worth 2 tokens in large language models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:4c375756804c54add01860919084352fc022a773dcafa5e41bd20c5d2603317c

Observation d5eef721-4e82-4b5d-8470-f6b1b6d05811 · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Boosting multimodal large language models with visual to- kens withdrawal for rapid inference, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:80d4db9cfb8aeb242d0cf2057ea25e94f771936d18b9efa21bf335d912eb635f

Observation fb79f693-dbb4-4cf6-ab19-5298c61918be · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Improved baselines with visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:c644db7a861bbd09ec30f8e58699125440bdc1ecdaa2b73d07b36441c1b68096

Observation 6c014e12-a6f6-4551-b5aa-80c8205add77 · outbound

This paper cites Visual instruction tuning.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:40da8e9782c905df1f175f6b7bd97b97d4193b78d014fac73e15cb3904e9effb

Observation 088db68e-ca28-4070-b332-7c9a29264fb8 · outbound

This paper cites World model on million-length video and language with blockwise ringattention.

On the Limits of Token Reduction for Efficient Unified Vision Language Training World model on million-length video and language with blockwise ringattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:617229c3175393967e9e886eb62a10321dd44908b665eb8b5798f3666e5f76db

Observation c979e0a2-1a8d-47ba-90bc-5afa477b54bd · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Cheap and quick: Efficient vision- language instruction tuning for large language models, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:25cfc989575e7303a40fe873d678f8dfb2d1e0efb785cff4da706feefa523980

Observation a8922ebe-c7cf-4373-a6e2-f50dcad78f6e · outbound

This paper cites Unitok: A uni- fied tokenizer for visual generation and understanding, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Unitok: A uni- fied tokenizer for visual generation and understanding, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:f67f62122ada8f99375859479b264db5be901e676442668b4b51c84929d76e53

Observation bbe35b9d-02d7-4e02-a0e8-9838de9541ca · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:01a463a9d38a406e42a9bd13d050405d3da45571bc561b99029472cb818ce90d

Observation ce0a3167-fc2c-44e5-aeae-5cdc2e55d129 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Learning transferable visual models from natural language supervision, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:f4a2b3659fe107136168e5dc15245ba37476a6d76a2fb9649df59c7b648bc937

Observation 591136e9-8ab0-4bb3-aecd-f36072565446 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:3de28d54de3d7b064128b1fccbd8adf52c9b7cca395b33268fcb695555e3c1a1

Observation 785be7f9-7aa4-4659-93e8-89dcb2eaaf4a · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

On the Limits of Token Reduction for Efficient Unified Vision Language Training High-resolution image syn- thesis with latent diffusion models, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:a4247b9f6dd0710d5a3b88f729847fb78ea88c6094c76ee75a8cab6f93d78629

Observation e7be0623-05e1-4c85-9800-da718381a4ed · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.309040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:f28c27f7aeeefd287b7a19e2e3bb66bb115adccaf5620bd8bdce4a8651ab82fb

Observation 37ab29a8-aa8c-4d0a-9d2b-2eea94331399 · outbound

This paper cites Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:907824bca5bc0dba0c56c159c7f616fb5e57ad6fdc2ee7f30f07bf19e5f0c88d

Observation fddad9bf-4462-490c-a628-42401bd01eef · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:508d533004cd25e387ea70f16f9ee171f74eecfb9ff5b77030834afab159c1ec

Observation 5698df56-b8a2-4c2f-8be6-58ea4f5cf83e · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:4771332071ef7ad11dc0ce8763ea9b64053608e872e457c1e889a3ca37f48f3c

Observation 066be97a-8f72-4f14-bf20-9d31fc2136d2 · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llama: Open and efficient foundation lan- guage models, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:b5258da0bacc86a0cb901ffe6e504b20a3eebd9887e3ef7852253b2b529ba6a4

Observation 98f9b6ff-325a-46dc-84d5-ef746b640c5d · outbound

This paper cites Neural discrete representation learning,.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Neural discrete representation learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:1d37773069405b91e6d8438fbb94a9c8c6cc972d293424a86d543482e414bc0a

Observation 0f816c86-b247-4de3-8ea8-20d3a015286c · outbound

This paper cites Emu3: Next-token prediction is all you need, 2024.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Emu3: Next-token prediction is all you need, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:0462a03156b861bf10a2990767c23199565e9afd03f676384f819fd37e8bf921

Observation 2ab1316d-da58-49bc-8104-3f5b5c10c13e · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.303829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:41e4d92695f2c1723403a73e6d9eb64bfdc541c2239a3c7b04eb2b13e1b8e366

Observation 510d55b9-c96c-44f8-8248-8422ac28ebf1 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.295561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:8ab4558fa5ce626f1b017588533aac4eccc0282cd6f2cc6e026510517ffb2ca6

Observation 6761f90c-f3e8-45aa-9e7c-f7abb418ae2c · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.298245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:b5d4cac3b9e089d2ac82730f65cbba4e9def107f60fd1f7f8886f5d9bac3c43f

Observation af1e59da-4e95-4b36-a7f6-51e418c045df · outbound

This paper cites Efficient streaming language models with attention sinks.arXiv, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Efficient streaming language models with attention sinks.arXiv, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:1ed28a9f1d0a7a55b31bb3757df1f42833bf785aa35d1b4fd99676b324376a51

Observation 85ede03a-5ba7-4186-8ab4-3cffc56e2544 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:02:24.300518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:8e7faf47bafea9a58ee83035e384c10fb2f52455f00a196300354968f956208c

Observation 248ee663-a3b4-4695-83a8-cd9e1cb45e9e · outbound

This paper cites Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:7ed71046fcbbe13f9d58c15f88b74330308d615f697aafd0a7e1c1015c5ef774

Observation 5bf6ecac-c1a9-4436-afa4-14e1ea516849 · outbound

This paper cites Anygpt: Unified multimodal llm with discrete sequence modeling, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Anygpt: Unified multimodal llm with discrete sequence modeling, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:ea64026e2911a7e0c48e4ebac9680afadc6a7953a77c3e9d52e7cb411854aa5d

Observation 3fcab27b-cfc4-4226-84c3-cffd89c4de0b · outbound

This paper cites A-vl: Adaptive attention for large vision- language models, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training A-vl: Adaptive attention for large vision- language models, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:2e0679aebd7892676e332340f79a95788874c3884db9a231e25a5bc96fefe73c

Observation 3202d961-ff2e-4165-b8cb-f45f634bbcd5 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Llava-mini: Efficient image and video large multimodal models with one vision token, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:8c87941205f0f455b65b5d14a25ac6f9c725548e5647a89390c5f97a4f4229d4

Observation 24525677-03b9-475a-a718-cc04d98432cd · outbound

This paper cites Himix: Reducing computational com- plexity in large vision-language models, 2025.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Himix: Reducing computational com- plexity in large vision-language models, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:3cb5fe482b5e6400db15be1940a42af1db5f345749e17f5146c23b7ffb5e6e7b

Observation ce075547-819f-487b-a606-2e0bc49fa037 · outbound

This paper cites Ar- gus: A compact and versatile foundation model for vision.

On the Limits of Token Reduction for Efficient Unified Vision Language Training Ar- gus: A compact and versatile foundation model for vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T16:57:56.333595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T16:57:56.333595Z digest=sha256:31a9e2b152c44ece39c752a6ba311e3da1acb096483f7383f1dd339635de8d3b

Pith citing papers

No inbound Pith citation observations are available.