Pith. sign in

Paper Citation Record · LEDGER

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.12309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12309 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T16:53:11.708868Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:23:16.226474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact18
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8925a8f0-1eac-4acc-9206-b0aa14c0c37d · outbound

This paper cites GPT-4 Technical Report.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.355927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:3271f1621e4d6c5fe6fc0092a43700aca76ea6b64848115c28d9fe70655103d1

Observation 57835968-8203-488b-85b5-8fb4a9e81436 · outbound

This paper cites Improving image generation with better captions.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Improving image generation with better captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.063119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:3e6235e81db4491c6b4f730185eb56ea9a893061ef2033de46e3b5fb0b8b22d8

Observation 93c6eadc-433d-43c1-8739-fb5db997e5ad · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.067222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:46912b4f7f2dd275b74298a10498b8f3645f8e77139c8589d3939f41bfb0e41d

Observation 2fb7338c-1f67-4d00-93f5-a5513ab3418a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.361439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:5bc233a60b616a01af2abb25ffbbea4d5daadc4e734764ff85e1beec92202d53

Observation 409de80c-f4d7-43a7-a20a-96dd8d20815d · outbound

This paper cites Diffusion models in vision: A survey.TPAMI.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Diffusion models in vision: A survey.TPAMI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.065027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:9c4207628443a5c19d5b1b6b8b203c71e1c0216c417e4958f687aa31f3ac9878

Observation acf3a854-831c-4530-ad46-3314c1458c8f · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.069150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e38286d34f23f9b890d29a276edd02cb35fd6a09652e177249901ee4af826bc9

Observation 2edc038a-cb93-411e-b432-42c1babb6af3 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.358740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:83b69e70041ea537a6d5d8cf4e5d7b95a6b84d074deb0a1d2fe1f12a0169cc76

Observation fb4f0b5d-4128-4980-9b24-fdb4e39065a6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.091236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:1376a610ae5272dcd98fffccf00a91e055547891d87158e7e85c179b0bd757d9

Observation 92304e4c-3f92-4a58-89d6-eadbfdbb90d2 · outbound

This paper cites Gemini 3 pro image (nano banana pro).

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini 3 pro image (nano banana pro)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.089363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f3810cbf157d11d3e6b493562c6ea6129cb66b9c18765aa1a060c46bc71b382d

Observation 512906cd-a925-4aaa-a898-578662eee7c6 · outbound

This paper cites Understanding and harnessing sparsity in unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Understanding and harnessing sparsity in unified multimodal models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:720518788d4a91d076fe45a7fab5eb65d1d1a73579fc656069e72d1bfd006d09

Observation f7194ef7-c9ca-4ffb-a6d5-292cffebf5ef · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flux.https://github.com/black-forest-labs/flux

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.095103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:9952efd4b6240a39cdf6d1a75d87ea91444ff98d768975754811ad42d98b9c02

Observation 1f75a905-337c-41eb-b527-ae6874f6692a · outbound

This paper cites PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.367552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:4177052f38c6c409bad5bcb654e1718a75d853a7aff60ed3bf36a3ab0559aa57

Observation a51cdea8-1a5b-4a2b-a019-274e5fd45dca · outbound

This paper cites Dual diffusion for unified image generation and understanding.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Dual diffusion for unified image generation and understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.105285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e879df984907eb2d5373ae8a741eb03accd4b3eaed2244dd2e8ea2b9c5aa3c11

Observation b7f3a2a2-d2e0-43da-9e26-f4b372086679 · outbound

This paper cites Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.400396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:cc28ac2a311deb45e1125a4c633b8d7214dac2addf39ac062e93649459851ca2

Observation c2acda8f-bb3d-4509-bc3f-7f23171c38f5 · outbound

This paper cites Visual instruction tuning.NeurIPS.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Visual instruction tuning.NeurIPS

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.097712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:6d21b09f36af8da8d255d6dec7b1a6d8dc36a23467a70ef527e350426972c3b5

Observation 702ed9fb-c920-4ae1-8f1a-07f44c4fd09d · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Step1X-Edit: A Practical Framework for General Image Editing

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.379475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:36aace0da68c271aa4308776303f25da8b4b9ac485136ba494e28560674e38d9

Observation 396631f7-befa-4636-88a1-800ac4718700 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mmbench: Is your multi-modal model an all-around player? InECCV

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.109278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:b2f102e5ac5fd4c2b73f0bb6e89ec261ff22a45764a7ca1250327e4cd1c1cc4b

Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.364921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:8caa947bd977bc676ba8b9d2d29ad70ea4fc5cab2081bce1468f465f005e6d77

Observation b6f3980e-1133-4e7e-b1c0-a35a00f4643e · outbound

This paper cites Introducing our latest image generation model in the api.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Introducing our latest image generation model in the api

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.081011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:8492a1cfc6b50b8e3407692c76de1705478a8e511929accf7b70a4d743da7421

Observation c7dafb59-dbd9-49c7-8c81-e274df65986c · outbound

This paper cites Wiseedit: Benchmarking cognition-and creativity-informed image editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Wiseedit: Benchmarking cognition-and creativity-informed image editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.370572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:65c61b7ed07293b97549ba2752ab1441d5e1d282bc7d5e9e9bfe564845c55179

Observation c257be5a-160e-4068-9868-a496e09957f2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.073244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:5cce6a9113e1a3e7d029658067fa469b5b4c5f4c848c59c8f9bfaad92c4c80c9

Observation d9698088-0486-4840-9bb6-dfd79191744f · outbound

This paper cites Holitom: Holistic token merging for fast video large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Holitom: Holistic token merging for fast video large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.075719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:d4788465fcb0ad550de7e3b5dc3fcaf1ea2adf545383fdc30777f605b06113ea

Observation 5279849e-2857-4484-85a1-a1549090b85c · outbound

This paper cites A survey of token compression for efficient multimodal large language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models A survey of token compression for efficient multimodal large language models.TMLR

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.071140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:1db637ca164f05156b9be2ac3893672096c03b8048e7ba2efe002d3c68edb0d0

Observation 6648d7ce-bd9b-43ba-a4a2-2f24da9d50cb · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.079133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:472d11c0b765c0e927f82ed04013edb14c48e49bf409873c9b8f1dd28a3ddd70

Observation 65fe70a0-d0bf-4b63-8b6e-dc285978a3ab · outbound

This paper cites Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.376802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:bb52fdaeaa91dbf251877dc9ddfe2068b9d297996870e069f7934c3810e78400

Observation 90089d6c-0c9a-4f69-a0a6-6688a39f8eb4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.394111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:2891a0d054269129ae089fbb04fc255d66cb7bc42f58136fec62006d868b3c4e

Observation e8fc5633-70f0-465c-9a87-538748b6a09c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.391340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f53c907081a3a313f8b06cb7d8e77e7dfd5eb39585d4a1b1ecf66ded8a8f2ed2

Observation f4b54679-5e4b-423c-88e0-ec99631350fe · outbound

This paper cites Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.385307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:97ed27f196e88578eb0ea990dc538b32f0bb199cc5300d82365ce061d44efc8e

Observation 8b167d6f-4d5c-46f9-9db4-6bdd10dedf49 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:225beab0f161c56722074ea760f6f70d73b48b1c69fed29cde7d38cc2428873a

Observation ce83964d-a53f-4935-bb8e-7df1de413c51 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.382343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:a05e7176d6d3b42c8fc6f1111075f536212b47c6400b974db9a9d3f6b8572142

Observation 4b4c5688-3ebe-4e3c-a403-ffca9069dd12 · outbound

This paper cites RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.403042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:12798104bc873f42b6712c109c9f00187bcf65f7a7f2d7d2c9963797d8c82561

Observation 9dfbf07c-325d-4a9d-88ea-608045d23a5c · outbound

This paper cites Emergent hierarchical reasoning in llms through reinforcement learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emergent hierarchical reasoning in llms through reinforcement learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.406050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:264d8eec10f6f6ba0be13006c16c73fd1c1d64289453cd3632c855d5303a57e3

Observation 263065b5-48c9-4c98-83be-91bec573a427 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.388168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:04696a7264488a26bc9667817b6dc0e058ad7c6107691799c5d633b90ca0c204

Observation 1b789c3a-4809-4c74-8ac0-7956950501e0 · outbound

This paper cites Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.107312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c236601852a1cc2fbcd2c9846e123e8cf838cce8f0488f1490f5b8db4160a740

Observation 8bbc5f82-5fe6-497d-9c08-99d64483cfda · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.111113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:aacfdea73e74b5edbf060f9a6d172cadce55d6d6376501691bac172847a839f5

Observation 68a2200a-c987-48c0-ad26-275c56e6a2bc · outbound

This paper cites Kris-bench: Benchmarking next-level intelligent image editing models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Kris-bench: Benchmarking next-level intelligent image editing models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.103247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:926d6674a022f8dca9cf9dcbef93d580214f1fcd0ca5e488dbd095258df7818b

Observation 06362bac-ed3e-497d-be0b-3e558285b072 · outbound

This paper cites Announcing grok-1.5.https://x.ai/news/grok-1.5.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Announcing grok-1.5.https://x.ai/news/grok-1.5

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.061118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:fe81bfe459713d8a9a8971cd9b2adb174244657ffa9aa3d4dd34b381541622e1

Observation 0e9686c9-6da3-40ea-950f-e098201c1668 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.093044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:6a2ae959db0792e696c790b661fc32431f21f881be24536e82e20b8ac8f22a8f

Observation 653e866e-61f3-45bd-8d94-dff9c366794f · outbound

This paper cites Show-o2: Improved native unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o2: Improved native unified multimodal models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.099570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:caec15b015bb1c99034640355ccf4d957a105d6eab86554b6b57a266e9355f51

Observation ddfbf650-9dd1-45aa-a626-43c53268e365 · outbound

This paper cites Conical visual concentration for efficient large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Conical visual concentration for efficient large vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.101372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e390a2480690472e6adc697724df05274959a92805591292313ddf087066f080

Observation 369b972d-879d-4699-b2f2-69434741e0c2 · outbound

This paper cites Rethinking visual token reduction in lvlms under cross-modal misalignment.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rethinking visual token reduction in lvlms under cross-modal misalignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.112952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:5f48631a7041a8c6da92ca9525a3ba8fc212739a090d59e7a889c57a4c35171b

Observation 333c58d3-8dac-461e-8621-c8f93a8cf2db · outbound

This paper cites Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.083000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:7cd325291e13c213553b9eaba4e5e9634261cf1ad1433463f05b9f6ee5e3181c

Observation c0103691-5230-440f-a923-24eebb1f3db9 · outbound

This paper cites Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.087497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:12383122fcd0c87800e02697d8bb42f5181d214eb79ad18e707367538649f690

Observation f49e966b-04a2-403c-9ca8-e08a3a2ac678 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.373501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f8d3bf5603160bfbed76ab86319ad8989429a46e2926530feeca164329b38f42

Pith citing papers

Observation 66fef0c4-b662-4749-8b0b-bb26e25f41b4 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:41:47.727930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:41:47.727930Z digest=sha256:efe163e5c9269c5ea3104c1a21489903cc1b3857dfbce7a1ef1a75c239b92d48

Observation cb1dcba1-9b29-49f2-8e36-bda0ef0f3212 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:17:41.520667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:17:41.520667Z digest=sha256:a20ed6e4a2a661c2b581fe8854f44c004736b0f5e2a520b9cc591fb2690d71ec

Observation f6b5f369-d12e-4297-b3e7-68f9c8070272 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.226474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.226474Z digest=sha256:b2a7cfa2baa85af03be24646565d1dfd418c431dc010a54c1810c8a4aeb6668f