Pith. sign in

Paper Citation Record · LEDGER

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

As of 12 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.12309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12309 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T16:53:11.708868Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:23:16.226474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact18
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8925a8f0-1eac-4acc-9206-b0aa14c0c37d · outbound

This paper cites GPT-4 Technical Report.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.355927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:b9b0ff9ca309948903444fd6a3e0133aa437273a885d6ac36b7fc00825a1e406

Observation 57835968-8203-488b-85b5-8fb4a9e81436 · outbound

This paper cites Improving image generation with better captions.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Improving image generation with better captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.063119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:64327164a6da94c31f83f3a2b25803aedc40324156a72f22ff030b4b90f8aab6

Observation 93c6eadc-433d-43c1-8739-fb5db997e5ad · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.067222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c561a493c3086bf0bfa0b59e5bd4908eb1af774cc0c7d1c424c60ba243cb6c31

Observation 2fb7338c-1f67-4d00-93f5-a5513ab3418a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.361439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:8316f6ac518148c4cbbb5db9cbdb3ad2f8335c452af2f114682655371c2f35db

Observation 409de80c-f4d7-43a7-a20a-96dd8d20815d · outbound

This paper cites Diffusion models in vision: A survey.TPAMI.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Diffusion models in vision: A survey.TPAMI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.065027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:ad33095eb4597e9c91e90462bb9528359d99b9c481f7a78b1e6a3f60ebc66cc8

Observation acf3a854-831c-4530-ad46-3314c1458c8f · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.069150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:cd3d57ff0a3061784e6630e6e9b80db87ab219c8916d318f3b5e608f20f6dcd9

Observation 2edc038a-cb93-411e-b432-42c1babb6af3 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.358740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:4508efcc80b1428b01954781e462073b5ea0be1c6313210d1ff6a8b53b20cc89

Observation fb4f0b5d-4128-4980-9b24-fdb4e39065a6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.091236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:05b643d31c61ca96d2aab69b686c072d078705e87454942ed63dcab453d1a8d5

Observation 92304e4c-3f92-4a58-89d6-eadbfdbb90d2 · outbound

This paper cites Gemini 3 pro image (nano banana pro).

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini 3 pro image (nano banana pro)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.089363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:81075245dafe43fc1c5227ec9522a3c2efbca195862ef8fe012ac6afec793737

Observation 512906cd-a925-4aaa-a898-578662eee7c6 · outbound

This paper cites Understanding and harnessing sparsity in unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Understanding and harnessing sparsity in unified multimodal models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:b0b0313868fdb6b36db04c05a862dc5fda50a335e573cfdcdd3bfd2d61b8e01d

Observation f7194ef7-c9ca-4ffb-a6d5-292cffebf5ef · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flux.https://github.com/black-forest-labs/flux

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.095103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:11b430e981e112879f0529feb15b922e722fd783f909b46c40c0dbcdb2c94481

Observation 1f75a905-337c-41eb-b527-ae6874f6692a · outbound

This paper cites PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.367552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:1ea415b73d5d3201448d35ec9bc9187d98bd812a4be27825c91e958a69a53808

Observation a51cdea8-1a5b-4a2b-a019-274e5fd45dca · outbound

This paper cites Dual diffusion for unified image generation and understanding.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Dual diffusion for unified image generation and understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.105285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:701514aee1317cd998cc2a5221f0f36cfccc4923ee12101e4c849253846848b3

Observation b7f3a2a2-d2e0-43da-9e26-f4b372086679 · outbound

This paper cites Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.400396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:d8299b2804576bcf3e2391d9dc3028eaa002bc39908f1dc671794a2fd63cd2a6

Observation c2acda8f-bb3d-4509-bc3f-7f23171c38f5 · outbound

This paper cites Visual instruction tuning.NeurIPS.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Visual instruction tuning.NeurIPS

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.097712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:62e8c6f1081fab2266203d6e5dc70908f76750071082c45f319416c3778d02c3

Observation 702ed9fb-c920-4ae1-8f1a-07f44c4fd09d · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Step1X-Edit: A Practical Framework for General Image Editing

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.379475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:533a0c6a521fb79874a7f3bf2fbb4514e73d8b30b355a6cef81a819481ec677b

Observation 396631f7-befa-4636-88a1-800ac4718700 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mmbench: Is your multi-modal model an all-around player? InECCV

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.109278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:a7581bcd3f62c067ed23feb440d08b19856e601b773abe233a6e86dc1631cbe0

Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.364921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:69dcc05f9757a9eddc6ae928c0ff1e67e7b1ae1504b120046adac0212d3cee67

Observation b6f3980e-1133-4e7e-b1c0-a35a00f4643e · outbound

This paper cites Introducing our latest image generation model in the api.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Introducing our latest image generation model in the api

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.081011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:675028a79225e62158edc3c83154ee186ccc3fe621160c21b1c2c50c18dfcadb

Observation c7dafb59-dbd9-49c7-8c81-e274df65986c · outbound

This paper cites Wiseedit: Benchmarking cognition-and creativity-informed image editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Wiseedit: Benchmarking cognition-and creativity-informed image editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.370572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:d7bfa1fff7230730d5171bf099751ba2be68088ad30960aef3500ee0f039980c

Observation c257be5a-160e-4068-9868-a496e09957f2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.073244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:2e2e48efc86b5641d4bd084f989d4013f176283082b68b2ab787ebd44c4a819f

Observation d9698088-0486-4840-9bb6-dfd79191744f · outbound

This paper cites Holitom: Holistic token merging for fast video large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Holitom: Holistic token merging for fast video large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.075719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:12b75a79b3f004fd46ab8a225fdd7397503fc6641fc0815ce3d92d58dfad36cb

Observation 5279849e-2857-4484-85a1-a1549090b85c · outbound

This paper cites A survey of token compression for efficient multimodal large language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models A survey of token compression for efficient multimodal large language models.TMLR

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.071140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c8f764f08582dfa1fc78ab3d646918c2b3b6fff7e03d8855bb6dbdc9f25302f1

Observation 6648d7ce-bd9b-43ba-a4a2-2f24da9d50cb · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.079133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:96b960d0655204714dfdbbf289eae6c09a8bab910307812f2a949ab6b96a1d7d

Observation 65fe70a0-d0bf-4b63-8b6e-dc285978a3ab · outbound

This paper cites Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.376802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:2da64eafeb8263d5a7ef8a5bb67756c4f57a27431d27a85728716a62e7ceff35

Observation 90089d6c-0c9a-4f69-a0a6-6688a39f8eb4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.394111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:99f0fb36bbe571bb4b5cee7cd10c3195c71cd2cde0aa07297890a843deb60abc

Observation e8fc5633-70f0-465c-9a87-538748b6a09c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.391340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:ce2272854a469f4925ab43dc6f4536e0296bc201e5045eae21e3a654d621b079

Observation f4b54679-5e4b-423c-88e0-ec99631350fe · outbound

This paper cites Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.385307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:66a5a99b5b2c4bef0c731bac1e6986b42c6cc2fcf3579fe1c94f499619ab6408

Observation 8b167d6f-4d5c-46f9-9db4-6bdd10dedf49 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:20974387ee25210a8a735c2b204ce7aa442bee05dd81d16fca67a8153670163f

Observation ce83964d-a53f-4935-bb8e-7df1de413c51 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.382343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e20889ea59b162f4430e8683bcba3a9dadce2e982f3f9ee43b33b6965286af60

Observation 4b4c5688-3ebe-4e3c-a403-ffca9069dd12 · outbound

This paper cites RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.403042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:8268c3f47f07422e3bb6d6d40ccc0d181e1c316c523f9a4015d36e9e48acef42

Observation 9dfbf07c-325d-4a9d-88ea-608045d23a5c · outbound

This paper cites Emergent hierarchical reasoning in llms through reinforcement learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emergent hierarchical reasoning in llms through reinforcement learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.406050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:43ba6784e7e96439676cf53129bb88507d39f046091c5aabd4cc42dda76e9264

Observation 263065b5-48c9-4c98-83be-91bec573a427 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.388168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:3fbdc40d4161b63a7cbd35c0e9159eca90fec263d875582981a7fd4a5e2fffa9

Observation 1b789c3a-4809-4c74-8ac0-7956950501e0 · outbound

This paper cites Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.107312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:fdb9abfe21453f0ff1bc196dfbf1273aad0f58d2c4aa1d4b9d77689b5041f961

Observation 8bbc5f82-5fe6-497d-9c08-99d64483cfda · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.111113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:5cde1244aaf1d735c88d3f167322510a923b85f27c8b133737d7d48ded5c0c3d

Observation 68a2200a-c987-48c0-ad26-275c56e6a2bc · outbound

This paper cites Kris-bench: Benchmarking next-level intelligent image editing models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Kris-bench: Benchmarking next-level intelligent image editing models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.103247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:a93a04c26bf06463cedf3fe53a9aa940b65531f84aa54c1bc874d6b16ce2bfbd

Observation 06362bac-ed3e-497d-be0b-3e558285b072 · outbound

This paper cites Announcing grok-1.5.https://x.ai/news/grok-1.5.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Announcing grok-1.5.https://x.ai/news/grok-1.5

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.061118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:143e1f4ac09141d3259ffbdda56dd7f8dc6805a6cd4f4bf179d2aafdfba0f981

Observation 0e9686c9-6da3-40ea-950f-e098201c1668 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.093044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:91ead56651387437d32c0db5bdb618e30f2dac28f6fa7b518cfa753f4e027d3d

Observation 653e866e-61f3-45bd-8d94-dff9c366794f · outbound

This paper cites Show-o2: Improved native unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o2: Improved native unified multimodal models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.099570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:9a6f306e8dea3d9ed84262344669a3a650b8ddbd594e8f35cd510f75d58f52e5

Observation ddfbf650-9dd1-45aa-a626-43c53268e365 · outbound

This paper cites Conical visual concentration for efficient large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Conical visual concentration for efficient large vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.101372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:fa4c0cb519ef9b0b7afb5b0305f213c358b56dbe58229fbc5577273e1f4db98d

Observation 369b972d-879d-4699-b2f2-69434741e0c2 · outbound

This paper cites Rethinking visual token reduction in lvlms under cross-modal misalignment.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rethinking visual token reduction in lvlms under cross-modal misalignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.112952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:4ec09c0f6c760f4bcc83185728c417f2b439113cbd735e5252303f5346e87ae0

Observation 333c58d3-8dac-461e-8621-c8f93a8cf2db · outbound

This paper cites Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.083000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:84a9e596ae5cc62790ff5ed765f33f210d5496ba4315073141e470426aabb6d7

Observation c0103691-5230-440f-a923-24eebb1f3db9 · outbound

This paper cites Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.087497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:0507e90169b167c9a7c390fe5116dc9e614232ece43bb0782dcdaa89e6d0d688

Observation f49e966b-04a2-403c-9ca8-e08a3a2ac678 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.373501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:cb14a8509dbd7390675d2a78542a676872cf1692c5e4f813436994b0af848dee

Pith citing papers

Observation 66fef0c4-b662-4749-8b0b-bb26e25f41b4 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:41:47.727930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:41:47.727930Z digest=sha256:efe163e5c9269c5ea3104c1a21489903cc1b3857dfbce7a1ef1a75c239b92d48

Observation cb1dcba1-9b29-49f2-8e36-bda0ef0f3212 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:17:41.520667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:17:41.520667Z digest=sha256:a20ed6e4a2a661c2b581fe8854f44c004736b0f5e2a520b9cc591fb2690d71ec

Observation f6b5f369-d12e-4297-b3e7-68f9c8070272 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.226474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.226474Z digest=sha256:b2ff56c0d70afdc17425da90f0d90255e9074f6d4ccf7c0ffa6d8076c9b3ae34