Pith. sign in

Paper Citation Record · LEDGER

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

As of 19 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2501.16297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16297 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:38:18.919283Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T23:09:32.594194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:01.665133Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cecb0cd-669b-499c-b5d1-611b54e4602a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.660902Z digest=sha256:9c44a4067d1c3333f53416a7f0f75a86e521c6ac7216db6f913d97b9f6a1c2e7

Observation 2930fe71-38c5-4dd1-96c0-cf4eb91f3fa8 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.650731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.665980Z digest=sha256:8566aa6d79b86cda10cdc6043d88882ddafc3dbcd2651f2f3014fc45723e6e3a

Observation 00b8b4a7-9f6e-4a5b-9383-1862af359e04 · outbound

This paper cites Less is more: Empowering gui agent with context- aware simplification.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Less is more: Empowering gui agent with context- aware simplification

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.639220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.670578Z digest=sha256:258e8c5608901c3dcbf840d49c65152185ac76f5fa1334b0f0939ff8c0f2336c

Observation 06f4e848-242b-4446-baae-33a9dde9139b · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.674964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.674964Z digest=sha256:58ed138c343d45ad8ef37a7cedf61d959ff8a43a21b9bc8cd1c5e3615fa9cd6c

Observation 48e9b2e3-80b4-4576-bacf-75157864c99b · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spa-bench: A comprehensive benchmark for smartphone agent evalua- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.628063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.679118Z digest=sha256:a94582f6aabf634217e76fe70c40dc17edbe4db39164fcbcfd7d2e27e8db5ac0

Observation 4a044303-7ddb-4633-a91d-549cd3158b58 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.683251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.683251Z digest=sha256:2aeb2bf1441a75a4fa036d418f7b2a34e7663943d197c428eb00965c77c5e328

Observation 189dacbe-1e7a-47f2-b301-7d58e427a451 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.610545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.687230Z digest=sha256:55a5734dd3e0982273cb94cf219e42865cb57b2519cd76d10bdc6b8136419acf

Observation f4d8f2bb-447e-48e7-b1a4-6d90bfe78e25 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.692030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.692030Z digest=sha256:bb19e996875dd124b3dae925f10909cd250761dccd43dd9fce9227f52098e920

Observation 86a2d164-30f3-4be8-878c-2ed042bca62e · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.592111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.696051Z digest=sha256:9c48c0ec5a20b23b22de3ecb292ba3a9442748d58238a78df2259007f804f8a4

Observation 54b7c54b-407f-42a3-a4d7-cae40c478b81 · outbound

This paper cites Vision transformers need registers.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.580813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.699899Z digest=sha256:9a064fddf75d73e07ea1b64f4753f6c14bc908ce33703d2bcc75c0ec5a6e6099

Observation b05491e6-12f0-4fae-b022-4e138d0b0fde · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.703723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.703723Z digest=sha256:7ef22e3f1b7b08e2e6af5aeb5f86b69a3c50f582251fa8f8f37a48bdb461dc99

Observation 86885e08-9b0e-4839-b105-0bd5e9780b93 · outbound

This paper cites The Llama 3 Herd of Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.707825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.707825Z digest=sha256:f26506d2a6c9be4d273d2173535f8252128ad8d980846c7eb2b1e7db9b7c41d6

Observation f0b27731-56bc-4aa2-b96c-c85fdb6faf2f · outbound

This paper cites Gaussian Error Linear Units (GELUs).

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gaussian Error Linear Units (GELUs)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.711903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.711903Z digest=sha256:b98f2886ec0ae4b3394b8cce52c040f54d9bd45e7b2a0bcaf97ca2c543a70f10

Observation ac01ab06-dbdb-4897-a3e3-12b933609d0d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.716137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.716137Z digest=sha256:0d22b2fb2a71ff719aa5fb43a1a0a03230e38dec1d8d37b4462c103b256f5d12

Observation 33395e58-9e1a-40ff-842f-eee1d6082bc2 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.563433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.720321Z digest=sha256:34df8f8f44e50c84539cc6741c1fab8c86def12e4e6ea4184364f4937f2f8131

Observation 905014d2-2ac6-4584-bced-f09eaa7de993 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.551976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.724048Z digest=sha256:e867f4a67af9a2506da238ea472b520669e7fa595037e122924be411259dd1ba

Observation 156e88a4-8f94-4638-83c2-989d28bd2371 · outbound

This paper cites Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.541358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.727665Z digest=sha256:07599500a4ef88d49d63d8a11da043139639315a48fd2bda76f5382acde3b2d7

Observation 6493d439-31b1-4e22-815a-5bd29d58c30c · outbound

This paper cites Openvla: An open-source vision-language-action model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Openvla: An open-source vision-language-action model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.530445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.731361Z digest=sha256:3b8547b3ebc6d870dec5711e739be64f3adb2333ed8242c06a50449bbc617092

Observation 057861c3-7ed6-4acb-943c-0fc7d84e2852 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.735094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.735094Z digest=sha256:9fd1dd138c4fb540efd029d5287eb6567c58e78d74827ac65fb5ad85cc59fe26

Observation 01559c1a-9d54-4159-b24d-a22c4e2cb7e2 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaV A-onevision: Easy visual task transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.739410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.739410Z digest=sha256:3e1e3a3a9348d2f59c2bcf71b8d11b1e6aee6aa8a0a662f53103ad920eb37084

Observation dbaafdc9-d0bb-4a5c-8a55-5eb8ee60a24b · outbound

This paper cites STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers STAR: Learning diverse robot skill abstractions through rotation-augmented vector quan- tization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.513557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.743525Z digest=sha256:e4d7fbc777368e74ba83cb6c4057f3b6177c0387db85c09c31b15a3e1b01357d

Observation 9773fdbd-b6ca-4420-b586-1ed3a5cdc433 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Blip-2: Bootstrapping language-image pre-training with 11 frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.502224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.747331Z digest=sha256:2d7e677ca062c013ae06a1d27fa46eae01b0cfc38716b97c2e66c8ce6ff9352d

Observation 05ab51b1-b59b-4d17-88f1-e36a767cf6f2 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.491062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.751038Z digest=sha256:4e59ba8d54d561bd84859a1b1f5ee4d1902af762583aaa3c02b8b7975cac4844

Observation 3898cb60-9eb4-4ae8-ba70-2caf4890a0de · outbound

This paper cites Lion-fs: Fast & slow video-language thinker as online video assistant.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Lion-fs: Fast & slow video-language thinker as online video assistant

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.480351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.754766Z digest=sha256:70537e7c7a998222cbddcb3322efc8364586582c2a1c3427f88046b509492fa1

Observation 92305da7-4122-4711-a3e1-2bb26cb96def · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Evaluating object hallucination in large vision-language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.758353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.758353Z digest=sha256:ab5fc6af4e7255c1efe5aa0873dd2dd3086d33120e969f226e4763faaa26838a

Observation a1862721-228f-481c-a619-8369a8c60462 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.761893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.761893Z digest=sha256:fb3b3a690b6deb5f9cab977422636e5bfcdef6182690003c479114b1ec7a5a98

Observation aa7da1ae-4093-4815-88dc-9d1598f5cd82 · outbound

This paper cites Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.463546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.766280Z digest=sha256:eba32ecbb919869e417afbc33164a0eceb948e292d17d55dd315a1642a73fc5f

Observation 0aa62c3a-5c79-47be-b161-5eb2c705ed67 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.452524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.770199Z digest=sha256:4d7a261d1e4abc18dbc3f7f8976904d5897414a347a732ce23e3d180f3a5eafc

Observation 75e8e79f-8633-4bed-bfe3-4ff0824dd92f · outbound

This paper cites Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.441573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.773810Z digest=sha256:39c55fa17c744cc63cdfd52b8a481e6a7e8caa450862d38df1e2eab72ad4369c

Observation aeb7aca1-8e46-4cf1-bbf9-a55e1f82e333 · outbound

This paper cites Improved baselines with visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Improved baselines with visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.430737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.777251Z digest=sha256:503e8351ee5866d06c72a9db982beed8babeff6a816ee3b5bc9fe687d445c50f

Observation 94754f3a-0b0a-49e2-ab02-bb33b379258c · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.780931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.780931Z digest=sha256:9b259e766e8578ed39ec0ecb21e681fc07a341837cdf2aef48dba1e7516f214c

Observation 403f895a-a8e2-4982-b704-cf18734199f1 · outbound

This paper cites Visual instruction tuning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.784750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.784750Z digest=sha256:4103e411cfdf7b3cb38ede6e3a7f59f63de220ddc4a761b06bef1eca008e9d66

Observation c8afdfd2-f1ab-45df-81a9-258884492a40 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.788312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.788312Z digest=sha256:58a827e5c855669b580af6edbc21d78831e1c00450eb1fd2b03e414d9bd5b576

Observation 99251e67-9df0-41a4-970d-90e9a6a68bf0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.792560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.792560Z digest=sha256:77c721dd3dcc8dc9ca95707685fa52f81614e4bd8b4557f023497a183d0d67e9

Observation 718f2c8d-4842-4d67-9520-ebb223d1e3b8 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.796210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.796210Z digest=sha256:c8cf2c939384d0301fdb6072c45f401a1603061e574ee780e1260fb6b4d36fb8

Observation 80466512-1ad5-4253-a94a-96d2ca4629bb · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.800215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.800215Z digest=sha256:c31f33375623f67456ed4dfd97696e8486de45b215dff91e94e1392e78917684

Observation b5245f0b-10fa-4f44-b04f-5b197bbd3d48 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.394658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.804301Z digest=sha256:b4336cfc0cc5d26d6495f1fbce6add03aaa9dc9b9cc91fa157dce94d0328ddf8

Observation 5c6f3f2f-88ce-4b28-aa9a-a977052721d7 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.807983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.807983Z digest=sha256:48b4ba238f101febfcd3a65ae6e9923e301b5014fdb71ab6fb975a75b284a7ae

Observation 4dfc0037-2743-484f-bd91-0de059dd5dd2 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Spatial-temporal graph diffusion policy with kinematic mod- eling for bimanual robotic manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.383566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.812074Z digest=sha256:2ccd06506985164e8621094d96f565fa5bded8b6b9e43a87ab79dc2a7da0f18f

Observation b58f3cd1-f15f-4198-b84a-0ab867f706a4 · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.372698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.816038Z digest=sha256:015f06069374f3ed9663be57bb8236d8507af3561de628331cd7d10c72144d10

Observation fcf03e1a-d56a-426f-878e-f1710aed48e2 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Docvqa: A dataset for vqa on document images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.361057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.819840Z digest=sha256:2e1b5d546d6c9e35e19629136d64f5a7c586efd3ab95dd526cf01741d6adaa04

Observation 2223dd60-7822-4ab6-8de0-6f1d3d81404e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.823527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.823527Z digest=sha256:43cbb42c125bd68eb5acf599fada738cadc372bde3ddf908470048357aa74abe

Observation 68d28a54-be9a-4828-b622-3bef84431d87 · outbound

This paper cites Multi-adversarial discriminative deep domain generalization for face presentation attack detection.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Multi-adversarial discriminative deep domain generalization for face presentation attack detection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.827229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.827229Z digest=sha256:6b880bdd314c2c3732147859534055e2be01589b2db3bdf9f01a7919d4673ad2

Observation 5682eb4d-25ed-4d0c-909a-1a2eb617661c · outbound

This paper cites Detecting and grounding multi-modal media manipulation.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.337795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.831549Z digest=sha256:294b9e7e153e8c716d5e32127e211f3dcb6f7d5c21385af4d6eccdf9ee70a566

Observation 1ee6b5e4-3a72-463e-91df-12ac496afa7a · outbound

This paper cites Detecting and grounding multi-modal media manip- ulation and beyond.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Detecting and grounding multi-modal media manip- ulation and beyond

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.327055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.835497Z digest=sha256:8b47fbffbedb726f8a76c2f301bdcbd5ed53b8618abcbcf40f80b5f566203785

Observation 302dd13c-0a73-4068-a2b9-490798c2f71d · outbound

This paper cites Mome: Mixture of multimodal experts for gen- eralist multimodal large language models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Mome: Mixture of multimodal experts for gen- eralist multimodal large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.315809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.839306Z digest=sha256:da7bbcbc248e6b89169ec3cfeac1d62233adf3f8b369c38de7de7ee57aa72dcc

Observation 5de446c3-1946-4748-85a2-4b8b58af2f6c · outbound

This paper cites When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers When do we not need larger vision models? In 12 European Conference on Computer Vision, pages 444–462,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.304740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.843134Z digest=sha256:d886dcfeeff11cc02eb2cf7dbda9054eef6686ea75776e2479f83a1b7dfa68b5

Observation f9db3667-dfaa-4807-a952-860ebba09887 · outbound

This paper cites Towards vqa models that can read.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Towards vqa models that can read

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.847335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.847335Z digest=sha256:226223754b0c0e049f18df6a3c4468f359510b5d0d098f132624cf9419afcc8f

Observation 36153ffb-5247-473e-a02a-992f25713644 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.851125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.851125Z digest=sha256:0ba733112bec298fb62cde9329a13301142cf3cfb301981e26ea049f48fe2826

Observation 279cc74f-b201-40d4-95ba-30f69f6556f0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.855123Z digest=sha256:5c5bb2dd12ee096901b3842940aafa688f7d236e47da5b0c6e18508e3c2355bf

Observation 41e6be49-fde2-4170-950c-dea735c419e7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.858993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.858993Z digest=sha256:3c4618ddd520aacef8a00e81ccbb88f99978e103b710d7df9feb996c2030ce8f

Observation 5b9a5693-bae5-4f98-b923-48684cfba1b9 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers V*: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.287477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.863115Z digest=sha256:f60ed3367efc65018a77301bbe0a6072beef4b90d210e630bf42a347151b3eba

Observation 8caefd48-b66b-45fd-98dc-526a0336d167 · outbound

This paper cites Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Gui-explorer: Au- tonomous exploration and mining of transition-aware knowl- edge for gui agent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.276332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.867100Z digest=sha256:84c554e3ee11b7a4c4b39af0972bd791981b059f5666472ffd9c3de80474e579

Observation 369d4a1a-0c92-4c72-b617-b8716adcf56d · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.870758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.870758Z digest=sha256:8acf58fd7b4cd22e22adf1202b4528b34f7b59846531ad253f2803ea21263f4d

Observation 9bd10117-1474-4938-a48d-c544a1805316 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.875541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.875541Z digest=sha256:c61b4d50f5f0ef4c8a1952921391540307d46d8b8ebc7fe312995a72ed5b7fd8

Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.879391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.879391Z digest=sha256:a1b925943285450aeae2011b48514f54d0468dffd2070ee6ea00ff33b28b1678

Observation 57e8c7ba-2775-43f7-8808-f8bb8be2a97e · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.264231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.883290Z digest=sha256:220333d0cad89ecec12e69e28a543a6ea9037fca747c80ee9cb4e4a1f9e70023

Observation 3908d8fc-4886-4a77-916b-13a650794a9b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.886964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.886964Z digest=sha256:b644b1b6599ec12c94a00f66c1941a1f561f8420d01f7ca7dc28b7cdc625437a

Observation d570105d-11f2-4813-bad3-6a832874fa46 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.239916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.891094Z digest=sha256:9ea06d5fa99c374562452bf653fb0b5a672cde845355c4d7e74e3f4bfdeff9c4

Observation 95e99493-16af-43f9-aacd-7f06afa2fa0d · outbound

This paper cites Modeling context in referring expres- sions.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Modeling context in referring expres- sions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.228709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.895145Z digest=sha256:b95cad673a2799972a6a6c99056b1a15e82a7c740150e304e9e92f455a6fe8e8

Observation 288adfb4-a5bb-43f1-96d3-a928edd16fbe · outbound

This paper cites Sigmoid loss for language image pre-training.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.216519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.899089Z digest=sha256:fbb023ceeef1c1fb827c41eef25ce4a204897d4e8a22a124797f24ffa03d15c8

Observation 7f24491c-e4a3-4866-b6fe-c2441f230af3 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.903154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.903154Z digest=sha256:32c4ab888555b051b2f0b0846093496957b1f717535788fd465471adafbd9f18

Observation 8fbc5767-86a4-45e6-85df-5737da82d995 · outbound

This paper cites Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.907427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.907427Z digest=sha256:7d2d681d21024d75853961cde3110095e52508e359175ea76537c4b30692fbb5

Observation ac5d1f9d-a70b-4c35-8dde-30a233ecbc2c · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:38:19.199802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.911600Z digest=sha256:46780c438cdb6fb75fd86938ed4325612c849014cd870595531d1930b59a0e22

Observation 5a52d847-6c2a-401f-a6d0-37d6db3f6f4f · outbound

This paper cites an unresolved cited work.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:38:19.186362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T13:38:18.915446Z digest=sha256:34b23d5d7d67e19b0bef380b3eeb4664d6e6f4424f6e700ef5d9e017fd76d220

Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.919283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.919283Z digest=sha256:37083169fbf4c72c8be45129f8843f2e690907a35e3d6a55feca50c0670e7923

Pith citing papers

Observation 1e720cc7-46b5-45e6-8500-60bd9775742e · inbound

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries cites this paper.

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.790283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T08:14:28.540629Z digest=sha256:3f1ed55433b53c6961f52941b06703977c65528256de68418ec426aa403468ad

Observation 308599d9-d0d2-4720-a799-43695fdad954 · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.667217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:c9690341ab44b184d524b338db9b585f0592cb932c8e36e06799c68b8d5239fa