Pith. sign in

Paper Citation Record · LEDGER

Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2403.03003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.241899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:37.024796Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 776bfcc4-13b6-44ce-8048-2b0e1629f187 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.148522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:d0264ec6ce047205e542f54d98fd37beadfaec3ed1d09fab1447742076ecbaa9

Observation 479d4a04-a8e8-44c7-89c6-c9552dd96ba4 · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.241899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.241899Z digest=sha256:2117a3bb0537d66bdf2a60021c0dd2403fd3b9df7288926731bd19902315a1b5

Observation 7d0232be-36ae-4f25-a436-a166fe6a183e · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.509499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.509499Z digest=sha256:d17215458df79a305d96f292dd82285c932944de6987fe4e0b71405af094cbcb

Observation 8f9db7db-44f9-475e-9bfc-ce67c36f08ea · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.176832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.176832Z digest=sha256:079e20935d1f749ce3ad8dbe99d0bbe4fa5c8c722f5b5ee3019322405a0280d9

Observation f46ba898-242d-4dcc-a67e-660f1cd1d0a9 · inbound

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs cites this paper.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.837013Z digest=sha256:444bb182340614cd84f4795bcdc822d140ea2af230a9ae33d026a7980fe4777e

Observation 56555408-5b6c-498d-a63f-ba3f12f10b05 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.688789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.688789Z digest=sha256:a63b9fce29bed29bdb7cd356008254ddb630674defa0d3591152f0fc8c6039fa

Observation 4e6d7fd3-13bc-4b76-952b-ad5ea10882ba · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:52.717672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:52.717672Z digest=sha256:c95f21a3888c5d48f30b08feb701e4cfcdd05d3bafe9de356ade4e1c963b20ed

Observation 5c50d509-3051-4a26-a553-5419dcc3eb98 · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.441648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.441648Z digest=sha256:24c07fe2cfa8a9760e81534eb887715475eb4c522b7deece8bf8224d973e2ba0

Observation 771b976f-f12b-4fb9-b996-ba86b6d8470f · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.606698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.606698Z digest=sha256:890b2d59ab692e18cf7379c4536df469d4ba49d3009a881942f1205920c5e8bd

Observation 0870f533-a5ce-4bec-b390-43a3a9d4e9ab · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.018551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.018551Z digest=sha256:c32ba5ff29c7107d0f6ba25cbe4c61aca21439ea5b11cd60e56a9b52303c91b6

Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.312239Z digest=sha256:210bfb1c50041935b8459eebbfda38b918c95e25bf9b2097edae98dd70b2158c

Observation 3ff3c3d0-80d0-40bb-9d61-d79f4b14c792 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.557208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.557208Z digest=sha256:ccb49b1bbdd1c97ddb5457319734a26993037a0e875b418eb2ac8caf4fad7553

Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.794185Z digest=sha256:b4740f0b6fc7e94f7f56b31a8c03d5be8fbb7a5034a6565a3ad3d699a7c9ec17

Observation 382d78d9-c7aa-49df-baa3-9797b2cf1ffb · inbound

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning cites this paper.

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:47.820050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:21:47.820050Z digest=sha256:278eaaa0cc97ba8fba41ab0d908b6b3e040bfeca8b30892f9e1fe6cfe8564495

Observation 04f92453-e9e4-4bf9-8fd9-dafccf4c7bdc · inbound

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images cites this paper.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.230169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.230169Z digest=sha256:35dd766879c48ff3208067e025fcafa2b61e0c802327532353b99b7d885df525

Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.463091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.463091Z digest=sha256:64dc5462c933bab311bea983424d4b1a7a63121db866be54536f394055fd1cf0

Observation 1a662a34-b84c-4208-82ba-1679f9af911a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.211532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:e662e9e756add1689ba843d67c522c4e1895e90ac838d1caa3024e3099f7c5b5

Observation 4a4d3bcf-c754-407e-90e5-db416ee141c3 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.680272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.680272Z digest=sha256:ad938a3d0e1da43ec76a33a5ce8a1aad978a29e92780ecf8a7ac54908c3f81b6

Observation 2b9fae99-2fcc-43ea-90d7-e24bf41d54fc · inbound

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization cites this paper.

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:24.213930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:24.213930Z digest=sha256:97efec20f9b6e77fa75a4bddc53769333d46e32f6a81c480385a9f804fc0670c

Observation 5d4fa2c5-97c5-4a3a-ac41-3c84ed8ea622 · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.587093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:ebd6f59ed0cfcf8153303a7b189a3a94f146e462b8a47f42e584c14c1ec01d25

Observation 31968c7c-57bc-4110-afea-fb5180460cad · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:52.974101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:30e9f4d5d0b9102b89d5f9ed5f2da5854d2d415f89faf5e6d772015091484b17

Observation f5ab41ee-f22c-4eaf-a99c-796d81535f0d · inbound

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception cites this paper.

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:36:02.563120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:25:46.990772Z digest=sha256:5b63c5d501c5aad36498b6dec1014246afc80ca2508e6b5bf57a15f462973c32

Observation d7e78842-bf47-44bf-b04b-89513cee8a1e · inbound

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing cites this paper.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.929452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:fd1bce72ed6de4410cc9539638abed47d4678f89e23d14d92cf153bfb2dc00d2

Observation 3a711b48-abd4-4610-b541-3cef14ebaf95 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.116103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:15156af731d8d1c3fcfd26d9cdbea3a14b130f5b1ce446b0a04347ce2a1d1a36

Observation 2189009a-8efa-45d0-a5fa-61956e228177 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.004034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.004034Z digest=sha256:88de40e56b1ae0db4f15c8adc8d112ad0a1a77581f4211ac85cb9d8511526ae6

Observation 6294fd61-bcca-4fd7-947e-787f214c9949 · inbound

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation cites this paper.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.279432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:33:12.116833Z digest=sha256:413d64d478fb562943b206579a2a6fd7916c7eb3ae1bfe175364a8bddce83e5b

Observation 8dba0300-11ef-4796-891a-62114535752d · inbound

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs cites this paper.

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.520798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:08:36.451147Z digest=sha256:2b383e86658758c5897b7d5686c194eb839a4e95c82e5394209348d5cd9c1738

Observation 84f4dd75-9489-49f0-af22-a9d0a2a86fb0 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.026443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:adcc9e8819de37448291d1382b0eb3c4e98b94bf65a2c1add1bcfc6d5663568b

Observation 44400a52-b8cf-4792-98e7-05d0ca7b99e1 · inbound

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models cites this paper.

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:48:44.720562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:48:44.720562Z digest=sha256:aae509c9a5da404c5b8acc5d165e57f21a89d11820fea3dc3d8ca7ce33e825c1