Pith. sign in

Paper Citation Record · LEDGER

Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.03003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.03003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:46.509499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:37.024796Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 776bfcc4-13b6-44ce-8048-2b0e1629f187 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.148522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:5aecb21138fa7792ea7220f6ed5e290d6b791c1829793e44a02936bcca2e5277

Observation 7d0232be-36ae-4f25-a436-a166fe6a183e · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.509499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.509499Z digest=sha256:1311b59c028081ba9dbe8cd37402f0b65a0f14b81d8cb8242967e6d9cf471681

Observation 8f9db7db-44f9-475e-9bfc-ce67c36f08ea · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.176832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.176832Z digest=sha256:5aa80077ab4442a26483a2ad3cc2bb9e27aa0848ee01d7d91fc891e923a55bd8

Observation f46ba898-242d-4dcc-a67e-660f1cd1d0a9 · inbound

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs cites this paper.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.837013Z digest=sha256:14402d21a7654f96b0c43dce9690dabed4e18d95b0dc25bba51c1f3fe086e3b6

Observation 56555408-5b6c-498d-a63f-ba3f12f10b05 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.688789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.688789Z digest=sha256:02010a4dea94661def7dac757acf17317e72ca05a0f67d21bb37ba2bd3c44aa8

Observation 4e6d7fd3-13bc-4b76-952b-ad5ea10882ba · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:52.717672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:52.717672Z digest=sha256:f06b7402656b794c3681950c0c123231cc36edfab00e280917cc3865cc9096d5

Observation 5c50d509-3051-4a26-a553-5419dcc3eb98 · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.441648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.441648Z digest=sha256:417984f699c2d2d4a67a050ffdab9b93a68d39f77e16b0d596030bc363f51496

Observation 771b976f-f12b-4fb9-b996-ba86b6d8470f · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.606698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.606698Z digest=sha256:890b2d59ab692e18cf7379c4536df469d4ba49d3009a881942f1205920c5e8bd

Observation 0870f533-a5ce-4bec-b390-43a3a9d4e9ab · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.018551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.018551Z digest=sha256:dc8a8fe7dfe6162602ecbd70d7e0462715819ad82d6842eba47c8b671960eb20

Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.312239Z digest=sha256:440bd495fb1f42535d52ceff005aeb4e26abe8bd0c36d7e938951294a751a71d

Observation 3ff3c3d0-80d0-40bb-9d61-d79f4b14c792 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.557208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.557208Z digest=sha256:ccb49b1bbdd1c97ddb5457319734a26993037a0e875b418eb2ac8caf4fad7553

Observation c64fb980-aa3b-45df-be78-a21e7f0a2fe3 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.794185Z digest=sha256:309efdd9e4eb70be39da1795da8e25886beb05b6f518049d7aab5c6f0d6e837a

Observation 382d78d9-c7aa-49df-baa3-9797b2cf1ffb · inbound

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning cites this paper.

LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:47.820050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:21:47.820050Z digest=sha256:278eaaa0cc97ba8fba41ab0d908b6b3e040bfeca8b30892f9e1fe6cfe8564495

Observation 04f92453-e9e4-4bf9-8fd9-dafccf4c7bdc · inbound

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images cites this paper.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.230169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.230169Z digest=sha256:35dd766879c48ff3208067e025fcafa2b61e0c802327532353b99b7d885df525

Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.463091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.463091Z digest=sha256:f0824d1c35c0ebe627b2882ca8753b8428c00bc465f26471eeca38253172214d

Observation 1a662a34-b84c-4208-82ba-1679f9af911a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.211532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:3c138bb1b59a08b714d1f18a7979b43619efddd0a1cd2139b7e4f5d01dd77f7c

Observation 4a4d3bcf-c754-407e-90e5-db416ee141c3 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.680272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.680272Z digest=sha256:e4cf5cca8eec573bcfc3179ba396697376ed6d4d481c03217f900bd2031a98a5

Observation 2b9fae99-2fcc-43ea-90d7-e24bf41d54fc · inbound

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization cites this paper.

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:37:24.213930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:37:24.213930Z digest=sha256:7de80855aa83d2b4ab29dfcbed184f9704af3c22ac6a524645d70265bc6eb417

Observation 5d4fa2c5-97c5-4a3a-ac41-3c84ed8ea622 · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.587093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:05037fc67e59f122b4e5392782d0bb257ccd0845386951dcff7a6d7a6abd7bd4

Observation 31968c7c-57bc-4110-afea-fb5180460cad · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:52.974101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:36e7962a134cb68d5e0ec762cbe235e3bf004c04bcab339844d85ac5e2f306fc

Observation f5ab41ee-f22c-4eaf-a99c-796d81535f0d · inbound

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception cites this paper.

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:36:02.563120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:25:46.990772Z digest=sha256:8967e294db1665acc79077fbc7f954aa4115d3f9c797f7cdae2d614ff2d1e393

Observation d7e78842-bf47-44bf-b04b-89513cee8a1e · inbound

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing cites this paper.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.929452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:7afa442994889fa1df135702af01c44e383f3166c39adaedf570c51c0f1d2a70

Observation 3a711b48-abd4-4610-b541-3cef14ebaf95 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.116103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:ad63676d24653206ae570101da309a896b7e344d423610220e330a2e938d1b2d

Observation 2189009a-8efa-45d0-a5fa-61956e228177 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.004034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.004034Z digest=sha256:88de40e56b1ae0db4f15c8adc8d112ad0a1a77581f4211ac85cb9d8511526ae6

Observation 6294fd61-bcca-4fd7-947e-787f214c9949 · inbound

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation cites this paper.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.279432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:33:12.116833Z digest=sha256:b65b2b25f64579eda4e8e9037596004b9389a345654a85364195d5ce7ab811c7

Observation 8dba0300-11ef-4796-891a-62114535752d · inbound

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs cites this paper.

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.520798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:08:36.451147Z digest=sha256:d20461c7f2d2f618f791b4a4469b02bf54e91252417535905a994dd17e97c94f

Observation 84f4dd75-9489-49f0-af22-a9d0a2a86fb0 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.026443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:2366a00682b22d570e9abf421917a0688e522a8211e93b0cd44ea94af732dbea

Observation 44400a52-b8cf-4792-98e7-05d0ca7b99e1 · inbound

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models cites this paper.

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:48:44.720562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:48:44.720562Z digest=sha256:a5f4a5a64d2429129b03103bfe8f80a21a9e0f6cbf64f47f12b182a7f74e517b