Pith. sign in

Paper Citation Record · LEDGER

Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2311.06607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.06607 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:55.752574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.479106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cb96e4f-8e4b-4928-83dd-191236d49a52 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.471878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:6cb3ca9494447579e480adbfe2c12ab474ff35969a3decf755eec9869a508188

Observation 77798ffe-5e13-4a9e-9b18-b33c9ab5be2a · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:53.896871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:4576453d185479b49591646ac1e7edca33bb37b48256e61131b055e393e1bd37

Observation 2dc1f223-840d-47b3-83b6-30d50bd8bf64 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.003085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:33f40a1387a7bc87adb3d7c86cf987514ee10cef478ae59f47e5b56b07395712

Observation 7194c1c2-9fc8-4a1f-91ef-9ce83b2e3d93 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.732796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:3c7b10105bff442507d4148a2ef4df70b5643fdbb27b5d5a0489ec157b459cfa

Observation d49e1eb5-1728-405e-ba7e-138d7f79f98a · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.326163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:d8b69551b1dec89fcc6d7c8932b8a6d6b9989d54ce3ef6e732ede4d28d80d371

Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.493562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:e867a923be4e0d843c05122e311c18bdbd43f8a099bf4ba11dc81db89e8f205f

Observation 245f14d7-b4f0-49f3-ad45-72fc13368c17 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.535674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:18df10ec9ebae01438bc8c65ecefeee82b8c61bf1969402f74494740ac13e6ff

Observation 1ed97519-ff46-4555-8a73-4419df871bfe · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.053339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:37eba9c8edc1c80669be202e913e0f44075bdce84a0068a818b70ca673cc4f23

Observation 013f2f03-743b-49cf-bbf8-339868f8f464 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.688589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:10a8694399140dd9fa85314c1abef88bfec46ccc6fbdfeedbfe9de501c3a7259

Observation e2c0c614-df4b-4e47-8fa7-185b0efbc18e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.808098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:f64ca86e3a211d75a06a9e26339dc611aa7e8f2984a5d5f537ab950be08f50be

Observation 9d3a7fc1-83a7-4d6c-9880-70a9f42332fa · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.703956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:20c4b4e6f0d93bb452f189193ab02f31009efaa36bd1137977d6530f36baa0d1

Observation 27a21b8e-3b2d-4ad6-8b6b-c8b3cfe7171e · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f9cc97a4d16ace8ac23606aa2b56c1d416849c2f61e93a0cc33b64cb0300b216

Observation 31365a7f-c765-4263-a232-b30f7ed0088d · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.752574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.752574Z digest=sha256:996341e7a099e7fdc14463d7975915a07e7852a7fd26cc47174bd6543cfbee68

Observation ea2a5ef6-e5f0-4db7-bc73-a67dbf644f8d · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.854770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:7ba6c02c7868c46a431478663e8bf5b8ccee1433f4d42ba3ca0196859ad7a14e

Observation 3984b339-064a-43db-9912-503af55cc6e9 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.362333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:adb001e683e4c4792433e08ab26e269c0b0027874a2052148173ef773e75fdf6

Observation 810e2824-ad71-43d2-b5fb-4cbea895bf9b · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:21.845958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:21.845958Z digest=sha256:3b79768f6722c477b238caa34975e6ba3e5c38eed230decbcfab41f894ba56f7

Observation 8c015068-9151-440b-b307-0fba713fcd06 · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.979543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.979543Z digest=sha256:d407b5c39f9a36bdf0c8d3ddbeb8bb4842f42f71e1eb6eec8290c419b0d9d5d1

Observation cc1567fc-3ba2-4fc9-9c42-1ccde88ab3d8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.542586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:52.542586Z digest=sha256:24f26a4df125fb0428517bbc3976c351bb636f6d8935066571e78fa37b26adc9

Observation 543ae786-de83-4dbb-b149-4ecbc8f6fb43 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.134795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.134795Z digest=sha256:bded6b5d3e91d95a0f83d43d288a7c2a37b220eba7287e4c3a35ec6c1723572e

Observation 88447597-4568-469d-a634-9e20ed3c5b47 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.931321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:923270e07e01b2d08b2689491b7ac1820c4ea5eefa75036ba8d21b7fbd669050

Observation 57eecde9-6e1e-4448-9fe3-a94055097536 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:19.334924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:19.334924Z digest=sha256:2a4df78160186f59695a773db3d431ec4687888fe48da6dbc3cb353dbaf70c2b

Observation e15e1f92-b278-4d04-89e0-b91bb83ed2e0 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.922398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:2f126ee69cc718d847fd021d35354c042e1578f375b2e41ea933f99d3097f90d

Observation 0ed651f1-a7c3-44ae-8b8a-c60a789c91ac · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 279

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.480709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:52ef1091461fd99d7b7c1d65b16abd118b8b20db3ade0850cd4f181fb0c4ba9f

Observation a6cd8e01-3f80-4ee3-86ce-87e3715168a5 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.426378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:abe54661d2fab9d9adfc59883b4144e5a20fff9054c65628fe305ad40c1f4cf8

Observation f03fc67c-d05c-4124-ba0c-b89f0ee94c29 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:b334c8a7c92bea43162c65d348897a0a5eb001efe7d0980782bb7cae84c15d1c