Pith. sign in

Paper Citation Record · LEDGER

Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2311.06607.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.06607 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:55.752574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.479106Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cb96e4f-8e4b-4928-83dd-191236d49a52 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.471878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:ca085b710d1df7e71ae77222204a9aaff8d75cddf70beaa992debeb56e8313dd

Observation 77798ffe-5e13-4a9e-9b18-b33c9ab5be2a · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:20:53.896871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:17a2ebf7f05e931796f9bb511c5c2d8dc4e96afa813fdb2cff79027746dd7f2e

Observation 2dc1f223-840d-47b3-83b6-30d50bd8bf64 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.003085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:0d4df1f57b42c03c65bfbb97ddd54619106ff7c70d7ee886526ca9327185ee07

Observation 7194c1c2-9fc8-4a1f-91ef-9ce83b2e3d93 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.732796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:8e2445dda1e12860ffe0b9aab28a9f3f64e5d16e4fd72a7ff87bed942eee1f2d

Observation d49e1eb5-1728-405e-ba7e-138d7f79f98a · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.326163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:13479b931c3b46197daaed933ae555f9eca01a3d27daede60e04014ffaea5835

Observation 6b69ddb0-f51d-49e8-ad2a-daef9811f7b1 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.493562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:d5a50d8563da21f144d9f49bb50b1f96937c70cd2eb60919606a764cee450a6e

Observation 245f14d7-b4f0-49f3-ad45-72fc13368c17 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.535674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:7680224d0b48f5be4ca981fc77fe19499ac4be5d4f84f7e988a65ae77609c847

Observation 1ed97519-ff46-4555-8a73-4419df871bfe · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.053339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:05b8e1ae92c192a1231f2b0f7d9ba86f7d348e9415bdbaaa5c5ef2450b19c0db

Observation 013f2f03-743b-49cf-bbf8-339868f8f464 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.688589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:8e4bb981f1aed6f7d6f87530ff7239f2d4c6337423fc12dce162c8ea9c9cfdd1

Observation e2c0c614-df4b-4e47-8fa7-185b0efbc18e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.808098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:94369e40038ca18ab06ef859dd97fbdcbafb81a66f5313d48aac3a733684fd13

Observation 9d3a7fc1-83a7-4d6c-9880-70a9f42332fa · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.703956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:649f48ccd878e6654af41cac236f2bb64d598a9b60a15fa3d14233a945e11b2d

Observation 27a21b8e-3b2d-4ad6-8b6b-c8b3cfe7171e · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.253757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:1b23887f737489141a9e4ad6caaf754759262a92e9ccffb57665eda515a92c8a

Observation 31365a7f-c765-4263-a232-b30f7ed0088d · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.752574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.752574Z digest=sha256:996341e7a099e7fdc14463d7975915a07e7852a7fd26cc47174bd6543cfbee68

Observation ea2a5ef6-e5f0-4db7-bc73-a67dbf644f8d · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.854770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:cf85603081da07c3b20b2ee440cfee041849e5f0cff2a4d8f79e8fe73e23a70a

Observation 3984b339-064a-43db-9912-503af55cc6e9 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.362333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:29fa09371a1f7065f395e0242016a17b21069358381a1e4598e0b954deebbeb7

Observation 810e2824-ad71-43d2-b5fb-4cbea895bf9b · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:21.845958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:21.845958Z digest=sha256:3b79768f6722c477b238caa34975e6ba3e5c38eed230decbcfab41f894ba56f7

Observation 8c015068-9151-440b-b307-0fba713fcd06 · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.979543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.979543Z digest=sha256:d407b5c39f9a36bdf0c8d3ddbeb8bb4842f42f71e1eb6eec8290c419b0d9d5d1

Observation cc1567fc-3ba2-4fc9-9c42-1ccde88ab3d8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.542586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:52.542586Z digest=sha256:24f26a4df125fb0428517bbc3976c351bb636f6d8935066571e78fa37b26adc9

Observation 543ae786-de83-4dbb-b149-4ecbc8f6fb43 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.134795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.134795Z digest=sha256:bded6b5d3e91d95a0f83d43d288a7c2a37b220eba7287e4c3a35ec6c1723572e

Observation 88447597-4568-469d-a634-9e20ed3c5b47 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.931321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:9a71e44c3e5a4012c7ed60a2b98d01a33b4433e5252ab45c427724bad0a685ef

Observation 57eecde9-6e1e-4448-9fe3-a94055097536 · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:19.334924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:19.334924Z digest=sha256:2a4df78160186f59695a773db3d431ec4687888fe48da6dbc3cb353dbaf70c2b

Observation e15e1f92-b278-4d04-89e0-b91bb83ed2e0 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.922398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:ec43975e154fdd5422511dd5773c60361429f7bc82484ba8146f9a900aacbca0

Observation 0ed651f1-a7c3-44ae-8b8a-c60a789c91ac · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 279

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.480709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:156e36f50efe97b42cc775757dbe617f2802eff2fc29e09ae3122042fcefed1b

Observation a6cd8e01-3f80-4ee3-86ce-87e3715168a5 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.426378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:9bc9bea02e77bba794dda8a55c4ce034f6a18af18facf7d00332efbd3bb10eb2

Observation f03fc67c-d05c-4124-ba0c-b89f0ee94c29 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:b334c8a7c92bea43162c65d348897a0a5eb001efe7d0980782bb7cae84c15d1c