Pith. sign in

Paper Citation Record · LEDGER

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2506.12776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12776 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:35.027792Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:58:53.702009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:37:14.448914Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0e82b7c-81b1-41c4-a0b5-6c052a6f472b · outbound

This paper cites GPT-4 Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.130945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.130945Z digest=sha256:ad3b5220fe653a649d453ea4fb70021696bc10a9c50829b7621fb13a41e89ab7

Observation 6369f684-748a-4e43-9414-9bc5a166ee8b · outbound

This paper cites Qwen2.5-VL Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.212811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.212811Z digest=sha256:d0fcbe994d91bf81c9cd0fc3f04c54bff3b5d7abd1beae965dc6c3cb67d8d23b

Observation 80ce4dc1-81d8-4543-8384-df02b0d133c3 · outbound

This paper cites an unresolved cited work.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:48:37.930240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:28.348014Z digest=sha256:77d5947a9d24ceea63c515ca816ce518ae57e2ec73b783c793ff0a5f474d7b7b

Observation 3a605ce2-499d-4e75-a0b5-7159b53b774b · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.448857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.448857Z digest=sha256:f844769833a83cd05ab221eff3261567fb9b5744cee6d99af59260b58ddf8e0e

Observation 92ffee03-72d9-4c73-9c81-4513847c6dcf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.559091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.559091Z digest=sha256:befcf28294868380870325aa644130a8c2d06e651f7610dbc4c20218ca946f6b

Observation f7d0afd1-97d5-4d80-80f4-e03ef8c6561c · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.686334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.686334Z digest=sha256:b76b4f1fab8589834a3f447034d981a9ee0b3b974654a20ad07b5e3265f49ff1

Observation b260ec6d-0bbb-4899-8e6c-51b9aff3a854 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.785387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.785387Z digest=sha256:7e7b800246fbcc13d9defe3a67b49e48a81dc3c2d8cc54e6f2f01ca3d59d592d

Observation 72c52b55-10e8-4431-bf04-c0040d6601aa · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.885481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.885481Z digest=sha256:f7cf979d6ada5cf9ae1e7e51f1f9bc65b927e72dc17367b4c963e73196f51041

Observation 15e0b855-cade-44c4-8722-8e0b0d8f28aa · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.989104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.989104Z digest=sha256:93b09759090f214f3fb4117bd0b70b3783f29b0adddfa6ae1a79ecde37dac31f

Observation afc9fdcc-c093-4040-8fc8-8c54c2b820d5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.154407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.154407Z digest=sha256:905aacbf854562b488d71c0ae3388d59da5d1fffe0960b884bc7de2fa1d37ee8

Observation 34d6d148-e214-4a23-b785-302d79a40750 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Gpt-3: Its nature, scope, limits, and consequences

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.749092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:29.255022Z digest=sha256:7d1bd9519fab42e4cdb4c7b957376834ae19038b53c1a3df9c071697df645310

Observation e32e37fb-45e4-4006-ada8-2c8766dc3bcc · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.381519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.381519Z digest=sha256:0d12a3fe8d2399a1563b9beff1e3fd5ca87875cb56cafd85e6098229b7635047

Observation 79bd0a28-93a8-4458-9fc0-5186a0340428 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.512536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.512536Z digest=sha256:373ae24634c9f6677d36d767607098fd55a17cb0ca6c64c347c07d0b2a7bf9ce

Observation 1d2629bd-4eee-4440-90d2-4227f3240ef6 · outbound

This paper cites Seed1.5-vl technical report, 2025.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Seed1.5-vl technical report, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.524920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:29.634978Z digest=sha256:bb4a00c6032c83afff5d2b768b6689f0a5bed34033ee3ad164ad80fb05de85e9

Observation 383e8978-26b3-4a4a-8418-e1043bfa123e · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.308102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:29.741955Z digest=sha256:a2d9e87b00e7e56265c78e1d5f42ce58bf133ac44ccd4f1d6c1ef5295f03f048

Observation e7ff2c4b-6a7e-4a70-a0bc-1de1c4fcf400 · outbound

This paper cites A diagram is worth a dozen images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models A diagram is worth a dozen images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.843867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.843867Z digest=sha256:a38b478e55c6f0b2e0d7b50a63362040a6e4a5c8a2dd6549d6e08ac721ee724a

Observation 757c8926-79bc-4b83-8eef-656cf9bd9112 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.931868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.931868Z digest=sha256:b9064a3b5d0de6bb1a727d32be631173fc6937513daefdc1c0e02c07fa81ce33

Observation f45a04d0-49c7-4083-9a7b-458e621671fc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.085311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.085311Z digest=sha256:c42a7f39c6a3a65e9797045f8b7c3f3171b3fd6cbb0897f1f6885e29fe88b0d8

Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.185030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.185030Z digest=sha256:47b15c433f6ebc06186b6f5f97f997a9db4aaedd4d74d77e68dc4acf560782d0

Observation ecc99e65-9cfe-4be9-8062-c03588f67105 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.316110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.316110Z digest=sha256:786ad6ba1a28e3ce1c5ae686ffe6bff21c2c74afe3e9c9b100a850d309caf1bb

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:2441b4689911afb586bc874d3f7f38aea21bc58fd7d735058de6b59a19de037d

Observation c8e8a796-7b0c-4d36-bc25-67623a38812a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.535552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.535552Z digest=sha256:f15d59f62e65ba4c65e92e7942db8c4ed25fec497e102eaec0922450a696d3a2

Observation 4544fb5c-1eb9-4760-9932-a995bd17dc91 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.624519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.624519Z digest=sha256:d1f400c15b3bc49f9c34551e0508d7a3ca121dd25f0207c22628fa0654ce1fac

Observation d39cb784-ac20-49f6-858e-3926719fd61f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.741494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.741494Z digest=sha256:656fe4b8b748757676a3ff832336c9cac1b59faf10a3db614854773a84c5695b

Observation bf36f4bd-1fae-47c5-85ba-3a755db47a5a · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.849435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.849435Z digest=sha256:1a758106b94851895882e91ca09fb79e7e0d41dd6e91743d695911a44de535c9

Observation 316f67e6-7fba-4b5e-84ee-da9e18e278e2 · outbound

This paper cites Improved baselines with visual instruction tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.977433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.977433Z digest=sha256:0ba616e17303a908d9019096c06426a4b46dd8bb93187fdc07a92ce4f251da2e

Observation c7c5f576-1235-4f76-ae9b-4cc6d7c4d51b · outbound

This paper cites LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.142439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:31.082464Z digest=sha256:694c7039f559dc70be0ac17dec210f8b73d53db62f992c86b044da87a3832178

Observation 50640e39-879f-48dd-9f2f-013a9676da49 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.182343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.182343Z digest=sha256:8453602f3ea5868e05827bfdb9b8c076f5b6c65238831b09360541b65947255b

Observation ef186e24-61e5-41cf-bf81-fda06bac4f5c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.252154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.252154Z digest=sha256:8e51704da9d8012ce781a08e750f8a575ef6b699b39accc42f48c0ee5771d8d7

Observation 7f8843cc-3055-4ec6-9c48-159ad1fb97a6 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.366111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.366111Z digest=sha256:610c19c8b2ebf070f06719a68be7fb569ee81bd063bcdb0448cf2e2106833533

Observation e39b5d73-5718-41e6-9886-f2394b594b86 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.424586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.424586Z digest=sha256:5b8a318588c8fdb58d563f25c07c2c2b55e87b92d8517e05f59493bc71458ff1

Observation a607c0fa-11a5-46ca-8516-3e96c1384d3a · outbound

This paper cites Decoupled Weight Decay Regularization.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.521261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.521261Z digest=sha256:b814fbe1dadacb5e461ccab3001d0a0c5d5ef6795f07ba359c307f87234dfdc4

Observation 5fb20c4f-02ed-43fd-a417-9700eb54ee93 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.608767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.608767Z digest=sha256:36e8f0be13dbd94da7f9ff0e62cfda754c8b79a2aa28fe7b959417d9f7264cea

Observation 56555408-5b6c-498d-a63f-ba3f12f10b05 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.688789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.688789Z digest=sha256:2f1b3bdc0a2abbd898f61315b548bb31f3f84fd0690247c2492343b9bbd904c6

Observation 888b9194-5eee-423a-a6c0-3abbdc2e8561 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.784648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.784648Z digest=sha256:9a639698feddff2279b9f18e26fbbd23a63a7079a129d030de7c8477a5a8a3db

Observation 3622c02b-b745-4a0a-84e8-765ad3c35cd1 · outbound

This paper cites Infographicvqa.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Infographicvqa

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.859231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.859231Z digest=sha256:b85d1670127a1443fe6ac30357313ee358b977ecd16775a2e91a08b01fcc856d

Observation f2035eb1-a169-47b4-a1c2-e9e7a65c513c · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.930932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:31.955507Z digest=sha256:508d182d43d5ca3f61728d9c38194030902254022f88e08fac9c30be820e4c9e

Observation 6a20a55d-df00-4e28-a867-2b132bc3d205 · outbound

This paper cites Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.759217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:32.047540Z digest=sha256:2212ea0739a6e29cff540eb49d7a56567eb8117cd343d2ba50dbda1911df95bd

Observation 9b0c52fe-66d5-49d3-8796-b11c09c98a05 · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ovo-bench: How far is your video-llms from real-world online video understanding? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.557679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:32.169757Z digest=sha256:8a0684d57e1a468cc8b13cd3d2fbd58605d91cf9c3a0b41ea638cdeede45f96e

Observation fd369666-9d8e-44d8-a40e-dd555445bd52 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Learning transferable visual models from natural language supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.297562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.297562Z digest=sha256:a439e5263b9de7dba679a67e38111a41bdc8fd28f10cf50766168481348fd682

Observation 396cba8e-2143-49c2-a0a5-5cb1159e68b6 · outbound

This paper cites When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.296805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:32.404618Z digest=sha256:d0bd65e60c00db735a140425e04a286956d1eb67a0bd9680f8ae1b65675b78f9

Observation 7ebac53f-50e3-4040-a501-c1ec785f2b0b · outbound

This paper cites Towards vqa models that can read.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Towards vqa models that can read

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.488516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.488516Z digest=sha256:e5c517dd8c05f14dfab829bc53be6a680e64225282305404bb62ee3878700625

Observation 7f22c7c6-e324-451c-bd75-660a10d00b58 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Roformer: Enhanced transformer with rotary position embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.571426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.571426Z digest=sha256:5d63eaf4f577ab75ed1eda7cabdfe9455647c11074196d094f00eb432854184e

Observation 26a6e65a-bf4a-4826-8a9c-fafbb8a22589 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.667962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.667962Z digest=sha256:28bd674476a7a6ceaf6ac92fc679e76e874691a9278176a9f6c0806bda402937

Observation 7d899ee3-e8d1-4120-8aae-aebb65305b14 · outbound

This paper cites Kimi-VL Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Kimi-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.740707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.740707Z digest=sha256:a046984015387ccd7f19cd359da1ed4952489f20e5504f3be4b0ffb03cdf7b80

Observation d2e9f43c-e693-4959-852e-547fcdf21f5a · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.855555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.855555Z digest=sha256:1642d44053f37b12188d56973c1849ebb4048b49a62f4e97d45b7b48572f9271

Observation f4b0045e-7682-452e-b341-2d0e1956b838 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.992231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.992231Z digest=sha256:8e20c14f6200dda61681cc0cce62b1d32dab5dbd4349b1668a679e9df2408294

Observation 2cc87ade-6243-4ef1-8afa-2e443bdfbefc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.114735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.114735Z digest=sha256:3d8cb36ba043c6f5ed87f5572666a40eeae3748e9bf28020059e6df64a357904

Observation 32f6e633-21d2-455d-a48f-02b01855bc70 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.091872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:33.213010Z digest=sha256:262dd0942503fa1c9f913c8d842077bf6bd0f39ec1845f5182c10db44cbb242d

Observation 3b412e13-b0c9-482f-9f09-a2dae75db265 · outbound

This paper cites Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.338434Z digest=sha256:f8fc9f1a454ea5c2201267bbbf9e00d241bbc3423f9d4e1b26ee92f0bed8af0b

Observation e80137b3-24e5-4535-8511-8cd6d336417a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.446924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.446924Z digest=sha256:45cd0a4cda02661f1eccd572d241580227578fb5b603428d2701babc6b5cba3a

Observation 1b8db740-51b7-4499-9a91-3e129af10f55 · outbound

This paper cites Qwen2.5 Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.535196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.535196Z digest=sha256:c988653fec271e66424b0d74765b3fdaa81be52e156f30a4d3c209c83612c4b4

Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.654963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.654963Z digest=sha256:26df180a654d7219086759d263154da991b2984988fdb2fd60b622bffcbff3c7

Observation 94d73980-c679-4485-84c2-0af603b7dc22 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.765571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.765571Z digest=sha256:b885d813d046d3a7f88d751ec6a2b414e27ce78b9e5943d8019dd8a1b4ced57d

Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.913826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.913826Z digest=sha256:6afb21e4d399882ac94784c01164c363f58d4eee1f505528d59edb3a5e23c16f

Observation 727ce860-a44f-441f-8fed-e0382e7af1ba · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Sigmoid loss for language image pre-training, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.049608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.049608Z digest=sha256:a0c0f7c951e1dbba5ff08439294d6a99529a6e61277c6b68fb2872162efa6419

Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.167998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.167998Z digest=sha256:b17749a0249e02ab41aaddb8c42f36b2ebad9cbc1e77b53ea8c39e66c8fce6bb

Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.329669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.329669Z digest=sha256:5f081030457ffadd9f4ff16a20cc8f5a4b00aea3b0807971e27773346dd1a4ec

Observation 3d866024-da5b-4073-b695-8bf8b0b3a9ae · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.447271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.447271Z digest=sha256:66b74c36977a80b30def3a8e564d2c0badba6362d24f28616084b802738a62bf

Observation 80f41d3b-5c2a-4624-901b-1e90b1c58f28 · outbound

This paper cites MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.538261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.538261Z digest=sha256:83ba3c0d88731378f62dd18ca646c850eef319377b60f2d08ddd66154b239068

Observation ebd9d7f2-d8c0-4b09-95a8-8012cb60fa00 · outbound

This paper cites Swift: a scalable lightweight infrastructure for fine-tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Swift: a scalable lightweight infrastructure for fine-tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:35.879362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:34.646887Z digest=sha256:9aa4a19022a900336b6d5f976ac41e28b24dcaecb99fc0910d37dbc7300de8ea

Observation d7e63517-4296-4df2-99a0-37102c2a6997 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.833805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.833805Z digest=sha256:8001c595975e3656591118e08c611d9e94ee1992be62add35bc6eff9c180546d

Observation fce9c156-8fc7-47df-8f60-e76731845313 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.911797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.911797Z digest=sha256:3aa70daa79a51d4e8823b2ab7c21fe4c6862607cbecbe80d49fdddbee3a81bd0

Observation 7a882fd8-dd3a-4dfc-ba9a-baa3e1f05b5f · outbound

This paper cites $” or measurement units such as “cm.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models $” or measurement units such as “cm

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:35.630405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:48:35.027792Z digest=sha256:213fe42cf2c24f12568caf19688481c22896aabb0cf0413a3a96fada7e57dc26

Pith citing papers

Observation e7ab0eda-b5a6-4a65-9a70-0eda9239920e · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.034155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f30846eb28c3194cbcfe07e343df3a9a23b21e525c5aa905a9654e3603abc4b8

Observation 112900c6-95c3-433a-b8c0-8835c23ea18c · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:52.439397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:c8677f41f66126a9fec27133dbdb7062ccb5541492b9429bb78e95d1228d8124

Observation 0fe9215b-8e3c-4eae-9b83-3f1e8e898893 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.450399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:4670269af6f6f4d089edd6c50043d76bd093c42d289c3eefe2f609b6ae5137bc