Pith. sign in

Paper Citation Record · LEDGER

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.14725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14725 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:03:24.836356Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29f63658-53bf-4aba-9f23-ce43457dc585 · outbound

This paper cites GPT-4 Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.770624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.770624Z digest=sha256:c5ee3d4979eb99ca2b4646f5afe110aea6fbe16a1f19364c9f1c2e5f8adc9a17

Observation 17e6553f-9f94-4f44-aac6-1cd611c55cae · outbound

This paper cites UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.786991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.786991Z digest=sha256:071690f538aeecfa9174f98084dc079b44f845a87c4a5cc7512e24d84df2b2cd

Observation 4b7f9521-41e0-4bed-bb9c-4d07204918cc · outbound

This paper cites Claude 3.5 sonnet.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.104020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:23.818926Z digest=sha256:9c7bbf9924b34a0e5be811d3152936c3ff6312dbd31d628ec43948a3c7b7fdcb

Observation 4dd63f1c-c7d0-4a21-b8ec-672b916c5bad · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Non-Determinism of "Deterministic" LLM Settings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.837472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.837472Z digest=sha256:8b3b6f37563f2122bebe5fbcb1c8481eb8686b1187ff4d5d2eeaf08c78d0ea6b

Observation 010a7473-fc9e-4e04-957e-d292964a54b4 · outbound

This paper cites Qwen Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.871962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.871962Z digest=sha256:f4d05733ac1d861ba6035c7ce34aa09594a34649574ac29729a3ae81db2070a8

Observation f6e0c276-55a1-40d6-a0cf-39bbb76421d6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.904276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.904276Z digest=sha256:3016e0da12455338de87e180c9eebc8caceca2c47d72d6c0938c1f52d91d8bd2

Observation 592d8c6c-0a72-40e8-b66a-d1f9e90797b0 · outbound

This paper cites Eureka: Evaluating and Understanding Large Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Eureka: Evaluating and Understanding Large Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.957999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.957999Z digest=sha256:ed3a839c8da7e3538a814e2632a5e014b2390dac1a9964e32664551030169032

Observation d75277d9-30a9-4b51-958b-d3ff16e000d1 · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.965262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.965262Z digest=sha256:c30aafdd0768420d1783ff8ce8f92b914d867a449aa012373730d4ea62104a74

Observation 9e811f0f-c327-4f70-9912-7d228f702d79 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.972836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.972836Z digest=sha256:6cd55b3a035d62319515d95e8796a4843eb84b392362cf27712074ed70fbc8c8

Observation 21a2f1b4-1b4d-4656-9d21-9cc70730fd1e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.978719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.978719Z digest=sha256:17ddec5c4b0b026c334d96e0850b577da97d99a593e403f89f8bba6d05f5eb70

Observation f054d048-c168-4083-8c5b-19c9e477b4dd · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens NVLM: Open Frontier-Class Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.984309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.984309Z digest=sha256:9dabf60ede33bc130d6a1a2ca01035eb0d4d7f93d22ccb1054617a4eab1e4bfc

Observation bd75654c-1b80-4804-bf6f-7f17fc094c34 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.989904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.989904Z digest=sha256:7e410098a15f77defba17639f05ef8a6940c3d8fe9704db5960a653e8bb09ff4

Observation 142f118c-2802-4fee-89e7-c344e8f7350e · outbound

This paper cites GPT-4o System Card.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.995882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.995882Z digest=sha256:b5cc80c36c14ddb95a57b3f7871e774f12cfc825ac96dcf25cb2e71937762cd9

Observation 8958e039-5a48-4600-8a83-0d73af34ad82 · outbound

This paper cites Editing Models with Task Arithmetic.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Editing Models with Task Arithmetic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.002580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.002580Z digest=sha256:058fb1b3d0567a246d959b7389533736da4b4c1a9a78df979d188f74605b79af

Observation 8020696b-2d88-4235-9430-16c73e0b7b33 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.010418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.010418Z digest=sha256:2ab049583690e579ea252e81ef115b7aec06a535d357ce342207ade9d5c28716

Observation 53f61553-04d1-475e-b527-02875aa567fa · outbound

This paper cites Kembhavi, M.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Kembhavi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.081220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.020558Z digest=sha256:89133de607b40b6e96eb8f26acf129dd36d3c527a1b29de35ac62827a4a01ed6

Observation e10dfaac-e4cb-4126-8a3a-aac64e9c0850 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.981473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.025849Z digest=sha256:37f6eee21573458a2e0934e4cca59a31d2a5d125465f4d8d279c6fb2fc044196

Observation 8d2e9f02-2eda-4165-9a63-59487947c66c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.956819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.031716Z digest=sha256:b023ab003e79ebf0ec5e8460d923da6593f18db5ebc9d5cd4c33257e7fc620b5

Observation e8c890f3-d616-4444-b28c-1e4c30500148 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.037367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.037367Z digest=sha256:c1b3bf6209e761005f227cad0bc6d4f2c0b77d5439887512b37b561eda0a0d48

Observation 021cca4d-b97f-457a-94cc-63aa3825ecdc · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.043058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.043058Z digest=sha256:2a1e5909f24da3658c44dc4038b6d20efb273c1aee98ff652ecd9ddc06ac99ac

Observation 6b522d49-fecd-4b49-a8aa-e01599866372 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.926888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.061742Z digest=sha256:c1ad155adcb624f872ccef84e7b5d487d39a01071e07ce3f1a06dc87dfacc4ee

Observation 5e19d423-3489-49aa-bfb3-53f14d09e3f6 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.077979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.077979Z digest=sha256:88fcd9efe4842ad798002269a7d4bb92fc90f70eee92007aed61ea0a00a437d7

Observation 2db77923-8e89-4ab5-b815-4b1b233f7607 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.105463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.105463Z digest=sha256:9509c0ed2214e83820d8c34a94c2053115aa547401e97d35dbf229e9b19ee55e

Observation 263e4985-f8a3-42ad-92fc-1f79f61dd4ae · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.886416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.129970Z digest=sha256:7f7187dc3130f7d304a30c5a02dbb50bfff422959e0003f370674039c9b1c978

Observation 0edbedc9-203e-4399-ba68-21537ae5f149 · outbound

This paper cites POINTS: Improving Your Vision-language Model with Affordable Strategies.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens POINTS: Improving Your Vision-language Model with Affordable Strategies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.193907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.193907Z digest=sha256:40f7a089340be5b0a2f73a04b086249e7acf3712b12f32e38f430b0e8afb3b9e

Observation 347de614-63f3-4d72-a2fb-a01538441053 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.261105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.261105Z digest=sha256:1fada224e062e7ac3529a8e3cce32eaa5111a5bda9038c0d9450fbd2a334f8a6

Observation 2599b4e7-255d-4427-bf93-cc868e3785a3 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.301248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.301248Z digest=sha256:6a576539ba060b0e8cd06aa97f3cfdd7069462c7f89e1a492fdef6da8c92239e

Observation 0cec2fa3-9d64-4e00-bbf2-9a0d6e90c994 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.319323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.319323Z digest=sha256:4584fa6adad220f0c391e3c36ae1ecf2b123835a9c2e37039aec4aa845540783

Observation b7f4f425-487d-4f19-9e97-541f6085ab88 · outbound

This paper cites Mathew, D.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Mathew, D

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.868949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.324376Z digest=sha256:94456ac1637a0b6ea78ae3073815d36aa8a4eca58b4f607770d4568d1f7be52b

Observation bc9917e4-889c-470f-abda-c12dc38b9e3c · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.329453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.329453Z digest=sha256:74cf4635e06e05a4a7c468146ee0b2f105bd0346d64c7543e90d8fbd365409f3

Observation 68d6272c-3291-41b4-a9df-6fe3b227fcde · outbound

This paper cites Gptv system card, 2024.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Gptv system card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.847137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.340481Z digest=sha256:b7223f61e33d9cd1b0989319a307b5b0c23cc866d8425cd0929467988d8f32ad

Observation 0765cc2b-5eb3-474d-a2d0-b26d576bdb56 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.345745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.345745Z digest=sha256:b10071cedbcaaca99bc06a4e64659bd5a8d8b3dcedcee7777ee2daad674d4ba2

Observation 1bc1e516-d884-49a6-a17c-5ee536a18e6c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.351317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.351317Z digest=sha256:0b4e728c35e2a460d32fdc7c513e0d23dc0c899761994ac84a6ce29404a704c9

Observation e991f5d8-d7e9-4a40-9872-834a9dc98c8d · outbound

This paper cites Radford, J.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Radford, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.356784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.356784Z digest=sha256:2440cb18021773d5b6fc6e5e04d2020256bb444f84f37d61f2af6c171c085979

Observation ceee076c-99d7-4175-8225-be5f4adbde4a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.362312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.362312Z digest=sha256:dad481fbafeb8d8b17b5924185c6e28c1ef16829988ec9bc2e6bd456c6e4902e

Observation caf8f6ce-7441-477c-8a97-eac1e07e505c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.367607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.367607Z digest=sha256:61ed8938161ae7e2ddd0ca17c8cbae72526740e97f9a35be1892b2d21b29c1fc

Observation 00d683b2-4ffb-4c06-b471-c4b5326f1ec2 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.373327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.373327Z digest=sha256:9873a6eeb5f8d33de96de44813cebaa4b3dd7780414e2214761ebafa629fa847

Observation a72697f6-6066-407f-bd7f-d8b07573fb79 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.378871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.378871Z digest=sha256:80c8faba7dad28b35173cf63a5dea473f4a5243e43c94df918a345992973ef93

Observation 224732bf-b199-4101-8915-c6a23b8bba40 · outbound

This paper cites Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:03:25.023717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.384210Z digest=sha256:6f9183f4c17112caf89dd93523ac83d67d17d00a5ac1cfe407ce768629d0438f

Observation 1c421364-0ecb-4800-aa2c-120399cb8363 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.389774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.389774Z digest=sha256:148c9af78329aff958978efa5dcfca4b0d4c70f7a4dba553657fafacfb40cdff

Observation 4a1c64b2-93a3-4586-846b-58f0c01c7b66 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.679958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.395596Z digest=sha256:16844786e149aea386ea71e415ce7f34a914118d59aaffb914352e59bbafe25a

Observation 61c2b739-a403-4914-a3cd-718564c4f772 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.618098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.415999Z digest=sha256:ca000b7988569bc28d97f54f3a4ffd48f6b6c3f1e2123196c96bb23fef702fd2

Observation 624ae08f-463a-4c14-b7e6-c1c29655e177 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.449765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.449765Z digest=sha256:9f1f962d49c37809d75f2d537710d817e022e3ac421492e0944a92920a836202

Observation 07c16031-c15b-45eb-8a98-14b19ba41a23 · outbound

This paper cites Unveiling the Tapestry of Consistency in Large Vision-Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unveiling the Tapestry of Consistency in Large Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.493884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.493884Z digest=sha256:d84e4ed712d40bcdc2e734896c772db06ef2c8be8923d584f0c1766d83773dce

Observation d1ca39bc-2e83-4c11-a2d1-f441b52e5a69 · outbound

This paper cites Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.558647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.558647Z digest=sha256:9630131af070dab40e615e6a808e860f5f5f4c6228e5d9b4ef1f523d77642115

Observation 45ff90b9-7d09-4f1d-8fd9-6c536de9e14b · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.598812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.636112Z digest=sha256:4b2eb0d66f737392fd20a2b5afed26bff4af52affa2ec6a792a221e21556cc4b

Observation f7916f18-88b2-4c1a-bc82-43740aec36d6 · outbound

This paper cites Limitations.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Limitations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.526008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.658335Z digest=sha256:da42d76077d86ae559710ec7f944457f9dcccfcc5c80b90609e8fc589492b85e

Observation e12fbe19-6997-4f7d-909e-b17297f5c9cf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.402925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.684847Z digest=sha256:531077c09827b3016079d24a62dc9063ff4675727f78e414b3ca516fe0e9c1e6

Observation 20bebfd8-650d-49d4-ac7d-ae17e22fbe7f · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.384883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.703238Z digest=sha256:459f0f17d3217e7c5af1c19452c0b02c88df8f72ec25586ad24331d8d4f8898f

Observation 64bf1c8a-3c2f-4701-b9c2-65e916ba274a · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.367041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.722927Z digest=sha256:59743688c717be941496c6d98e8f6efb1073118b3c160bca2d88f935e5ed485e

Observation f41e0017-36f1-4589-a6d3-7c23cf26867a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.347802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.742266Z digest=sha256:92f4177647bc3d3883eacc074c88af43baa0e7d6839120697a86ea380b3ebc72

Observation d627cf61-1a2a-44b5-b3d9-84ef68722d59 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.201497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.774916Z digest=sha256:3c8f48f1fac7654603b7860168bcb2244f493835df378798d3711dde64c52643

Observation 908586c6-847a-490b-974e-25670ae78301 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.130394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.781282Z digest=sha256:24d6a2ec0044e306cd182e324f1f6e7991cd189e8aff12425edf85c3ba51fee6

Observation d8408868-447e-4c1e-9f17-52ab8c050c08 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.104548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.788121Z digest=sha256:3c04631b68be283abf0c6a9552260c9691923b9da2852674067bb349e03d8b9a

Observation 359ef756-7ca3-479b-b8e2-0173901c1f08 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.979567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.793680Z digest=sha256:3c0fd74ddf478735134c1cf065adce4fdc1f73703560b1214688d07473e57c1d

Observation 9b67f61e-c732-45a4-8ca9-4e1c04f7cc21 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper poses no such risks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.879257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.800239Z digest=sha256:4107a73a32f37750e03dd875364752b641fdd366eaf7c564a8ef1c9d775605d1

Observation 203b5458-eb86-4803-895f-f06ba7791d03 · outbound

This paper cites We will elaborate on it on the Appendix.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens We will elaborate on it on the Appendix

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.856187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.806116Z digest=sha256:a34eae26f0118a1777f573c7249a270e998fe3fb8d8021a029d9abb90d540a1f

Observation a65abba1-5764-4268-b798-b0d52d976d09 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not release new assets

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.835417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.812627Z digest=sha256:ead26cff23192d34181dbcbc7653645725da4b2b213fc729b59db0c0f465882c

Observation f4685861-48c2-48e5-87dd-eb3e2c22db39 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:25.780005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.821604Z digest=sha256:b25a43166351d2f2240b36b52326bf3215c192ec08906a4dbca8110db6f1c7d2

Observation 634334d2-a7d1-4e3d-9c9c-61e4764e8d8b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.665905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.829140Z digest=sha256:ac649e6262686cea94621fd39504bf7aa2373c7242ecb0581450606f0c2bfb23

Observation b6332e21-180d-40d5-93e2-f7353a5f43ba · outbound

This paper cites Answer: [NA] Justification: We use LLM for paper writing.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Answer: [NA] Justification: We use LLM for paper writing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.647440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:03:24.836356Z digest=sha256:2332fc84d1dd7ea2599e708805cc09e84147d4c7064a8fc0233c5616a43b1639

Pith citing papers

No inbound Pith citation observations are available.