Pith. sign in

Paper Citation Record · LEDGER

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2411.14725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14725 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:03:24.836356Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29f63658-53bf-4aba-9f23-ce43457dc585 · outbound

This paper cites GPT-4 Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.770624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.770624Z digest=sha256:aa84134ac5be5920957c4ff8ebb2a7ed0e7d9ea950ce2077d8a14cd2c06cb802

Observation 17e6553f-9f94-4f44-aac6-1cd611c55cae · outbound

This paper cites UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.786991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.786991Z digest=sha256:a373d6c362ac50c8c35ccb017f9875512d20ad3a97b08a44ac3344c6d8c21994

Observation 4b7f9521-41e0-4bed-bb9c-4d07204918cc · outbound

This paper cites Claude 3.5 sonnet.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.104020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:23.818926Z digest=sha256:d2201f335183fce75b260523a3b18d42c5330d42426ba8f9ac01ebba9ad6669d

Observation 4dd63f1c-c7d0-4a21-b8ec-672b916c5bad · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Non-Determinism of "Deterministic" LLM Settings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.837472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.837472Z digest=sha256:c5695005a29e1ccbea9aa31f66bf859c36a9fa41d575106641d7858d55782e45

Observation 010a7473-fc9e-4e04-957e-d292964a54b4 · outbound

This paper cites Qwen Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.871962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.871962Z digest=sha256:4d14839497037ec4682cf4ff4fd53c428a69cd46eca0aecf46da2e6d1cefc20b

Observation f6e0c276-55a1-40d6-a0cf-39bbb76421d6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.904276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.904276Z digest=sha256:e372435f56fcc8eb0eb2fad63b88df4a97370ab175f48dd77fb49fb97dd81950

Observation 592d8c6c-0a72-40e8-b66a-d1f9e90797b0 · outbound

This paper cites Eureka: Evaluating and Understanding Large Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Eureka: Evaluating and Understanding Large Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.957999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.957999Z digest=sha256:4dad1e21801b062a9080f133a44633b2f74b774f9921567845c114079b797a67

Observation d75277d9-30a9-4b51-958b-d3ff16e000d1 · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.965262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.965262Z digest=sha256:49b7d39e320608e5216673a4ab34eced46f67331f432c7eee5d41ba19d69f9d6

Observation 9e811f0f-c327-4f70-9912-7d228f702d79 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.972836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.972836Z digest=sha256:ecde85ef936516c1abc6c7e3eb28a8af38ef36615377af283e096cff0998ba17

Observation 21a2f1b4-1b4d-4656-9d21-9cc70730fd1e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.978719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.978719Z digest=sha256:9696e7247af372c3baeaaa03cd711e35de568e5274f6e2ca631adb778a060c55

Observation f054d048-c168-4083-8c5b-19c9e477b4dd · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens NVLM: Open Frontier-Class Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.984309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.984309Z digest=sha256:163d7909d9b8c15324b112dcfc647817556a5f86a48ee3fc5d7159ed84815003

Observation bd75654c-1b80-4804-bf6f-7f17fc094c34 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.989904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.989904Z digest=sha256:119ac1dbac63208f13528d406f3b2fdfa1a9175d43fd38af6558b05721334425

Observation 142f118c-2802-4fee-89e7-c344e8f7350e · outbound

This paper cites GPT-4o System Card.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:23.995882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:23.995882Z digest=sha256:cf3f670024e95e7a1a809059a89fc9e084a41120cd2f14e0fdb8a0626a76cf63

Observation 8958e039-5a48-4600-8a83-0d73af34ad82 · outbound

This paper cites Editing Models with Task Arithmetic.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Editing Models with Task Arithmetic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.002580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.002580Z digest=sha256:118e7f0d13ac21ca26fe6444ed1ab7404a8047eb78d42d1c486c5662972d85da

Observation 8020696b-2d88-4235-9430-16c73e0b7b33 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.010418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.010418Z digest=sha256:91fa8188b52e7fa62deaeb1ccccfc7476f964cd3d667622b121373081b0506f7

Observation 53f61553-04d1-475e-b527-02875aa567fa · outbound

This paper cites Kembhavi, M.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Kembhavi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:27.081220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.020558Z digest=sha256:0a6f850cf728b45cd75ce21e05990ffc7c73822ff1dbbc9c6aa05d1c22f51132

Observation e10dfaac-e4cb-4126-8a3a-aac64e9c0850 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.981473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.025849Z digest=sha256:f26aa50129ce561fd6829c762dd8c36e70ea40a1d268e7e9253de6f443e862c2

Observation 8d2e9f02-2eda-4165-9a63-59487947c66c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.956819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.031716Z digest=sha256:7c1aa5c5647ac49f3e251b9d7d1bbcc6fccd479b01aaab6e66c0a095bfccb5c3

Observation e8c890f3-d616-4444-b28c-1e4c30500148 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.037367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.037367Z digest=sha256:aeea27f2b0f6e0e2ab47bf1c00f4da80c09c3d8304ed9fc76deed5e2263893b1

Observation 021cca4d-b97f-457a-94cc-63aa3825ecdc · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.043058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.043058Z digest=sha256:0aebcbda4c59be47763b0f9e67ebe0986fc2f0430b6f4e35b72734572641f06a

Observation 6b522d49-fecd-4b49-a8aa-e01599866372 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.926888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.061742Z digest=sha256:8c3420fe672d8841333f5de72b9fba7ab3162ea09851585860110279444e157c

Observation 5e19d423-3489-49aa-bfb3-53f14d09e3f6 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.077979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.077979Z digest=sha256:d785dcef7237f1250a92c4d6a7b1a9d9c8f8ffd3d320b7f222df18272f14970c

Observation 2db77923-8e89-4ab5-b815-4b1b233f7607 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.105463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.105463Z digest=sha256:c0e7731b3a751bb99c88bc2e38068576d4eef7735a8c052eba5df3c980a6f2dc

Observation 263e4985-f8a3-42ad-92fc-1f79f61dd4ae · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.886416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.129970Z digest=sha256:bfbe9bbce1f43ffb1bbc01b32c345f4b8a77b14c09af31c8aaa456092a4bc4b6

Observation 0edbedc9-203e-4399-ba68-21537ae5f149 · outbound

This paper cites POINTS: Improving Your Vision-language Model with Affordable Strategies.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens POINTS: Improving Your Vision-language Model with Affordable Strategies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.193907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.193907Z digest=sha256:e6cea9055157e58bd49905b313bb8181dbcebe241076bb15a7119bd134ae465d

Observation 347de614-63f3-4d72-a2fb-a01538441053 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.261105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.261105Z digest=sha256:11d41867a5d793e59163c8ffe28ec7806ae39be43ae2a4784abd54d2cf8e0b1d

Observation 2599b4e7-255d-4427-bf93-cc868e3785a3 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.301248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.301248Z digest=sha256:9989ed28aac534ea89342358f7398ad2a3bc6261506a14fad47443a9ba30be20

Observation 0cec2fa3-9d64-4e00-bbf2-9a0d6e90c994 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.319323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.319323Z digest=sha256:3de2e040e1611cb01070c8f9507f80e33af94e6ea612cde291a454bcf05838ef

Observation b7f4f425-487d-4f19-9e97-541f6085ab88 · outbound

This paper cites Mathew, D.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Mathew, D

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.868949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.324376Z digest=sha256:292af0efa4bc142ee5cfa3c48d9d225067c73da7ab204a7afb6dec4578a6c2fa

Observation bc9917e4-889c-470f-abda-c12dc38b9e3c · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.329453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.329453Z digest=sha256:88baa516cdfad08e39d62b78ad2e850c38cf74960c36555083d5d45d36684e40

Observation 68d6272c-3291-41b4-a9df-6fe3b227fcde · outbound

This paper cites Gptv system card, 2024.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Gptv system card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.847137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.340481Z digest=sha256:e032d16e35af3ecd211b8686f3af267b264f194e4bdf621de5f27a1503506dd5

Observation 0765cc2b-5eb3-474d-a2d0-b26d576bdb56 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens DINOv2: Learning Robust Visual Features without Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.345745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.345745Z digest=sha256:57e3d87f7d9537b798ec3338448b2055a7bb655ef445cd4ddb010fd71596597d

Observation 1bc1e516-d884-49a6-a17c-5ee536a18e6c · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.351317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.351317Z digest=sha256:62e5175c2a1f8e6e3fde22cc0096a4aac73b1d846e70572de69e9a2e47a7a9fb

Observation e991f5d8-d7e9-4a40-9872-834a9dc98c8d · outbound

This paper cites Radford, J.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Radford, J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.356784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.356784Z digest=sha256:f940714c97aea53abd2dbc4c56bf4461a6f23246ffa0590be50fba714a83de71

Observation ceee076c-99d7-4175-8225-be5f4adbde4a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.362312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.362312Z digest=sha256:7ac09e22bc73c908462b91b77a0841dd1a707ebf7ce234e774675d0fa0685abd

Observation caf8f6ce-7441-477c-8a97-eac1e07e505c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.367607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.367607Z digest=sha256:f5332fdef696ec2c19e6858f1d5425a01ee6ddbb289a83680f0681cdf6fcf648

Observation 00d683b2-4ffb-4c06-b471-c4b5326f1ec2 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.373327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.373327Z digest=sha256:878ddbc5a6095704e9e41047253dc2a8fd6171976ae61bd2625603bc612fc52b

Observation a72697f6-6066-407f-bd7f-d8b07573fb79 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.378871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.378871Z digest=sha256:d60c126b85e161046e2f99d7a40baed255e1338a5a61f69f968312bf6534036c

Observation 224732bf-b199-4101-8915-c6a23b8bba40 · outbound

This paper cites Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:03:25.023717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.384210Z digest=sha256:8b1e278c4c923b1dfd14a6592cd44a254fd6fcd2de2a5b5340122bd921623621

Observation 1c421364-0ecb-4800-aa2c-120399cb8363 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.389774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.389774Z digest=sha256:b4b7cd7cec9c120287fbb12a4186885452596646f415a26539d58bcc0f65bfd6

Observation 4a1c64b2-93a3-4586-846b-58f0c01c7b66 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.679958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.395596Z digest=sha256:d79251b203eb5bcf7852a0599d8e07a1ed3814af1a246c672127c0e9decb67e3

Observation 61c2b739-a403-4914-a3cd-718564c4f772 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:26.618098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.415999Z digest=sha256:6ae9eb2282792b140a8fb0c3c5216fe1a4f6355d68b0084bb8eb5012ff9f291d

Observation 624ae08f-463a-4c14-b7e6-c1c29655e177 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.449765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.449765Z digest=sha256:5b6f22e5d7e1e65dc811e52598f47c8f8332d8e55beb79f4961249dd8e145b16

Observation 07c16031-c15b-45eb-8a98-14b19ba41a23 · outbound

This paper cites Unveiling the Tapestry of Consistency in Large Vision-Language Models.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unveiling the Tapestry of Consistency in Large Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.493884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.493884Z digest=sha256:f9bafc9c2dba79fd56118206e124f5e384a45c3e23700b8266077ca3b27675f2

Observation d1ca39bc-2e83-4c11-a2d1-f441b52e5a69 · outbound

This paper cites Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.558647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.558647Z digest=sha256:8356507bcd6bf885ee565f6199a8e2e2662a52700bb0a2cb92b1e9fa97b390a0

Observation 45ff90b9-7d09-4f1d-8fd9-6c536de9e14b · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.598812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.636112Z digest=sha256:8a3ac93953fd38314c578a53e78f272cd11f8fcbaeb89308c7cdbf2d866023e6

Observation f7916f18-88b2-4c1a-bc82-43740aec36d6 · outbound

This paper cites Limitations.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Limitations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.526008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.658335Z digest=sha256:b6368c93f1108d5bd1fb7e5e71230038b979f151995ee336f36d7bd3439965dd

Observation e12fbe19-6997-4f7d-909e-b17297f5c9cf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.402925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.684847Z digest=sha256:cd6314c9f37041ad59125549d93ce3de4b21e1085471c2f1dfcc7f3fa64a8d29

Observation 20bebfd8-650d-49d4-ac7d-ae17e22fbe7f · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.384883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.703238Z digest=sha256:1f8ffd7a3819532a3e38d17c5caaa1022304ede0d834625b6780fe5aec1a1e26

Observation 64bf1c8a-3c2f-4701-b9c2-65e916ba274a · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.367041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.722927Z digest=sha256:534de342c0aac5220cbb744679b20f2995d9c83391b7d71f0a7db64675335d25

Observation f41e0017-36f1-4589-a6d3-7c23cf26867a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.347802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.742266Z digest=sha256:f706381ac508e33c438758089e74e130b514f1bfa8513c1ed36d647a5cf6e797

Observation d627cf61-1a2a-44b5-b3d9-84ef68722d59 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.201497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.774916Z digest=sha256:2b498c179da48e1d51472e25ca2743072a7cb5de1dffb516b9a83b68c66c8636

Observation 908586c6-847a-490b-974e-25670ae78301 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not include experiments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.130394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.781282Z digest=sha256:36179f4dfb352783df808b52644fcf9ad54ca1faa932e646b89dc45898d87baf

Observation d8408868-447e-4c1e-9f17-52ab8c050c08 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:26.104548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.788121Z digest=sha256:58e5e9fc8639db206250716be894375451c4803737ec477246ec06100e3940fa

Observation 359ef756-7ca3-479b-b8e2-0173901c1f08 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.979567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.793680Z digest=sha256:c9bccf4f9dc5a741be2d4aa2a36923c07bc1cdab211135bf49843f2f064abeca

Observation 9b67f61e-c732-45a4-8ca9-4e1c04f7cc21 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper poses no such risks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.879257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.800239Z digest=sha256:9f7ca3510190f30300938ba64c48e820648bdfd93f951539ffe37d696c40bd26

Observation 203b5458-eb86-4803-895f-f06ba7791d03 · outbound

This paper cites We will elaborate on it on the Appendix.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens We will elaborate on it on the Appendix

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.856187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.806116Z digest=sha256:92c634079c484b950577e5fb7e20a7cebf86c7fe4da18d36ea5d30bb05ad91eb

Observation a65abba1-5764-4268-b798-b0d52d976d09 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not release new assets

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.835417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.812627Z digest=sha256:0dc692d885b8c97e35bcfe4bb0ea41ab91ac87082a5b37de6091bf52d9a09732

Observation f4685861-48c2-48e5-87dd-eb3e2c22db39 · outbound

This paper cites an unresolved cited work.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:03:25.780005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.821604Z digest=sha256:6cef08ad4f12231bbf3961cd4a1fe2c0f6fb81a6e3d67416803f00ecc6ef90d2

Observation 634334d2-a7d1-4e3d-9c9c-61e4764e8d8b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.665905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.829140Z digest=sha256:d3c699d13bed5d09ea5533ec533a192d62f5fba2ad066bf9f7a4306dc4ac66d4

Observation b6332e21-180d-40d5-93e2-f7353a5f43ba · outbound

This paper cites Answer: [NA] Justification: We use LLM for paper writing.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Answer: [NA] Justification: We use LLM for paper writing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:03:25.647440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:03:24.836356Z digest=sha256:dd29ba359de972dd1938286e4756cb709c11dc4fca33074b9475c1f2bf5f3299

Pith citing papers

No inbound Pith citation observations are available.