Pith. sign in

Paper Citation Record · LEDGER

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.08468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08468 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:58:06.130584Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:45.290988Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:02:48.057736Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 911d7246-3d83-439d-bbed-4e24168cf271 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.923359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.923359Z digest=sha256:9fd054c94ef8f44330654fc4c310df7662a6ed7d2e6d16d2426129262c596d1f

Observation b5310a3c-1ac4-44d4-b42c-e48ce9ee47be · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.928089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.928089Z digest=sha256:ba1690aae0ffdb57564a4ab244facee5f1e93b141c7789bd3415ae406d1baba8

Observation d7ac8ac5-2587-44d8-9baf-4058724cd976 · outbound

This paper cites Qwen2.5-VL Technical Report.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.933486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.933486Z digest=sha256:609f56f99ee5d6f4e62f2116f7618691d803fbd0d504f27a840e354b384a8065

Observation 69f8eb1b-f8fa-4510-834f-20c747f4c4a8 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.937884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.937884Z digest=sha256:02b0764af42b00cc8b7fe95b417885892097c6d62591a47821d3d4960cf6ecfe

Observation 882e82a6-4310-4bbd-a09b-6de2c6f1d089 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:07.020503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:05.942429Z digest=sha256:8a2afef9076f143fb9d8220ebb41fad2fcb515cbac91dfebffe36122627516e4

Observation d728a538-1e8a-4dd6-84bd-1e43c385483c · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.946807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.946807Z digest=sha256:22dad3e6fb5b3eb22a1a72c8633d09ead4768c66a6fb5632c2338375ce751407

Observation 91ea524b-07b0-4780-a018-f30c03f490f2 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.951451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.951451Z digest=sha256:91d40220011a6467b4accbc966041c224fda304d01b4a81671dc9968082e8e1c

Observation be86b068-3cbe-4d36-8fce-ceaf4b3abb3c · outbound

This paper cites A Survey on LLM-as-a-Judge.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? A Survey on LLM-as-a-Judge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.955548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.955548Z digest=sha256:0e271ea299031f0a60507365e5bcbc1060cdb927fc4c5851d11debf014c61a43

Observation 21fd0795-0030-4042-95fd-8b3ba8e44139 · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.959951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.959951Z digest=sha256:09eb1168819b4aa9bb9e921d6dd5bf76c4c448793126402bbfa59ca734506cd5

Observation 7cdcd40c-403d-4b0a-98c6-57d4ce1f8550 · outbound

This paper cites Hoque and M.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Hoque and M

Reference 10

Resolution
verified exact
doi, observed 2026-08-15T21:58:06.247095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:05.964150Z digest=sha256:c00f407ba9ed22bbc77e62bba59891f1ad16193cf557868f3de1d408189e6fdb

Observation ffb56156-cb08-417c-85c7-bf991837908b · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:07.006758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:05.969576Z digest=sha256:e19b94adc792b50972b85112b14cd6cdb9785a0dc87e3fc31ffac5f47d3634a0

Observation 44c7f213-3f1e-4290-b67c-a900d10dcc67 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.973425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.973425Z digest=sha256:9e40911cef440ee83a8292a2c378d5e035a1b27fa2a66fd4580b92a419117064

Observation 3a36e0e2-ad8a-4ae4-95b6-fa5e1d0950b5 · outbound

This paper cites Mistral 7B.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.977367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.977367Z digest=sha256:b0f2e02e60c95bfe9f8a2b991f9d1a7ee7b628c6644da3de797438d10c4fd0be

Observation b153f563-57fa-4131-b90b-41326f814989 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.981258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.981258Z digest=sha256:c5929abef98df31495a9114f6d72617071485c735cb88eef4528a47d48509dfa

Observation 80b714d6-bf15-4dde-91fe-d98e1739be0f · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.985004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.985004Z digest=sha256:3f35331d3cebd58d1d051698dc2d9b039129adba724ab0021cff3d727a67dd1c

Observation c87e1fe9-e02b-449f-bbe2-52655b0fa87c · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.992078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:05.989042Z digest=sha256:d610a3957098c170017a2a4b1065339b16d1b56d7993a0ec0ca0918fd0487de0

Observation edf98761-dc47-4ea9-a969-a1ec0f721605 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.992942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.992942Z digest=sha256:f5e1695c6185d522f3e38720cd258b4d261a02b1e3922156623cd7b1758fd27b

Observation 9c78bcd9-93ab-4a71-8524-85ed308c8067 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:05.997220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:05.997220Z digest=sha256:99f0450e84c685bda771c9740eee192540a947bdfda84f8dc79f5a14510890c8

Observation b32b164e-97b8-4be4-9a96-83d50a814fb7 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.002215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.002215Z digest=sha256:2775eb29101e3dc0f4c089445d86df8b867f493b368513e7fdfa7e80a1257b50

Observation 01b8dcde-723f-4ef2-b8dd-fe2312499b51 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.969819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.006772Z digest=sha256:bcb6e440f59e204d1e6e765164330476a5123c48b8117ad9d02ff7b6695be474

Observation df86cadb-78b7-4efe-90e4-e53eb8598e5a · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.010932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.010932Z digest=sha256:1b4eb4e267adcb7fa9045c3340b61f639dee18c40424f0f38dcceaa00d3e5f80

Observation b1e2fc81-1c6f-41c5-a5a2-4c78e6d2336a · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.015449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.015449Z digest=sha256:a161c440fc5484e3f8a46321785eb3cc7d37d626aa6b85f68e9d3085daa42ca8

Observation 24891392-0bbb-4fa0-beb1-b2f55301937d · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.020334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.020334Z digest=sha256:c8f27de20d4ee8ee14ddcca2f51de4a94688a987a835567fdaf83d7af36de843

Observation 0f50d615-9781-47bb-a4d8-fce889e197a4 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.025627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.025627Z digest=sha256:c5f96c679dd4d877c9d3dc28655f01a7dcc7b72f044537f1e0e60bfb31438f75

Observation 72fefe5d-a095-496f-af1e-9c1b2ca1c032 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.954660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.029528Z digest=sha256:f62c39fc6c493981cc93449950b777b2c93a839e1ff7dd0e22761d65f4b51b0e

Observation 436205ab-49df-48ac-89c6-7ca65dd5cd4d · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.936030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.033641Z digest=sha256:0be070e70c39808fdd81409aad28b3df7163906c2faa93f4b63d677cf9fd0c7b

Observation 3bd1e347-76a4-45f8-aecf-2390577300df · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.908848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.037822Z digest=sha256:e15b7ead2983d06e568ed574e33fc7665d9868d62da9a204065282ee2ce9b0cd

Observation d1bc2630-7b9e-439b-b684-6a22cc082512 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.884765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.043558Z digest=sha256:920f449931f1504596628e10dd6a8dffa46e32da900820d1567ef46fe2c21aa7

Observation 27859357-9ebd-4609-b14b-ccd9e942291a · outbound

This paper cites GPT-4 Technical Report.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? GPT-4 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.048106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.048106Z digest=sha256:9f9ea56cc94c09dc7ac0f736642cae9bbccab8baa2b97848fdaad918f4ae3fd1

Observation 860e3d6c-1ab6-46c8-b985-3922da4d0dc1 · outbound

This paper cites Tahmid Rahman Laskar, Md.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Tahmid Rahman Laskar, Md

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.052106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.052106Z digest=sha256:9dc53a92ec01478ea37e30787c6abd1db13db2b01bec229343fe73a2b15f2844

Observation 8507f25a-4403-4ba6-a37d-679f67fd02ea · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:58:06.866099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.055761Z digest=sha256:25f4a8cca7a243a48a5452babf4599016bac07206ec6eb6ff302ac5088610010

Observation cc6bca43-bac0-4e0f-8457-a8e607d929b4 · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Tang, Angie Boggust, and Arvind Satyanarayan

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:58:06.850877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T21:58:06.059666Z digest=sha256:2d42b439a444b037f66bf24144435c66a6fe5da64606745efd99cb4eba52f774

Observation 18b58b9c-f4e2-45db-a138-fae3e206dc21 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.064144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.064144Z digest=sha256:bbf5bea4d561fabca531c6cff456409fa3e1b5d63862d412241509123cdb951f

Observation 4ecd3fbc-bed7-46a1-8f08-6a08785fafba · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.069575Z digest=sha256:4a26c21c895c257a55d786ac02907d0844bef8c26fe06a1efbc1c3a5cf98044f

Observation a3deb986-72d4-4420-86cb-f567bde32c2a · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.075092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.075092Z digest=sha256:e43a45456b70b09b345789179e6653ada66013b1b37e7b775331a7de1edcd7ff

Observation 0381991f-33b4-4d98-b74a-379a32734184 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.079699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.079699Z digest=sha256:aa57d2f4b8d56b18e68168a8ea052fe625243f059a57d567fcd10547183d5df6

Observation ba27bc36-7c99-4308-afdd-6a82dec2b5b7 · outbound

This paper cites an unresolved cited work.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.090035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.090035Z digest=sha256:5c2429b64259db374b711879d32f5ca26735eb2b7be2052aec231f2559ad4f55

Observation e7a8651a-0138-4793-90e0-6830be302215 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.095164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.095164Z digest=sha256:b7bd1fef8f4dff19574a23feb1134ab537721f4cd189a98381facf19a6819065

Observation 5f841f93-ca19-418a-8f90-a1e425039be0 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.100133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.100133Z digest=sha256:0fe3ab34cef3be4f803865d32d739c178335891597136b3d0a6c59523f7ea4b7

Observation 52f49ded-e6fa-4408-a823-d99b8446e625 · outbound

This paper cites TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.104239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.104239Z digest=sha256:60ea604169aa5b453c61e382b092b1397879316fdca5bdebceecb27dbb6f8d00

Observation cd943d81-dbec-4632-9b13-ba85c9954024 · outbound

This paper cites A Survey of Large Language Models.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? A Survey of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.108749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.108749Z digest=sha256:1b2c6833eaaa9ea3bda2a3c3ccf4d087b01c16f7ea8d78e380dd9a52d0d6e070

Observation 348f4967-1323-4c34-8cd2-19619d73d37d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.112939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.112939Z digest=sha256:1f790fc15a89355d048b2170cc2f93d883f79b6e0b44cfbffc0dd9e56538e44d

Observation cc318073-01fe-472f-b1a5-d9194d15ef0a · outbound

This paper cites URL: " 'urlintro :=.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? URL: " 'urlintro :=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.117284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.117284Z digest=sha256:4927e9e1001c03ad151583acb496ac70ed91f3cb99b8eb97b423d9bf4923b219

Observation 26aa9106-fbab-4363-9da1-a37dc869d716 · outbound

This paper cites write newline.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.122004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.122004Z digest=sha256:47fca399c4b1271c28b3e41e5a2e1155efe8734f73f6bb40567cbb4783eab27e

Observation afc087d4-7c45-4c60-8921-77849caef84d · outbound

This paper cites online" 'onlinestring :=.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? online" 'onlinestring :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.126227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.126227Z digest=sha256:b501850a4e87aa21a2189880cbb40f06f2b519106afd2a3dc53eeb3f3be1ec7b

Observation 30ff7b4f-832e-4220-a2b0-a7f133e46cc1 · outbound

This paper cites write newline.

Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning? write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:58:06.130584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:58:06.130584Z digest=sha256:6c59028e8400d65066e645d8e9f5d3fe6ce5c525813a08d10c1574f4ece71c1a

Pith citing papers

Observation 4a2ef83a-53bf-4b90-80a8-e7d056c11299 · inbound

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge cites this paper.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:48.093541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.290988Z digest=sha256:f83443088f9428884eb6857a9ee223b2c58e949e082fa48946f938dd0cbd8652