Pith. sign in

Paper Citation Record · LEDGER

VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2503.23064.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23064 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:32.535373Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ffbe223e-d9c3-4ae9-9117-c5fa30c22be1 · inbound

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models cites this paper.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.535373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.535373Z digest=sha256:4d99b2ddba9f25ee2ac0729121184170d60366c9529c3bb68032ad3081c5a8df

Observation e694965d-e68e-4cd1-8aed-ce8277157796 · inbound

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL cites this paper.

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:18.617428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:18.617428Z digest=sha256:dfdbe5b8bdc33e86ba52110933e26fdc019ea6fcf6e5ae0535b58eeddd379b24

Observation 8d876b5e-022d-40bd-b173-4d0e16cc8068 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:58.020886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:58.020886Z digest=sha256:3f7f675568aff2c882cbea839e16250a275bb687389ad574e2b1dfa1e883b485

Observation 820be703-276d-4e5e-a29d-d32f4f8c9e90 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.516932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.516932Z digest=sha256:41b7a138d677132773828a4d36a8a95d812a24b52a4e8fb105cabd012a38c1b8

Observation c4745894-f539-49e8-b423-da4d3c2e6961 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.576642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:180a27dbaca542b8c35bb8c78ee741f1e7e50633c03c261810821fc6e2d00bd6

Observation 5c555827-bff2-497e-a72c-51c0a2fe2ba9 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.363121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:bd423cf45611d45f29c9f7917bdf2c3245cfcd50d16d2780399f4e37c495221b

Observation 7235af6b-ec94-4326-80c7-96d88d53c2e4 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:08.370312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:08.370312Z digest=sha256:ef4f9b964f1770e06e7a8e62146a1d878a79b063714ecdafb5f2d990167f18cc

Observation f8ba29c2-58f6-4551-ba96-2858f788498f · inbound

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs cites this paper.

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.242098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T09:18:33.234354Z digest=sha256:d2385328fe4587702171b167b5c4e81ef9d9c5d743388741ff683ad16cffd2a6

Observation b3fe7372-8644-4f0a-857f-b28b22be425c · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:25.099997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:ebce8df344b90737edd136f5450c3f6758db302e7bfa895d72d17d04930ff97d

Observation a15b774b-cb81-4ad8-a409-51396aa81da1 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.591007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:b60fbfd849b17ddae066c282861f4b733eda8b497c7c6813588d0e38a3ed9d2f

Observation 4e369fa1-6a0e-4b11-a4a8-4a508d15ca4e · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.512354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:a5bb97163f1400eb1f2f393a653f4a323d0796a63f355b0a11e2bc95e50d672a

Observation 98026f46-0f5a-4fa3-9851-1b2ecc01cdc5 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:08.781223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:38f7d45bebd739203bd2b5ffe6bf11016d4522b6c595ff5eab4997262e354dd6

Observation 05b8ec7d-02ef-49b8-bc20-4a3b338ace98 · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.848985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:9a9b9e65b12cdd61cc04d8bbb315a972a2a5f65b519aa6c09badb5e35bf777ab

Observation 31e4795b-8560-4d8a-9b08-2ad1f7908213 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.644799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:2f52e493e5959c32f19e0be3061076af1e8958d7d1ed130cfbb8e771112914b5

Observation 0114579a-b248-47e7-a0ab-997e05354579 · inbound

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models cites this paper.

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:26.128499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T18:37:27.171215Z digest=sha256:1e217803bd4906a27c61f43732069f2735648458a2d3f63bd656052731803516

Observation 54c48fe6-78a2-42b3-ac40-38996f921129 · inbound

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles cites this paper.

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:39:41.733830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:39:41.733830Z digest=sha256:3b65697b87568c61ba169f9ce408013bef2c07e675770fffdc1f3e90a32a4066

Observation 96159599-af9e-4b36-88b6-34d8ad3f2012 · inbound

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles cites this paper.

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:25:21.511122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:25:21.511122Z digest=sha256:ecba437f36837f747fb7ed2f4218425bf1443d45a2171506fd11e897c21fd23d