Pith. sign in

Paper Citation Record · LEDGER

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2505.02018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02018 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:42.512437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.794953Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3b2dc24-843c-46a1-9923-1f07e03446ec · inbound

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs cites this paper.

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:42.512437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:42.512437Z digest=sha256:e8406bf459e0d95f532ace9986407c31adf3d7bbc8b218486b4fe7bb750c4dc9

Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · inbound

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems cites this paper.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.587173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.587173Z digest=sha256:a2434b07deec3ceeab4af04a3bad7029ed2dc8ff3efdf669f39df36367af58e2

Observation 8bfe9a92-6741-43ec-9626-7cc730452232 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.710933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.710933Z digest=sha256:4ca07a7660ab87d960c286106c52e8906e7fb560f15a49e7b31e18b9183a5a03

Observation e2498c15-1ae6-4456-91ac-4b9828478eed · inbound

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models cites this paper.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.031094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.031094Z digest=sha256:6759f3eb98dc11eff5f5f5a027c05d032a0f164921e9a8b82a3877361c6c33ad

Observation b0663a81-39ac-4ee8-b7e8-d3f351d37055 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.234154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.234154Z digest=sha256:c1214980026dfa63c7af08051c8d421acf97dcc1c0f10a772069e244ffaae632

Observation a84e9cb5-7ffb-4bf3-afd3-a6f9e5cf63c3 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.420473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.420473Z digest=sha256:39714ae8a9d780c8aa354ec2648dc0c4ce3b1ebda7f87f15c31598bf41ba279d

Observation 64549480-e240-4c10-8fa0-a9dfccbb5966 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.617724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:671b78a58a9e7bf478adbbed1d8e0ce2f0562c051671869ff5ca76d0ea836433

Observation 797a0210-c121-42a9-957d-0453b00c769d · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.093227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:37b1466d0cca93f142bd27c80d6925da64d079fa9962b60a4ef2f548fe42929c

Observation 7e63eb40-1a4e-4c5d-901b-a879d875aed6 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:05.988418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:05.988418Z digest=sha256:6fcc4f63957eba7ff7f61d82be56cd63824371e4e3ef2dc41dd3cfe8780a9e03

Observation 5bc41635-32bd-45f0-a7d2-1591273abba0 · inbound

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video cites this paper.

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T23:47:18.378887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:47:18.378887Z digest=sha256:c81eaf251809e012c9e2b9811e9cf801d47bb1c184cead634044aa89970485c0

Observation 58be337b-c331-4d8a-9e11-ca6c1b02246e · inbound

ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control cites this paper.

ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:30.430849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:24:47.014162Z digest=sha256:a12ec3f57907c65bba96afd2700ec9021e88253e08297900562c745fe6995359

Observation 5463af0c-4957-439c-86c4-2a873252e8f6 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.709984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:39c22361259531d37187cfb2b2e53a7e6bd783cce5c0ffe1abbb65f3dd44de09

Observation b397e0b1-1b2e-4a2a-bbc7-90c5013a1e00 · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.590378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:23dc9c16a9cc9b8f7c9e1736c255da71f3fd834ae995c18a6a0fcdc240aab600

Observation 4f51371a-7990-4445-871a-7f81ef2b07f0 · inbound

SFBench: The SciFy Scientific Feasibility Benchmark cites this paper.

SFBench: The SciFy Scientific Feasibility Benchmark R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:04:21.322289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:59:19.857066Z digest=sha256:263d28948826d5233260a4ef75381a7a3b148f0a6ba3cdae9675ae1e9e775867

Observation 1bac3f50-ff3e-4870-b37c-ed52de082641 · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.796446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:59899f5abb50bdf326776cda187f597a7b1fab2168c036eaf9c2fea69a26a339

Observation 0c812abd-0d67-434e-afb2-cdd9590282e0 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:6296cb137e4a406d52646861cf400941e7c2bd22920a6bf1f11ad4121bf2a501

Observation 9f756069-7f36-4809-bd2d-a468123d613e · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:38.710219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:38.710219Z digest=sha256:27349b84f4a86b56bb9ad59d58d73a3eb2a232eda2e6854c880fed7975058a76

Observation 96a55390-e8ef-4e21-a9a3-17ef0e834002 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:13:54.686646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:13:54.686646Z digest=sha256:ebedde72e1680ed210cfa87c94b73ebbbf93298306b06bfe7242ff8d612147b5

Observation e38f4f09-afe9-4b91-9cad-06f220e0180c · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:39.422268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:39.422268Z digest=sha256:ea3cbf5b1e2c80891baab7a77fccd080fea66bde8c669bcd4ac702caf8499b38