Pith. sign in

Paper Citation Record · LEDGER

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2507.12806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12806 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:46:02.635196Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:09:59.639299Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:10:16.398252Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965e8d88-5920-4fa0-997d-18671108562e · outbound

This paper cites online" 'onlinestring :=.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.805851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.805851Z digest=sha256:e63ec051a2505c79d9b3b1d32df647808a2d8137cadb39ab0ca70e263cb74919

Observation f48c8dec-5f67-4714-8de1-2e2eb14d62c1 · outbound

This paper cites write newline.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.871577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.871577Z digest=sha256:0321a6571da2f86ced63ccf8c66b6c78b08235999f93f936f7fae22aecf10e3e

Observation 54a68f72-8b63-4ba7-9b90-4afa6c213840 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.424251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:58.971979Z digest=sha256:8cf32e5b5323f7c48f56ded12572a1ac33e856f86bc35757c8cb19d207259286

Observation 0ac6f1bd-aad2-4886-b45b-2511928efe4b · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.326853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.050121Z digest=sha256:0f2b06dadcb3b72538c533c8141aac65d9b0f279f96cf12bc8fd895e3e0e7295

Observation 7f2ccccd-f0fe-44b7-975a-a4917b8daabc · outbound

This paper cites On extensions of the Jacobson-Morozov theorem to even characteristic.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models On extensions of the Jacobson-Morozov theorem to even characteristic

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.463822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.124683Z digest=sha256:161e4ea6c0f66c889a0b834f74565f837c815292f5a5508d2042ecfef27fda1f

Observation 6ddcdea5-a3c5-4978-b407-ccd568d67cb5 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.218192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.215761Z digest=sha256:2acbac59d9cea8419fdf13a85b9dd3fdaf2abe74b769f1b9b960950f9869cb4a

Observation ff70678c-e4b2-44c1-80ea-76173553b1ad · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.117258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.308374Z digest=sha256:a64a5408d6fef7f3a34df422534924d78440759a5ad4cc809de101827cd7e64b

Observation 382aea7f-4f3e-4d31-8b83-8ccda811beb7 · outbound

This paper cites High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.280548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.402697Z digest=sha256:3eb700d7d8cfdfa132cae143e1551e7d8f89439299188df2d84a0bf589178930

Observation 9963f502-26c4-4230-be21-15f6fc5160b3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.503158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.503158Z digest=sha256:3d319d99d089f783980e8e06947dd757930d16387893771835f463a37af716f5

Observation 8a58901e-2fd1-4e76-b9f1-3f512b7b46ab · outbound

This paper cites REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.600438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.600438Z digest=sha256:2a832c815ad413590c5233f2c83bfdfb0084c0e380ab3967276def26ad10ce57

Observation 52fde6de-9293-4f89-9276-73de583e86ea · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.689261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.689261Z digest=sha256:a5986b05f632fe81813a963f752cd2d44ad857de06dcdddf04c911748363c0dd

Observation a05bf665-2307-4f98-84fe-e0c110d1fff4 · outbound

This paper cites LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.002763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.767986Z digest=sha256:f0ad22904b820923593d0ade0e3ef6566cbdf01526a89a6ef3fbe2e1b9a0a152

Observation b413d103-a494-4f61-978e-3a676697e357 · outbound

This paper cites Recurrent neural chemical reaction networks that approximate arbitrary dynamics.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Recurrent neural chemical reaction networks that approximate arbitrary dynamics

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.888765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.835670Z digest=sha256:d10178dc5989818024d252d52710bd12c9f594834b87a377785c3cfc94d93552

Observation 0c2a0c8e-840a-49e8-a759-c6788b9e8af9 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.972064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.913761Z digest=sha256:70f8afbee87451c78a84e6960ad7933dd1e497d98050d27e3e96ad8b82a26b90

Observation d5e270d4-5040-4774-9917-b0b1f1891b6b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.003696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.003696Z digest=sha256:a9de434f4677655b1b114e7d7531b21f888d4a6392b06eb4978dd3cf1b314ceb

Observation da1ec4f5-2cea-4d02-9714-852c90562a85 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.105921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.105921Z digest=sha256:54c4a3e0481bc0da65d926135c837745aa9c02009eb72c3d5200e13d8fcdd8d5

Observation 0e3145b2-54c9-4e9e-b271-36a5d95ca7b4 · outbound

This paper cites ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.165311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.165311Z digest=sha256:40e7411dfbaf8b0f5a18edb51484afd1e877054a931e8204654f23f4d3af2884

Observation f16da7f4-79b5-4715-9121-e8d03a7cfd06 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.829666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.236217Z digest=sha256:dd85c96629f8b2c1aca32b6e928c80f5057b140e6647a51c7128e517841e4b8c

Observation 39716159-e2c8-490a-8800-1e82ea79b651 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.317956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.317956Z digest=sha256:ef0ad9a0a4f30c5651b449c647185e22b090734c6e49f19ab19d65deb01e8b86

Observation 10d30c61-c805-483a-aca3-821be6d5c8e2 · outbound

This paper cites PRACT: Optimizing Principled Reasoning and Acting of LLM Agent.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PRACT: Optimizing Principled Reasoning and Acting of LLM Agent

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.694873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.387306Z digest=sha256:3f4c2bfe8d094fca260e81312f6893749ab14b760eaa764359ef47ece30db369

Observation f1500648-c6dc-4fb5-bec1-bc9d37bc47dd · outbound

This paper cites BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.503290Z digest=sha256:3073763dc9881a3b9c78d3f89057dd2d9791dd6b6a92685a2bcc183297a6056d

Observation 3b85cbb9-9ee2-4265-8954-2dbfe0457389 · outbound

This paper cites AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.571687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.571687Z digest=sha256:b9014654cec5c19563b7548e4509ad525ea061eed7da1a31621af58d172af13f

Observation 3533abeb-e1d2-4a2d-a2b4-fe04b0e9cc98 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.691587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.643162Z digest=sha256:44bf90ae91e77ea2fa8924cb0b2d525af1070c3e966add9d23620d53924c8b93

Observation 5bacc67b-035f-41f7-9ede-cda04be96e11 · outbound

This paper cites ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.702329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.702329Z digest=sha256:e83e0719ed9e5e8e5e4b389cbc49c702fb24fc191fcf01b87f175e156bfe4e12

Observation 59d6be4b-64f8-49cf-977e-956a9ec79f4d · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.781635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.781635Z digest=sha256:aa0a1bc367680d806cde15064e1ed7bd294ff877e3767f42ee61f027d1ed311d

Observation 601a5aac-4ffb-4790-836f-c4219bfdb3fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.878106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.878106Z digest=sha256:f10e11ad58254ce699fd4f21e96806b4cd038b3f85c2bbb6d3ac478f81a7eae4

Observation 4b9de82d-5851-4695-a7c3-af6f7b0f7fc1 · outbound

This paper cites GPT-4 Technical Report.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.956599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.956599Z digest=sha256:c19ba186c972eae06f26a3a0500f98fd88efb5124cb48f7fb3beba10a2f61430

Observation 1f61e1be-832c-47c5-ab38-ffe3057f52e3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.487334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.032125Z digest=sha256:14f9a168a7cd55fc0f3b79d509a93616ada1dcad0706fb9ec602525c41aff55d

Observation afb76a43-b3a7-42b7-aa5e-5c8f1f69623b · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.118506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.118506Z digest=sha256:01052d0c4f60800570a050dd2c473b94cdb31d4bf069ad87a558164178b5b2ad

Observation 98a46a96-655c-4937-a520-f84420cf7735 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.319496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.176261Z digest=sha256:517edebe3b52f7a6ea14b749bd83e663a12a5b910ac0cb58637423da6292291b

Observation 277c3b40-f282-435e-8481-dc53cbe330fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.147837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.251698Z digest=sha256:136796ca725f070325b1b6834abb7f0fd17a6b6d3d82a73fcc3c07302a8b77e6

Observation cea0ddfc-24f6-4de2-b2ae-41327a4f6009 · outbound

This paper cites PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.340983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.340983Z digest=sha256:a74854d316b08071ed3ed2cbb25a5ed8a867997a7a1c25fba5d3f4aa3bc6687c

Observation c31e5581-2880-462d-af0b-097ac6237e88 · outbound

This paper cites Peak Age of Information under Tandem of Queues.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Peak Age of Information under Tandem of Queues

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.339657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.420089Z digest=sha256:088fb1b5909751daae3abb8b6f790c8f56518dba7d813f64f3c1737aa48e4489

Observation 91de0b95-21ad-4e1d-ad2b-f137266aff4e · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.019555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.505456Z digest=sha256:ef0ed5edca18cc5437d2cb7e6014d9a49e682fcd82fd5ab3a95d0b4f2f911a5c

Observation f13d3aeb-25cb-4548-b755-eafdb31c03dd · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.909834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.570593Z digest=sha256:0cf702d7cc5b46c0cb6a22acdabac52052ef93957955f518984b62eff12e3ac1

Observation 0fe92b4f-1fe3-41e0-8586-4fe378d8f71b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.623365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.623365Z digest=sha256:024eda43e9f775634bef5afc6d876be45559ab3e72fcf4ccd1dc3be41572c3ec

Observation b196f587-3af0-4831-97d9-a95dceff3733 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.750655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.676364Z digest=sha256:227637cef8b0abdf695083a9e3f905418aff8012aaba2b77e17bc7ec6521707d

Observation 6aaf830b-3ec9-4809-b1b8-55ab9db3030f · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.738182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.738182Z digest=sha256:60db613493366a8eba67528a5579b394424e56b3b87475e9899b5adf820be61b

Observation 24645806-899e-471a-9159-5064ce9a8a91 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.826594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.826594Z digest=sha256:486cc6af88ba6f0009d6a1b2e3043cff63b0f25f53d2855f5f89bb402b7228c3

Observation 1eb775e8-6a66-43bf-8819-92406bbd28fe · outbound

This paper cites MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.921739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.921739Z digest=sha256:3f474ca930c56cecb0eb789362bf05bc443022109bb839852ee46e40a9835d9e

Observation 8cdc92fc-d07f-4d04-9a2a-05d6c3adb33f · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.613287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.974886Z digest=sha256:2945cf95702fa00837f59637df58e3dd86be05cc05cb4ec0a57d7b9974dea4b8

Observation a3b0ae8a-5fc8-465d-bbed-990c8b4dfcc3 · outbound

This paper cites A Survey on Large Language Model based Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models A Survey on Large Language Model based Autonomous Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.056791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.056791Z digest=sha256:209c6bfac5885b06c41bf2ee8f8861e8e0c155c8744510d58f843024e9b1df3e

Observation a242dc43-64ca-404a-a3e3-fadb42c1a74e · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.127936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.127936Z digest=sha256:049a516e578ceefaf846fe4679bcf3dbb414ea9c9d53411f5019474b24293f59

Observation c607c54f-dc16-4cb5-bf78-d021f4fff877 · outbound

This paper cites ActionStudio: A Lightweight Framework for Data and Training of Large Action Models.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.088040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.199468Z digest=sha256:2b0f5ed8b87d2b8c0e6f8e29826d5d1d306849552d43da4d8dbbc8fffa25801c

Observation 380140d5-000a-4b40-ba1e-8f10e9c4d317 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.276368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.276368Z digest=sha256:5024816fcc15243dcd360d9b4a4f6344ef4b24d2b093df4b3b88d98ddd6f556c

Observation dbf4f137-366a-4b92-904e-d90b7f7e9806 · outbound

This paper cites DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.956374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.333562Z digest=sha256:6c889a86bb44612e47223a7f6030cf489801734ea41a1f8c4a594560201dbae3

Observation 247dbebb-307f-41a6-a3ad-76cf635dbe64 · outbound

This paper cites Laser Printing of Silver and Silver Oxide.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Laser Printing of Silver and Silver Oxide

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.814754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.418806Z digest=sha256:e5a12fab13ea39aa3e9fd004ea11b62fa0d111ff911ca14ae21f587857e89eff

Observation bdbe1501-a723-46ee-a294-bc95b9e765c1 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.488733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.488733Z digest=sha256:85112d39fd8df1d0751976360321c57492442349fed55884d4d48f9073aabbc7

Observation ccbb6752-8574-4506-9e62-d15cdad40e39 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.574065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.574065Z digest=sha256:aee02652cf1d3aa939538a459d74beb30c8684368dd8fb7eab5044ee9355a627

Observation f51c148d-fa70-4e16-babb-41243255782c · outbound

This paper cites Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.635196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.635196Z digest=sha256:a5448f6659af762aa6872bd29a3ab9bce062bbfe38967b26fb458489a498d139

Pith citing papers

Observation b57c8d05-524a-4962-90ca-703cb0946f1a · inbound

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools cites this paper.

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:09:59.639299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:09:59.639299Z digest=sha256:c4a648476194afa4a091efbe048b2113d286ff9137decc52ec880ada621c670f

Observation f50b43b1-2286-4179-9d51-536a3c89175b · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.621785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:8340963b47843377157a71e13e66d67ccc5298c172a00bcfdb8e86de8572d2e0

Observation 639cab6f-2318-46d1-b552-f2ad2c5d1ae4 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.401136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:9dbfab3ac286efe55a5201966556dac30792ce31773ef96af88b468870de19c2