Pith. sign in

Paper Citation Record · LEDGER

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2507.12806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12806 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:46:02.635196Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:10:04.253250Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:10:16.398252Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965e8d88-5920-4fa0-997d-18671108562e · outbound

This paper cites online" 'onlinestring :=.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.805851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.805851Z digest=sha256:4c2497b143cd4dff47a1b2ce1dcfb305b697f2094a18315227ac49fddc9a4645

Observation f48c8dec-5f67-4714-8de1-2e2eb14d62c1 · outbound

This paper cites write newline.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:58.871577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:58.871577Z digest=sha256:39673f4493950ad3b6f7643d5bf6240542387db3e321a7c7bdbf01bb3099ca3f

Observation 54a68f72-8b63-4ba7-9b90-4afa6c213840 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.424251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:58.971979Z digest=sha256:b3cf48686df1839ea8748af363a3e9d6a07728e0becabfa342e89486787c2e19

Observation 0ac6f1bd-aad2-4886-b45b-2511928efe4b · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.326853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.050121Z digest=sha256:2f35409f6146b16664960ce4f6419c7ac0db12f7f59b774ce405c1f466f91ce3

Observation 7f2ccccd-f0fe-44b7-975a-a4917b8daabc · outbound

This paper cites On extensions of the Jacobson-Morozov theorem to even characteristic.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models On extensions of the Jacobson-Morozov theorem to even characteristic

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.463822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.124683Z digest=sha256:4f257bf53e143d3e607d27eb828ed4709a6b896ae9d7b9693d0e9adf15a3dca2

Observation 6ddcdea5-a3c5-4978-b407-ccd568d67cb5 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.218192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.215761Z digest=sha256:767ccb53304bb28203f864b5ac1c6caed3c60025ce32ae4507ca08cc2681ae43

Observation ff70678c-e4b2-44c1-80ea-76173553b1ad · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:06.117258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.308374Z digest=sha256:a73e67f9c2d959258ccebf8696816f609da7c69209d030200b755acdd22e829a

Observation 382aea7f-4f3e-4d31-8b83-8ccda811beb7 · outbound

This paper cites High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models High-temperature oxidation and nitridation of substoichiometric zirconium carbide in isothermal air

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.280548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.402697Z digest=sha256:a72186edb09c022bcb1057de7578ddfeb29565269947714d31c5125b45f930fa

Observation 9963f502-26c4-4230-be21-15f6fc5160b3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.503158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.503158Z digest=sha256:66a03ed2bfa56f168b1a78a5dd36874fe246cce5d0e479f49ce8730e47257120

Observation 8a58901e-2fd1-4e76-b9f1-3f512b7b46ab · outbound

This paper cites REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.600438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.600438Z digest=sha256:ce33bc3a240d8ed9c315b8a4ddc3e5bc77e34017bac27239d5de086512bf7f38

Observation 52fde6de-9293-4f89-9276-73de583e86ea · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:59.689261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:59.689261Z digest=sha256:037a09d84100e98ec5031806664b599996014b9ad3850bfd97db966170bdbec6

Observation a05bf665-2307-4f98-84fe-e0c110d1fff4 · outbound

This paper cites LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:04.002763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.767986Z digest=sha256:06f254bd2e18df5d6308a86112a73deb525405e96baa958e9f6d73a91aad5a53

Observation b413d103-a494-4f61-978e-3a676697e357 · outbound

This paper cites Recurrent neural chemical reaction networks that approximate arbitrary dynamics.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Recurrent neural chemical reaction networks that approximate arbitrary dynamics

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.888765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.835670Z digest=sha256:35ebf885bdccfc27dd6fbb821e344e0b22f2c4195c5dd3e06edcd7a15c8afa91

Observation 0c2a0c8e-840a-49e8-a759-c6788b9e8af9 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.972064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:45:59.913761Z digest=sha256:83aa2be3e1a59b42fead31ccd4ff0770be1e2f0f036e467ec481a3ba8161f71b

Observation d5e270d4-5040-4774-9917-b0b1f1891b6b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.003696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.003696Z digest=sha256:34146242bb65d80b1db2629fb62ae8e7d57bfa409048efbb166a0f142401430e

Observation da1ec4f5-2cea-4d02-9714-852c90562a85 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.105921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.105921Z digest=sha256:5f809290d85af4c6bad7d6d3e0409d3b0122bf347393e874c70bbe6d217d127c

Observation 0e3145b2-54c9-4e9e-b271-36a5d95ca7b4 · outbound

This paper cites ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.165311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.165311Z digest=sha256:0f30f0ebfdb6e40c8d307c6e6d730e65138919d8631b4eba1e3d3f5961370739

Observation f16da7f4-79b5-4715-9121-e8d03a7cfd06 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.829666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.236217Z digest=sha256:e3295d53409c979c201a24efab121bc46ab0047d8be894428f643aac14d84a94

Observation 39716159-e2c8-490a-8800-1e82ea79b651 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.317956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.317956Z digest=sha256:afbdcb5fb623a60f19a754f0c1e24196f89afd39d77853f293551722ef95fbd3

Observation 10d30c61-c805-483a-aca3-821be6d5c8e2 · outbound

This paper cites PRACT: Optimizing Principled Reasoning and Acting of LLM Agent.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PRACT: Optimizing Principled Reasoning and Acting of LLM Agent

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.694873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.387306Z digest=sha256:36e99c18c075063082b035869dacbc281a46471d027a6bbf65736220ac66c66e

Observation f1500648-c6dc-4fb5-bec1-bc9d37bc47dd · outbound

This paper cites BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.503290Z digest=sha256:751f4a1e1a382d28e7b1d19787a8bbfe448cf4d7bb07c8f349fc8984edc05fec

Observation 3b85cbb9-9ee2-4265-8954-2dbfe0457389 · outbound

This paper cites AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.571687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.571687Z digest=sha256:a2f7607c1847fade515ccc2613a582df37653747c2fc915dfa59a3d127f07556

Observation 3533abeb-e1d2-4a2d-a2b4-fe04b0e9cc98 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.691587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:00.643162Z digest=sha256:d80e32bc17db9168b5d1b745269b4623b3ad20ffe6f6b09795a26ba6d00f9990

Observation 5bacc67b-035f-41f7-9ede-cda04be96e11 · outbound

This paper cites ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ScaleMCP: Dynamic and Auto-Synchronizing Model Context Protocol Tools for LLM Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.702329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.702329Z digest=sha256:a43410acdd1bfad5ea21b5644fdfae12ad11e083837d6995e437956ee6380286

Observation 59d6be4b-64f8-49cf-977e-956a9ec79f4d · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.781635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.781635Z digest=sha256:05e7fb434b390741d15c71d26f3eec95ccb0c13bc891881eecfe8c05c43a6942

Observation 601a5aac-4ffb-4790-836f-c4219bfdb3fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.878106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.878106Z digest=sha256:81b94da6d38d5ccb693745c73d6ec1f677fed79284382eb0c36162b463a9572e

Observation 4b9de82d-5851-4695-a7c3-af6f7b0f7fc1 · outbound

This paper cites GPT-4 Technical Report.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.956599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.956599Z digest=sha256:25752040fcf430c38d7cbe91598fb7a9023300bdaa8a875bcbd33c4cd9d352c5

Observation 1f61e1be-832c-47c5-ab38-ffe3057f52e3 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.487334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.032125Z digest=sha256:1a2505e5f9765f0c0a0811939259d20af5820004d6a2b4b1296d8c6907664ded

Observation afb76a43-b3a7-42b7-aa5e-5c8f1f69623b · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.118506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.118506Z digest=sha256:33d93238d90864762beb16dded12856a6e97064902f88ffda4cf3601739826ab

Observation 98a46a96-655c-4937-a520-f84420cf7735 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.319496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.176261Z digest=sha256:755c3688c4a5ea047df06632560d059313cf0fd1e32d38880dbae1fc04f5a98c

Observation 277c3b40-f282-435e-8481-dc53cbe330fc · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.147837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.251698Z digest=sha256:febf1561e6959a403a1a867bad47b696d2ac354d95b212263eb51fb29821873a

Observation cea0ddfc-24f6-4de2-b2ae-41327a4f6009 · outbound

This paper cites PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.340983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.340983Z digest=sha256:d71a24810378075de7dbe353e48c6321ce3f90be59f2c6bc023d8afa10fdca87

Observation c31e5581-2880-462d-af0b-097ac6237e88 · outbound

This paper cites Peak Age of Information under Tandem of Queues.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Peak Age of Information under Tandem of Queues

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.339657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.420089Z digest=sha256:af23ac5f3e61119f5bb7a37902eded390f40a5b914a5d079a48bb29e20cb054a

Observation 91de0b95-21ad-4e1d-ad2b-f137266aff4e · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:05.019555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.505456Z digest=sha256:7970e0d460ced98dafab718ed2ab01d5c6deb1479c313c3e5a34997d3133144b

Observation f13d3aeb-25cb-4548-b755-eafdb31c03dd · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.909834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.570593Z digest=sha256:8a0c85ab25253527a79086d3d90b4e16cb684f20707d8f8523052c5485151207

Observation 0fe92b4f-1fe3-41e0-8586-4fe378d8f71b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.623365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.623365Z digest=sha256:221be55a54925fd1af1242593701efc9572597a7bf65def7794b96300c6d99e1

Observation b196f587-3af0-4831-97d9-a95dceff3733 · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.750655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.676364Z digest=sha256:fab0e54ee4c309a9004016ded62f7a602510859c5886cb361c67cd09d459fc15

Observation 6aaf830b-3ec9-4809-b1b8-55ab9db3030f · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.738182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.738182Z digest=sha256:869c2f8b36bb2533dd395e4125b6d8c1f412ee7d5fdb306d4c34630831ca465d

Observation 24645806-899e-471a-9159-5064ce9a8a91 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.826594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.826594Z digest=sha256:ea0cdc5e12dd07b89e3c220653cb597c21402bda173975e7e672137639792a07

Observation 1eb775e8-6a66-43bf-8819-92406bbd28fe · outbound

This paper cites MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.921739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.921739Z digest=sha256:4d07af12586203250fc10f1d482f760c5835e361a8fc7229604a369a029604fa

Observation 8cdc92fc-d07f-4d04-9a2a-05d6c3adb33f · outbound

This paper cites an unresolved cited work.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:46:04.613287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:01.974886Z digest=sha256:8b7ab0e1cb3751e7d4264b1f4fbf720a43b4e752669cede1ffd1cd9b22ed518d

Observation a3b0ae8a-5fc8-465d-bbed-990c8b4dfcc3 · outbound

This paper cites A Survey on Large Language Model based Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models A Survey on Large Language Model based Autonomous Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.056791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.056791Z digest=sha256:f14e945207987ccb1e5f4f7ef5175f3a59b4e027f33c39a11caf9fa9dc4fbfa0

Observation a242dc43-64ca-404a-a3e3-fadb42c1a74e · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.127936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.127936Z digest=sha256:e1231c37a5e99ed814ddf8943f0ecff75fa6355d483e099ba93c7a62477ce23f

Observation c607c54f-dc16-4cb5-bf78-d021f4fff877 · outbound

This paper cites ActionStudio: A Lightweight Framework for Data and Training of Large Action Models.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:03.088040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.199468Z digest=sha256:c63c4ec74052bd370828c5decdeb47835569a25448566f6bf90ea4ddce89cf19

Observation 380140d5-000a-4b40-ba1e-8f10e9c4d317 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.276368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.276368Z digest=sha256:d7b5174c05cdc58d2835794eb8d2f56c65026f403f4ceb98639d0e8d5a5eaaf0

Observation dbf4f137-366a-4b92-904e-d90b7f7e9806 · outbound

This paper cites DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.956374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.333562Z digest=sha256:2fe8973514fde2050c862aacbcb7999ecdea229390988bf34a52729a020a8784

Observation 247dbebb-307f-41a6-a3ad-76cf635dbe64 · outbound

This paper cites Laser Printing of Silver and Silver Oxide.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Laser Printing of Silver and Silver Oxide

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:46:02.814754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T16:46:02.418806Z digest=sha256:1a294628dd2bbdb8f2644a3c24e3c2a71669308c94c27b489ed05fe144f06289

Observation bdbe1501-a723-46ee-a294-bc95b9e765c1 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.488733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.488733Z digest=sha256:4a046c643ec853d4d73b476f24a4ed3f832209f90cf1d79903e89c71d338e72d

Observation ccbb6752-8574-4506-9e62-d15cdad40e39 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.574065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.574065Z digest=sha256:8eb75492a69d961ad578a324c2bb1f418bebcfd546fe6515b7c41c65669ff7ed

Observation f51c148d-fa70-4e16-babb-41243255782c · outbound

This paper cites Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:02.635196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:02.635196Z digest=sha256:25be18b8897dbd285284dc514f171057ff0c7cc60b660a64aaaa608dcc3c0e92

Pith citing papers

Observation f50b43b1-2286-4179-9d51-536a3c89175b · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.621785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:6a585442d5c1e92e4ee4ac8182b9389a79e09d61dd4e0ebd9c11d22ecefcc0d4

Observation 639cab6f-2318-46d1-b552-f2ad2c5d1ae4 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.401136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:d110cb06cda696b5e26815804565cd7c9ab65c05f1fec828ab1a9fe1e756388f