Pith. sign in

Paper Citation Record · LEDGER

AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2401.13178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.13178 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:31.241598Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f76ff87a-3d72-4686-8967-5ffdc37f8e4c · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:19:32.581967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:712fc986db2549467ea105125f0277b3aeb8a74202f71ee999662b54f355051c

Observation 8fb94b4b-3c84-4b6a-85c1-06bcb76127a7 · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:29:27.515329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:4f0cfa90617b6a803bf15b24ee42aba4ee166ed935aa0989d0c478a23aae8f26

Observation f6202b31-7864-4934-af61-b7d2544c0570 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.416304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:ad3b66f29bf5fafaf3af60d7a6a76819e093041117294c1cee06d50f6c2de5c5

Observation 78642a46-cfb1-4bda-b33f-2afb0c9e06cf · inbound

Make Planning Research Rigorous Again! cites this paper.

Make Planning Research Rigorous Again! AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.241598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:32:31.241598Z digest=sha256:753d6fc1a25f544d28e43afb3d7f26e127d6714e8f84d435e5d393ecf1024b56

Observation 07df1bbd-0184-437d-ac61-5330f2269e25 · inbound

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback cites this paper.

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:55.654047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:55.654047Z digest=sha256:95cad88b1252490a6ef4d195c16b437da48c38014faa48db141e035f21bddcb2

Observation bc6f3ffc-5f94-4b77-aaed-b64b99c6b270 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.228552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:8ddd688bf3a3aee85d4fb6deae59c30bbc416641bc2dcdbc9fae43b9e8b97d43

Observation 581b98ec-4edd-4610-96ee-2b871f2fa734 · inbound

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems cites this paper.

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:59.063879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:59.063879Z digest=sha256:e181b23ac4cc43821ade7f49edfadbef8d353d2487449d7ae2ce5dcb96c74f03

Observation 3a1d57a7-a4b6-40c4-a890-fc0c4be9b11a · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.104695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.104695Z digest=sha256:7e448a155b67590f2e26c47b3958beb7c53398e631c2a2812f12ea0cb0b32e1b

Observation d775f98c-79ce-4fda-ab4b-d2fc7398bc4f · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.905361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.905361Z digest=sha256:9bab835836a34c6329826945c15bd74fad3faf942957400a1a9e6c61dd62a480

Observation 59d6be4b-64f8-49cf-977e-956a9ec79f4d · inbound

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models cites this paper.

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:00.781635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:00.781635Z digest=sha256:05e7fb434b390741d15c71d26f3eec95ccb0c13bc891881eecfe8c05c43a6942

Observation ff4c10d5-7239-4858-9037-39477c88bd91 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.699335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.699335Z digest=sha256:9d50a652633fe83e6099c5de1629dfa9c39d88645790cccd3275a0bbdd89d943

Observation f1ba59cb-94fc-4e95-8896-a3cff2265162 · inbound

Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents cites this paper.

Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T01:01:38.735705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:01:38.735705Z digest=sha256:d454bd9ad4844b21334b1ec92b1bc034e60ed4abbe3a04f9b444c878183da2a3

Observation 14eba690-db7b-4d04-92b7-fd4fb6d5f73e · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.243421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:a3d610530eea92bfd670bce504898baf884eef9c1d66d944cf51e774e5dca90d

Observation e9607fd6-e224-478a-ba99-489caa8123e5 · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:18:20.359761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:0ef35dcfb2db0ac3cd2049f3452fe0aa4b8c9c9c3859028b1ab25dc102b1cfbd

Observation 9a1aa186-e393-40e7-ac4f-7548b4f78508 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:57:24.553918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:f1d5198a07a4ea3e543d6cbf4a40b3159f3e20c922cbbbde36ba56472f1d214a

Observation aa8b3b2f-1d2b-428a-8fa9-a65bb2bb6c90 · inbound

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill cites this paper.

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:02.516415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:36:10.725278Z digest=sha256:3a1b70b827cff14cfd259c498485af02ff8241d64f47f08e1884ba867bf0e82f

Observation 6045f6b9-eb30-42c8-a7bb-07f156e290bd · inbound

Holistic Evaluation and Failure Diagnosis of AI Agents cites this paper.

Holistic Evaluation and Failure Diagnosis of AI Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.221202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:17:28.794622Z digest=sha256:f1157a5b96d3fd9f1d00a329c72126aa21045e2dbfec86c2ee87c1b2466da6c2

Observation 2aa9f43d-5d04-4293-b1fc-66dad36290ca · inbound

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design cites this paper.

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:23:03.208096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:22:40.544305Z digest=sha256:94e5cac109ffddeee419940d0384b923265c606ab757c85c961190f95991f906

Observation ddced076-1a45-4ff7-b2d4-29f0fe116742 · inbound

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design cites this paper.

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:34:59.702391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:24:57.376385Z digest=sha256:c9beac99bc5aceb4f54b601a59d727392974476c1a191ff3f9c4361266e14569

Observation ba622d8f-9203-4871-876f-a11f9308063c · inbound

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents cites this paper.

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:44.106800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:34:43.686189Z digest=sha256:c464348045c8ae6beddca5184d7fd210cee19fed07840a03f95d0d3cc4fe7acc

Observation 1508867e-de20-4be9-b4e3-fbfadd577f49 · inbound

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents cites this paper.

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.825258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T17:52:14.944875Z digest=sha256:a389243bb6ec6563a801a6732455440e759d218794847befc3d30530249cf64f

Observation 6de4b676-0d25-434b-804d-9b7c2be9a0eb · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.094083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:a8e7ce64ade7bab4c30efa8d7de7e1580316525942b933f2c0e1c52f56c03397

Observation 702924c7-5ae9-4c68-8ece-9d8b6d499abe · inbound

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows cites this paper.

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.284768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:42:20.527063Z digest=sha256:177549ebfcc44023c15a3c61fa7990e62b52d581ca40af16e9fd376427c4dcc3

Observation aefd2fa4-d369-4759-95c5-8aa39d835a58 · inbound

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents cites this paper.

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.634908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:28:08.382201Z digest=sha256:c3f8a539103cff97b24c06360b6e12d0f5520ac7dc25453b35a0aa327cf86539

Observation 929c8d43-2632-4205-9d91-610cbc9c1482 · inbound

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness cites this paper.

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.015798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:04:29.791558Z digest=sha256:1fac8f7ff14b9bb6273f6a3ec3a072074fad7611ec92a3e3dff104f67dd29fd9

Observation a38e9815-8bdd-43b8-be5b-62b0e1aa79ac · inbound

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development cites this paper.

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:41:32.038494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:41:32.038494Z digest=sha256:418ae6d08710f7b3c18d2962bb669e96eb7f5a3e94f52f12071dd78d7083d78a

Observation 285ed43e-09ab-44c5-bc7e-271c70d3b5ce · inbound

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows cites this paper.

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:35:43.053786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:35:43.053786Z digest=sha256:ee3768b18482cdc05ae79dd506aafbeeae7a0bb40801ef6188b1d3e533a2e8dc

Observation 078adbd5-018a-4b12-9862-500131aaa417 · inbound

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems cites this paper.

Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T16:10:44.381736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T16:10:44.381736Z digest=sha256:d2f8e5b239a4a5c0ce3cf59d2dcbe3a4183fdaf9d4f653c0112286196860ddf8

Observation b53841d8-7de1-4d38-8c00-2d2f42968e7d · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.674463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.674463Z digest=sha256:20244a7ea53aba8f83882c691eee5aae8d3f42339c1311bc00eae70c7cb0ba2d

Observation ecde07a1-275c-4b25-ba4d-c7856806209c · inbound

Beyond Component Testing: Validating Agentic AI Systems cites this paper.

Beyond Component Testing: Validating Agentic AI Systems AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T07:41:24.314016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:41:24.314016Z digest=sha256:9158ae0c5845642aab47606b449b6c77702d01bf883df0bb34341f7696c786fe

Observation b9eb8009-1da9-4cfd-b2d0-f72eff0b0898 · inbound

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning cites this paper.

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:44.867963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:16:44.867963Z digest=sha256:effdd1a744fdb3abb6efa395e7362278639009da58d33e2d46f03d5d7eb35290