Pith. sign in

Paper Citation Record · LEDGER

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.11443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11443 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:31:14.453979Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 33ed73b7-1e5c-4092-b25c-9eb290242497 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:16:41.801484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:7e95f426f46627d5a36f1ddba110e9035087c61598bcb59c962fb150de2f3a1d

Observation 74c15749-b6a0-4e28-b9a6-2cb4d3f95d08 · inbound

Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems cites this paper.

Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:32:19.615151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T19:32:19.405615Z digest=sha256:e70123e9049c1e50cfa97e36def2cce9d9e79ad4005455ff779a748389f03082

Observation 31a6edaf-1644-45c9-8edd-118da0149679 · inbound

Multi-Agent Collaboration Mechanisms: A Survey of LLMs cites this paper.

Multi-Agent Collaboration Mechanisms: A Survey of LLMs Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:54:54.343887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T15:54:54.146003Z digest=sha256:1ed9b5fae67408f28e5d8f14ce253978a3a13b8520e94f7fdd9fb5d5a5c68ab6

Observation 8119e3f5-740b-4f32-b7f6-91d3a269085f · inbound

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation cites this paper.

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T05:31:14.453979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:31:14.453979Z digest=sha256:5ec630fe0330e4d54d575cbf6ceec7b6cf50c79058a26a5e5f06bb53be4f6935

Observation 40a336ab-ae3f-4f9c-b23c-d1c5d9d6005d · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:52:10.039764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:19d13deff00b89cea885f297129f7599e9ac5ab3bc9bb0f9d8f1b196eb9a02d0

Observation 2590573f-3035-4269-b202-ff079e461bb5 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:24.777981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:24.777981Z digest=sha256:b9bbb3a38ca91a025ada421a72a354a9bb77168d5cd1103f95a252dc97856386

Observation 27763c9e-edc1-48d5-abe3-84d74710dfc5 · inbound

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making cites this paper.

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:47.379383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:47.379383Z digest=sha256:b807f6326cc0d9fa579f0c6e79c62a2594cb188b80244ff08a00eff44d3a9d49

Observation e8efb481-ee2b-4016-ad62-e2f495fb91d5 · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.671307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.671307Z digest=sha256:f454b15ec17df8f13b9efb6324d9ca9dcfb9b72d60869c464bc06f739ab05f93

Observation 1dba8bc9-ccf8-4659-99c3-584e10b0cd81 · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:12.123748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:12.123748Z digest=sha256:5207a8c81f52cfa3bd60c558a0d3c721641712b3a62bcef5609648f5ba68d4ae

Observation 66e57e5f-cdb9-4967-b15c-286dece8a6b0 · inbound

Configurable multi-agent framework for scalable and realistic testing of llm-based agents cites this paper.

Configurable multi-agent framework for scalable and realistic testing of llm-based agents Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:37.742865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:37.742865Z digest=sha256:967161334ccf3003f25238c5be31f82bf16ffe9a1d9085658f7550884b67802f

Observation 9298ab66-356b-49f7-9ae2-0973d5c6213f · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:14.814860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:6dc2fbe1b8b3d0db64ecf89116d0151848fb02323f6ee3d595995cae951a2814

Observation ef268480-dc49-480e-82c8-af662745a063 · inbound

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web cites this paper.

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T18:55:33.713403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:55:33.713403Z digest=sha256:683dd64d61c2f0014e3286d803c39231d28d16da8c7a4e1ca589596416d27fed

Observation 6e832b30-0aeb-4c57-bd0b-e3ec2431d0f6 · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 284

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:18:20.405886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:b3831e91192584984de27e56b31fc80faab7dd087f86fee762949c6c76a7d4c7

Observation 23b5d3c3-a6ff-4d7a-8c0b-de4f11de2814 · inbound

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis cites this paper.

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:49.933190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:39:36.819142Z digest=sha256:fc76fe47943ec82d16e8474a1a9901e2ca7ca7aa8dfe772af467996ae7a9ab50

Observation cae55505-ab63-40b6-a7c5-05d54f56cab9 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.749142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:abca3a82690faacb48cb33d89ca38de21bb437241ad089a9bb182bf7ba9d6ba2