Pith. sign in

Paper Citation Record · LEDGER

Dynamic Evaluation of Large Language Models by Meta Probing Agents

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.14865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14865 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:20:21.524087Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:56:55.163345Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 774029bf-e886-44f0-bf9c-2c7fb90d0874 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:16:41.768894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:e39c6166e6d74a02ec8d036711eb511169d9402f9750c71a8bc4364fd7d0d161

Observation e5667eed-8736-453e-a12b-878b49375758 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.165139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:25c37a183923ab88b308e99380754025a0f7137f360e451fbe395fd90e9c16e8

Observation 46ab20e4-e10a-4530-b7ea-acb9a9d7cd62 · inbound

AgentReview: Exploring Peer Review Dynamics with LLM Agents cites this paper.

AgentReview: Exploring Peer Review Dynamics with LLM Agents Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:38:37.076842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T23:38:28.005028Z digest=sha256:92ef3e351773806eedaaceccc55877810235f02f0a58c51bc8a4d01f3e13dd85

Observation 41996491-cd86-4db3-8569-cc51fdad9ac3 · inbound

Addressing Data Leakage in HumanEval Using Combinatorial Test Design cites this paper.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.524087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.524087Z digest=sha256:2d8653a15b3ec12384ec686dfb28160a6e3c96b03b14a341503966bb00ea8ff0

Observation 1ee7a73b-7456-4d9b-8939-e8246680a985 · inbound

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient cites this paper.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.717256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.717256Z digest=sha256:67c265b177f933412dc548fc865349e309bf5addeae7aa3c7ee8d4ba107eaa5d

Observation c5e53e17-3ea4-4387-8b95-c6b375353b8a · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 205

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:12.218289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:12.218289Z digest=sha256:56e7aa4aeb470e6bbcced510576b5e2266271a41e82d6296843d9c8ab5fee55c

Observation 3d656faf-f8bc-4921-a481-5c2ec615bb6c · inbound

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era cites this paper.

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:12.868181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:12.868181Z digest=sha256:bc41e5b476a5c288fda0ab2b904a817136aa3b3d4389a77f7d5a6b6297c9372e

Observation 697935c8-f6d7-416e-80ab-1f179bb4f393 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.747145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T00:06:08.545334Z digest=sha256:63adf3d8260d5c5531b72b270d6a0f3b9bbd30b0bd6ceb730419e79474e830aa

Observation a005f15b-86d9-4f86-bd55-bf92d3178a95 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.774630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T11:18:51.642405Z digest=sha256:b94118780527060fc172f2227893c36d64a524eefe95fc7f99d68436c808e48c

Observation e70dc3cf-cd69-4e8f-a7fc-6d8597385c83 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T17:25:30.971119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:25:30.971119Z digest=sha256:c032cfc7b639c72ad7c66594037d904f49383cc188a9a58e620d2357225c0057

Observation 62e7752a-db7f-4d7c-9e70-ab0064f0e99f · inbound

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios cites this paper.

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.164979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T02:49:01.527149Z digest=sha256:5eb3744b13566bafcecf601dc48bd29a79e1e23186d30a260c5b5d4489f64f7b