Pith. sign in

Paper Citation Record · LEDGER

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

As of 5 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2604.16706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16706 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:00:45.789649Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:18:52.865771Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T04:18:58.494760Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact26
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b93512f3-f6a7-4429-96fc-8429d87f3c1f · outbound

This paper cites Krisztian Balog, Donald Metzler, and Zhen Qin.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Krisztian Balog, Donald Metzler, and Zhen Qin

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.369234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:ecedb743656c09a266dc5c9f690b203a5c4a1fc573c925ed154305b89971257a

Observation a4f6cfe1-bb76-4158-a342-97083b4e70f5 · outbound

This paper cites Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.979582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:402132f0d8a8567908e017bc1f61c0a37c5ca638e45ad403a719b71ea673f02c

Observation c0443dad-030b-4911-9d9f-52715a0643d6 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.902869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:c8b114c1a38691b1980505aadb6c0c7ff0520f02ed2bf6bb2acc91390daeceab

Observation 894f4d4a-d92d-4f67-8db3-f8d7c631899c · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.964601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:ce3b8d23ac275a9baf5e30000e198b7f13360db9fcc93106b4c590f142c68814

Observation 1982637d-2fad-4555-9868-189c85bf8fb4 · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46

Reference 5

Resolution
metadata mismatch
doi, observed 2026-05-10T08:02:24.365360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:f59a0c4ac5168d47c0515b96df40e7ce8d3f2add1f4354b9698110957811bc5a

Observation 347df548-b830-467d-aa4e-2a4d1785a44d · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:37:40.926471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:01f5bd0bd57ac35e3ff90f1ed3fd4684321e479a2021ba6b49ac276601a801b5

Observation 722986dc-69df-4417-9c05-8f4d29ab3bb6 · outbound

This paper cites Farquhar, J.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Farquhar, J

Reference 7

Resolution
metadata mismatch
doi, observed 2026-05-10T08:02:24.374673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:c6369746c5443f8e67ccd84457997a81ffeb18d4c42833e50c70ea13fee9132e

Observation e19a740d-89fd-4719-92c4-033f2953b5be · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.969672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:c00c94c7a600808e9b9e8899c968fe1047f420a1186545fd8fae63e37435441b

Observation 96d1e742-97c9-4f62-8017-650134e2715c · outbound

This paper cites A Survey on LLM-as-a-Judge.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on LLM-as-a-Judge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:02:24.959501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:fc4719fef49bd6b2deeaa4dc51774ca1527efe7a1991109365ce0eab2ebbc57c

Observation 03c46076-2444-48de-b579-f21a78039509 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.888032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:0a5f0a4471ef6476e64bff4efc0d93e62bf9c1d78afd8907f7de45c8c6285d88

Observation d5830afb-2fe5-4599-9684-bf9b972c93f5 · outbound

This paper cites Richard Landis and Gary G.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Richard Landis and Gary G

Reference 11

Resolution
verified exact
doi, observed 2026-05-10T08:02:24.371924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:e47b131f5a44a5603aa89187051b31ed559d35c2d87fced505c5485ac91110e7

Observation 4bb8cc40-d144-4315-b4cd-0bd193dcff07 · outbound

This paper cites LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench LLM-based agents suffer from hallucinations: A survey of taxonomy, methods, and directions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.977159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:544eec23c56cd13429006dabc7942365fd73d69de3355c7d89facbfe854f83fe

Observation 747a18a9-7c70-4bdd-8932-3306f4902bb9 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentBench: Evaluating LLMs as Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:40:05.312349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:e7a4003b3fb4fd8bb1829345cabf1519d17d9ddecd246c6148e96715e263068d

Observation e9f13570-68cd-49f0-851e-225019763238 · outbound

This paper cites AgentHallu: Benchmarking automated hallucination attribution of LLM-based agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentHallu: Benchmarking automated hallucination attribution of LLM-based agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.987100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:fe837d46a17062de0dc1ba78f453a8359ad1fa165812f5f1db73403486a37a7c

Observation ce27df5c-9ea0-4776-841a-729af5d0bf7a · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:11:22.649439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:2084400cb579b8007f5dc806f9aa20a25b965ea68bd51b146fc16c383db03d3e

Observation 7c9cb805-072f-4416-b1a3-8c3a330cc5b5 · outbound

This paper cites Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao, Prathyusha Jwalapuram, Amrita Saha, Doyen Sahoo, Silvio Savarese, Caiming Xiong, and Junnan Li.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Ziyang Luo, Zhiqi Shen, Wenzhuo Yang, Zirui Zhao, Prathyusha Jwalapuram, Amrita Saha, Doyen Sahoo, Silvio Savarese, Caiming Xiong, and Junnan Li

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.972157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:6a54417fc545616ac990fa293f5db11f54612f8ad3e42c1b81191b1ac74e2e19

Observation 01c3f8a3-5262-4b46-abde-c66aaaf1f6b1 · outbound

This paper cites RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.962158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:76f2d3338f3cf28a633fdca6e385a9f77d0e6bfb15d8be86ee1c3b0c168e9fac

Observation f92234bf-83be-4c8a-a91d-296d73f7f277 · outbound

This paper cites In: Yang, G.H., Wang, H., Han, S., Hauff, C., Zuccon, G., Zhang, Y.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench In: Yang, G.H., Wang, H., Han, S., Hauff, C., Zuccon, G., Zhang, Y

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.354859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:c30db00ac42ed54ed11e8b3b0c0728c21a119252f11c29988349d23f8553ce61

Observation 911c8154-7ed0-4cb5-ae30-e7a78005e804 · outbound

This paper cites Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.347506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:28b5993f93de682651caaea5b752c84dfdc6e92871e5280590a2f863e927c3a6

Observation a506623b-98f4-4159-9a21-1e3775c73ae6 · outbound

This paper cites Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.999729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:0152f3040d2d8b2f1273b909ffa451caf88b7f905b115ab5f4b0dc64e04fbeb9

Observation d962f3e3-e2d6-43fc-bc9d-f23beb0206c7 · outbound

This paper cites Don't Use LLMs to Make Relevance Judgments.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Don't Use LLMs to Make Relevance Judgments

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.994461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:7c9ccfe36a07301460d080d64eee2c03b5df2906797f317cb55b3415a62e1a54

Observation 98a09ffe-cc34-4299-8b87-cb3e37423673 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.009796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:09146a18c4a7c42a597acc4b1430819fb49ad56671ab3c6a5950879fdc2b63d1

Observation 2cb13816-cab7-41f7-a857-1fdd14102449 · outbound

This paper cites ISBN 9798400704314.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench ISBN 9798400704314

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.358733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:6b35738acf1aa222d1ea331d97d280d74290ef5269589d24174d07a669d2d837

Observation 7e54429f-3d8f-4a0b-8fe4-de845fd7585f · outbound

This paper cites AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:24:32.777586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:1a4832e997f5355d23b1d7ad5cfee96bf0c62b27a0522a595312b7b1842008d5

Observation dd33ece9-5654-4f19-b8dd-fa39584858ba · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.004745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:0d2e46b898bad10985109feccf5b166c00ec5c81177fb8b3414360b1884d8261

Observation 0692cdc7-69fd-4913-adea-25be4156338a · outbound

This paper cites A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.997039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:96d7a2379cb4b6d211f70a044782395606209cf09ee584e7796d85e2785ac92a

Observation e4370c7d-19fe-45f6-bbb4-5d9bb82833a9 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:24.351219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:e160e9e84c28bd0ecdfa8124cd4ee3745656b7b8638e5cc4ee4e86df8b152052

Observation 85ac5172-c564-4a96-a62e-0131fc66c5d6 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.987380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:138302faca4ab9fc8d4f89225b463f98621e83ef6a802f224e58795e1039619a

Observation 0f1a841e-dc10-40d0-b40d-52b4d4e22e49 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Survey on Evaluation of LLM-based Agents

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:02:25.007134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:05a693dad4ac1407fce69c53b9cd360e1dd472bd67cfd8361083d0b5e33c8168

Observation dada0be9-080c-4b50-bfc1-8021b455ae3a · outbound

This paper cites Large language models for information retrieval: A survey.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Large language models for information retrieval: A survey

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.989481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:69b8a140f5f65218ab7b4fb3dd4434bcbfe63202ce05e78c349d2ca7106c4a6f

Observation 24978f4b-1c75-49a6-9540-898137b7599e · outbound

This paper cites A survey on the memory mechanism of large language model- based agents.ACM Transactions on Information Systems, 43(6):155:1–155:47.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench A survey on the memory mechanism of large language model- based agents.ACM Transactions on Information Systems, 43(6):155:1–155:47

Reference 31

Resolution
verified exact
doi, observed 2026-05-10T08:02:24.361394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:a51f27bd7f168b32016ea0c624181220d8db7a7c087fc7105e4a4a2dce92feea

Observation 5b395675-7da1-40a3-9490-314c7bb298d0 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.240297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:5ea986faed4597a57c528c53f7a4832e36c4cdec8dbdc5bf8bb2461548cd54f1

Pith citing papers

Observation 33f66bf7-ec2b-46af-b59e-761d2af92727 · inbound

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP cites this paper.

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T07:06:37.817238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T07:06:37.817238Z digest=sha256:8761939c580f94fc9554b6728b61872b3d3c3b64679a4aa164dbfedbe924e036

Observation b504a754-8419-45f3-a8c2-f4915f885fee · inbound

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP cites this paper.

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:04:24.398577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:04:24.398577Z digest=sha256:ef392bb38ec83649979728614aee58de4250d8e20c8a3156b07d90ba547510a8

Observation 6aa18014-c06c-4dd2-88ea-ea98b63070b6 · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:18:58.706418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T04:18:52.865771Z digest=sha256:537b78623999b7a72678722e3c16c18aad9c87450ce957951afdd755c732b567