Pith. sign in

Paper Citation Record · LEDGER

An End-to-End Agent Auditing Engine

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07346 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:52:48.176696Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5625cc56-5b20-443e-bde5-a3cc049f14d4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

An End-to-End Agent Auditing Engine Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.085855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.085855Z digest=sha256:dd960be22a0ee42f4cdbfc91e41e83e5aaf93ae7be1595d88cb7e5ec3f821bde

Observation 1a54921d-cc84-44e7-8ee1-3c99ee4bfb3d · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

An End-to-End Agent Auditing Engine SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.098487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.098487Z digest=sha256:f680671067ab10c43c63ec99f78091a4b6fd5be917e0f750661207407332c8bb

Observation db198a5b-35bc-4d38-a853-3cf7de2b12b1 · outbound

This paper cites Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI.

An End-to-End Agent Auditing Engine Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.102698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.102698Z digest=sha256:14864ba3f2f0eb67537bbfc2e4b743db252e0b6ddc1fe86dc139e3c49c0da67c

Observation 3eb138b2-b36f-4c69-8065-0659313dbba7 · outbound

This paper cites AgentQuest: A modular benchmark framework to measure progress and improve LLM agents.

An End-to-End Agent Auditing Engine AgentQuest: A modular benchmark framework to measure progress and improve LLM agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.539262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.105912Z digest=sha256:170e779e05d6501817b8cba58d8c004d9c905fdc88fe71255fe14e9fe8905e95

Observation 65bdc42a-287d-4d87-bf08-93c141cd6b4b · outbound

This paper cites doi: 10.18653/v1/ 2024.naacl-demo.19.

An End-to-End Agent Auditing Engine doi: 10.18653/v1/ 2024.naacl-demo.19

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-10T05:52:48.108994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.108994Z digest=sha256:2dcbf0a237f5cf917787ace3630090dfe4880986b61d3e946f98022f3cd4f56f

Observation 04d89948-296e-4634-a31b-577117c37761 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

An End-to-End Agent Auditing Engine Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.112150Z digest=sha256:0918e5ac435dbe6dba463576e7c1f14328fa116a0cce9f47e2bcc6f6d3a93b5b

Observation 408085aa-2f01-4960-bb53-84efac2b1c56 · outbound

This paper cites Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.

An End-to-End Agent Auditing Engine Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.122129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.122129Z digest=sha256:42f8b74267e7be44b5dad58dd70ba56d68a2f983f7479e7537ba9ed58a65ab3a

Observation 3817d661-2696-4aea-93b4-b902cef122b0 · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html.

An End-to-End Agent Auditing Engine URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.128965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.128965Z digest=sha256:194d4a6c7dc45a62bc972c3a442a50b2855e7a2482d31d49b960deedbf2466db

Observation d5c090e1-e6ea-44b7-b0d9-017520822c2a · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

An End-to-End Agent Auditing Engine Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.136857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.136857Z digest=sha256:c95e1662b42b27a01e76b7bc289b392e9a9b9eddf0009bf72fceb247a014f88d

Observation f3c75470-fc93-44e4-9e36-be13b64de80c · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

An End-to-End Agent Auditing Engine GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.144277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.144277Z digest=sha256:cbdc371b25c70b4c3a804e84596a7e3aae6e6b64261fd502c4f1739c4f2e5903

Observation 139e6e5a-fff7-4f4d-87e6-4455a1230951 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

An End-to-End Agent Auditing Engine GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.147679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.147679Z digest=sha256:90b3ce53e016cc3fe35af34689f19aa855a8df35cb2378a8d02906dc022b2f7c

Observation 1b5f96d7-880a-4d91-b219-8ca16ce2a11f · outbound

This paper cites Challenging big-bench tasks and whether chain-of- thought can solve them.

An End-to-End Agent Auditing Engine Challenging big-bench tasks and whether chain-of- thought can solve them

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.528742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.154853Z digest=sha256:147021bb3f37ee49e186af1c32239245de8e99cc302007453bef0bd30c71c909

Observation 11cf1d65-33b9-46f8-8219-5cb159b44821 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

An End-to-End Agent Auditing Engine Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.518772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.158585Z digest=sha256:67e4355ccdfd5d1b2086c49cf8608a66a5b928b479ea976427645de4fe7ec216

Observation dc862431-79e6-4274-bd6b-2e62e90a1f64 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

An End-to-End Agent Auditing Engine AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.161998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.161998Z digest=sha256:43ef8cee19e7e50537a242b869040f51bfab455385242ccfde5f4e1154584f8a

Observation 29cd5a68-1edd-45f6-b352-5db9ba220843 · outbound

This paper cites URL https: //aclanthology.org/2025.acl-long.1355/.

An End-to-End Agent Auditing Engine URL https: //aclanthology.org/2025.acl-long.1355/

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.165568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.165568Z digest=sha256:06115cb0b841295f90ebe2de93597caa47f3a2a86922b1604e70c2435af3c310

Observation 4cd9036f-7a42-413c-be52-d7ab72ab06a0 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

An End-to-End Agent Auditing Engine $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.169156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.169156Z digest=sha256:5452c71e9b98f66947ea4d14068199aab9d52140a2b6609484b8c67235c536e3

Observation 967a22cf-b289-4992-9ea8-4cdfa8037e74 · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

An End-to-End Agent Auditing Engine Agieval: A human-centric benchmark for evaluating foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.507977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.173232Z digest=sha256:153a05ea5518ae1e21b06ce3d36340a7fd4316f43087b8767257b4f609ac3386

Observation 1193309a-467b-4d65-9667-62c9f12175ac · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents.

An End-to-End Agent Auditing Engine Webarena: A realistic web environment for building autonomous agents

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:52:48.497017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.176696Z digest=sha256:01333aa8114f24881fc79c8c658efb7d90b6ad832abe4649dc212ab4c5ff4fbd

Observation 539b6acb-6091-478a-b9be-ec9b3644bcb4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

An End-to-End Agent Auditing Engine Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.094149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.094149Z digest=sha256:8fc07d55592f824a269d6893c68eed62aec78b420907addaed560952c1fc2619

Observation 0348cd40-320d-423a-bdfc-bc61eea4c093 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

An End-to-End Agent Auditing Engine Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.115676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.115676Z digest=sha256:9cc3967eef8b014aec9dd54c40fc64e83bd72f8f11ea645691956ce23829c89b

Observation fc4b48c1-a2ab-40ad-b5d2-933213c61d98 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

An End-to-End Agent Auditing Engine Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.089974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.089974Z digest=sha256:ad1240c5f865586a009ef0f0de922d85af966abdee461549740febfd488ba758

Observation ed8b3f05-daae-4ebf-92d7-bf465ad0e3ba · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

An End-to-End Agent Auditing Engine TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.125316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.125316Z digest=sha256:7976b9f6ceac895f301f0402002389c48a8eae6c836e5100e135239f8960415d

Observation aaeede7f-642d-4d82-8cfc-e039574f507f · outbound

This paper cites Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents.

An End-to-End Agent Auditing Engine Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-10T05:52:48.244282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:52:48.151220Z digest=sha256:4b5ef13911a98efab7119e529fe99957d4240b49cac6078d4b4bc4666ddb9313

Observation a8fbcf01-2c69-4ddb-8a7a-dcadbdc527ce · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

An End-to-End Agent Auditing Engine $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.081468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.081468Z digest=sha256:bb06e398348c808921ef4d77676ee776d90e11e844cebf331ac2c1832a274a94

Observation f020150e-188f-4d8b-bf9e-c106e2db9075 · outbound

This paper cites General Agent Evaluation.

An End-to-End Agent Auditing Engine General Agent Evaluation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T05:52:48.077579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:52:48.077579Z digest=sha256:1619739f8afe2e8aee368bee360a3d595632340541d12602ce73e73a42cdceee

Pith citing papers

No inbound Pith citation observations are available.