Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models: A Comprehensive Survey

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2310.19736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19736 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.695574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.777329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:a099376dc393cba0589a1101287d2cb9ac9b7a0dc48b2481d20c618df47af825

Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.614193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:c4b41f17b30fe428051fb3deee83394d5df2d8eb8cb57abbbe9353118b6d7465

Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:37:36.032998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:96120424cf7cef94b42a620b6b73ab4df143dc2f75ff6c6680ad09baf199371f

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:96a4617e016965e283bcb7378f63fce9dba7afaff713ee38e842be946b723089

Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data cites this paper.

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:03.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:03.725406Z digest=sha256:2280c5b403ec108ac7bb5dbd5ab1ab940ed64b0fb39c18e7bc7e132bdb269441

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:95fabcbfa5868c32724f754f57c2ec800eb866993cfd8110f63832a9d601404a

Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.263848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.263848Z digest=sha256:9ebe699fe2de94bc7c8355a807217f5115bf50f47acc0a3b3b1c9d3ae18f781f

Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.624960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.624960Z digest=sha256:374c088f9624fdc44f85b8bd7e7c97b8e91a230181bacd2d01a1d1a4ef706a38

Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:8b20b0cd20b0db2917186434b6ce75ae9e354c5038800b5d9b64a3c2b901b093

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:082c7d9ba5408d8b3251b9a0917f4aaad32fa3deffeb79196ca11dddadeeadfc

Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:39.388006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:39.388006Z digest=sha256:525bac89a87643ea8d20f55c3b7feef60dec5d996cc93c264c1359dfe60868fc

Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:31.666043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:31.666043Z digest=sha256:e3d6376805bc2c8c1b32f38be075ca7d1545ed9cc0eadb2a517b4e79670dc1e3

Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:34.962660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:34.962660Z digest=sha256:38a340ab433f3c38ef86ac90cacfc9f9dd4f0b34ca7d72b18a487077943f79fa

Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.530237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.530237Z digest=sha256:4a6c6e874876f2d7e4a5c1c5a0da526cd5e5fbd5b32498869c374c8da685f9bd

Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs cites this paper.

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:57.803410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:58:57.803410Z digest=sha256:987836632e58330f6dcfe8f5fe6536f2bfd281caefa996125966cc74abcd03f6

Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound

Cognitive Agents Powered by Large Language Models for Agile Software Project Management cites this paper.

Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:59:24.676630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:59:24.676630Z digest=sha256:50b370ca377ecfd87d045f2b77129394d329b21833c183718db2156df250e7b3

Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.813109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.813109Z digest=sha256:19cb652640954b92618d677fe65013da16e13bcfef9b3c788182874dbc8f51b9

Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment cites this paper.

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:10:18.500779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:09:29.669360Z digest=sha256:0360cd3bce8efee7fa2d8a55637d3b2bf6daacbc75c1a031f0759d9b2ea8f567

Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement cites this paper.

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.758821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:53:10.245859Z digest=sha256:d880f4f2e20ed5a27d8b3443fbaf637035bd5664c58024c58a3723eb8483fc25

Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound

The Generalized Turing Test: A Foundation for Comparing Intelligence cites this paper.

The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:26:18.923707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:25:48.145137Z digest=sha256:a69440bd8a5c7b092ea9c6ec8f0e6767fe62d41ce0887467d796b28af03cd556

Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.749420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:24:16.157548Z digest=sha256:6e37eb92d187383dc4580be6457a9dcb6866d5cfd392ca47625dc511e07e2b13

Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.695669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:44:52.491280Z digest=sha256:0c1e5807b98d79b2751fbd6a3c7ad9d3f4d77fdea3d4e7a07777927a39f58faa

Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation cites this paper.

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:29:03.254380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T21:26:13.698051Z digest=sha256:d1b8f814006b723aa3ba5d97abb1f26dce0f0a50ee4bcee047a7a69152d3b354

Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts cites this paper.

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.882307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T12:13:12.018506Z digest=sha256:be538e44b856242584ecd5c2aaa784ba07aa49c9ce40a626ff123255388a7c51

Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.500046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:1f5c0c5ad13102c2cb235710e739929c3d43c83484d9be40fe9f79df9faec12a

Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound

Efficient Sequential Evaluation of Large Language Models cites this paper.

Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:08:18.676413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:08:18.676413Z digest=sha256:c993b8028bbc205cb86d1a824d01b51f895ca5bafaa2686da2406b6d5b618b2b

Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.521574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.521574Z digest=sha256:c244f9964d7f3047b788c8bd881688ba43f7e1362fc6e0d31e1f3492932512fb