Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models: A Comprehensive Survey

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2310.19736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19736 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:35:36.355786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.777329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:3d064952abe8180d2d9658178da6a7c1c06c910ae9d935e5d25d1bc64a807410

Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.614193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:b2f9eb67f67049dfb78c9786c444d297723325faea401769cf7c732416590276

Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:37:36.032998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:d60dba238c8375f3bf894eefce5a0b09d945db2d67c3a5bc2ea879495a5c0b7a

Observation a24a0a9f-4ac6-4fea-9d7a-f8c312a689e8 · inbound

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact cites this paper.

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-08T05:35:36.355786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:35:36.355786Z digest=sha256:bb7408552ab63e58769fbc2eddc25413bfa4bf8af0955bbdb7ecc49c139b0a61

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:f2fffc478ae33f575c5efa7b0716d6624e81b560e70e9217b9d71467628c6ad5

Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data cites this paper.

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:03.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:03.725406Z digest=sha256:2280c5b403ec108ac7bb5dbd5ab1ab940ed64b0fb39c18e7bc7e132bdb269441

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:198e3673957177d4f8fe41183eca3ffe70a436038331531888cfe381ab1c82f3

Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.263848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.263848Z digest=sha256:9ebe699fe2de94bc7c8355a807217f5115bf50f47acc0a3b3b1c9d3ae18f781f

Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.624960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.624960Z digest=sha256:374c088f9624fdc44f85b8bd7e7c97b8e91a230181bacd2d01a1d1a4ef706a38

Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:8b20b0cd20b0db2917186434b6ce75ae9e354c5038800b5d9b64a3c2b901b093

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:b169adca9a593ff5e69268cc02d9a030ff740ad55d9f8052ed526c548ee754d0

Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:39.388006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:39.388006Z digest=sha256:525bac89a87643ea8d20f55c3b7feef60dec5d996cc93c264c1359dfe60868fc

Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:31.666043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:31.666043Z digest=sha256:e3d6376805bc2c8c1b32f38be075ca7d1545ed9cc0eadb2a517b4e79670dc1e3

Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:34.962660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:34.962660Z digest=sha256:6487becd4a23d14a6cb46fbadf8e6d8669a195560f7cdf4d00aa046389df7d2b

Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.530237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.530237Z digest=sha256:f3e0998d6f0cdc971f491764166ac22e4bd74cd311b8fe389eb3fe875f86a515

Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs cites this paper.

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:57.803410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:58:57.803410Z digest=sha256:987836632e58330f6dcfe8f5fe6536f2bfd281caefa996125966cc74abcd03f6

Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound

Cognitive Agents Powered by Large Language Models for Agile Software Project Management cites this paper.

Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:59:24.676630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:59:24.676630Z digest=sha256:5d147eb54f5a63dd90b8e79e8c84e736906d224cb5c577ba03f6a2c54f1816f7

Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.813109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.813109Z digest=sha256:19cb652640954b92618d677fe65013da16e13bcfef9b3c788182874dbc8f51b9

Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment cites this paper.

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:10:18.500779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:09:29.669360Z digest=sha256:893fba139b921d5e315b76fee2e9eb7808bf428b076246af0f9921604b893aba

Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement cites this paper.

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.758821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:53:10.245859Z digest=sha256:b93620917234a0efbcacb9bec622faab8dd2383da54893825a36562efc965826

Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound

The Generalized Turing Test: A Foundation for Comparing Intelligence cites this paper.

The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:26:18.923707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:25:48.145137Z digest=sha256:b69b2a360245692d2d9884b13cfc4cb7bc18a27c4ec3a7754e90ca5bac9fbdda

Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.749420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:24:16.157548Z digest=sha256:dba42a46c884cc24a2a16fe9e2946b556961bf22c9358f1ac32078e3eb1cf810

Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.695669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:44:52.491280Z digest=sha256:9cf71db6505b91c345e2759542c0ff93a02b49f75d179275be43cdeaf048cae5

Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation cites this paper.

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:29:03.254380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:26:13.698051Z digest=sha256:b10ab29dfe8d008ff251647f1e74f1df676b614681d28fcd14dc10ce5b239bd9

Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts cites this paper.

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.882307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T12:13:12.018506Z digest=sha256:4d7794f718c999ae5853af771ff5abb51dc98d9d766ce11cf77e59b3324c25a3

Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.500046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:3c59510b8402a64beb684470916a93dfa6e669d2b5b7f808e677881f901c089d

Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound

Efficient Sequential Evaluation of Large Language Models cites this paper.

Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:08:18.676413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:08:18.676413Z digest=sha256:c993b8028bbc205cb86d1a824d01b51f895ca5bafaa2686da2406b6d5b618b2b

Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.521574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.521574Z digest=sha256:52a76e58e27d2370863aa1dad0b71190551f04ec38fa647364da758419017200