Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2310.19736.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.695574Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound
A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound
The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound
Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound
Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound
Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound
Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound
Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound
The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound
Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound
Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound
A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound
Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.