Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:02.319589Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2509.04013.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:02.319589Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:08.844007Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T23:45:07.998622Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 128e5b78-77f7-44b3-8ba5-540d8af3ab63 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Program Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032f5bc1-5851-425b-9f95-bae93273db73 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e28eac2-5474-4fa9-9b64-922e1d3182a5 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Bailey, N
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 766cc208-33f1-4b9c-adb1-e312d203b2bd · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3b2f874-c6b9-4767-9bdd-7e6b6e164b3a · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Burnell et al
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c5a3774-d159-438d-8764-c88fab8a9d99 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, J
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 555fd79b-d022-455b-9f0d-cdfb09ccc4a0 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0952470b-d78b-4187-b02c-03e45a157225 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73abcfd4-d6c3-4275-beb0-e0e23c633649 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa1b7741-f64a-4045-92a0-ab886997637b · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Clark, K
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cad664e0-b2a2-46bc-b0f5-99639c46deb1 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97370bb6-f807-4ab7-86d1-458051a8a23f · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Training Verifiers to Solve Math Word Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e54834d-57eb-4a74-9042-3cf4dbc2279b · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Frohberg and F
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b89a67e-665c-4c54-b7c8-78dc26687eee · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Guiver, S
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ae1f1f2-014a-4992-96a9-b8f3e2c3a9ab · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 094ceee2-e613-465d-9876-fff9a937eda6 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Hendrycks, C
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04a26570-1182-4b11-bdd9-7fe9326ba056 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1701abe3-569f-4e95-ba65-170ea1c8c103 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kim et al
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f47fa918-c496-460b-bb29-1aa8ff18da21 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kojima, S
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0e77fad-3de6-4f99-b2c2-a5e45e072299 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 25a96312-2b0b-475c-91b2-649f60b911e9 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating the Robustness of Analogical Reasoning in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887ea66b-3234-48da-a262-04da90ae8d1f · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f0c23a-571e-4cbd-904c-acbbbaa255fd · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Lunardi, D
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96ccef58-4bca-4a9b-808f-ac8f72575dd8 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mihaylov, P
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1d0c527-9588-48b4-8bb9-4dbb8515e46e · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 226d59b7-7fb1-4c96-9e39-54c25d67787e · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33a6f282-cfff-4201-bc73-309166eb850f · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb30fad6-2087-4804-86dd-a88a403aa5df · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Ouyang, J
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd6581ea-42a0-4bde-8aa1-16a34cf5048f · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Variations in Relevance Judgments and the Shelf Life of Test Collections
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1928cce6-36f7-44eb-9b8a-6650c08abc75 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a6e8c2e-a0c5-4393-8137-486c69e9eee8 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Reuel-Lamparth, A
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9de30eba-ef39-4bd6-85ed-2ca5ebc5ff69 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sakaguchi, R
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8644bea2-f5d8-45ff-8543-8e3c73b0cf41 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f473b916-536e-47ec-a003-aa68c32eebdf · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sanderson
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c31df391-90e6-4fb8-a8b9-f824cca61ea2 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sclar, Y
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b87919ff-4415-45b9-9be5-d6d8977c26cc · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sparck Jones and C
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b27103e1-f069-4ad6-83b9-9fbcef21a8a4 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49213a02-9fdd-4668-a7b4-068e583792e8 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3053029f-53cd-4dbf-b20d-ef2e25be86d0 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ccf09db-f904-40b5-b722-508a892310dc · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d473cba-0ea7-4dcb-88e8-b00755e160ae · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Finetuned Language Models Are Zero-Shot Learners
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b7c9c3-2c39-49ea-bc90-c7c1ebba685c · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a247aea-7324-4277-b3d6-86a6d3e39e04 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Emergent Abilities of Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826811b5-17ba-4d78-a2d8-1eebc9fcec3b · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Welbl, N
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd8639d5-3661-49c9-9ae8-aeddd79afaae · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e282fa3b-bb90-4bd8-b2d2-6c7cae47a607 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zellers, A
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc8a2d74-d951-4f55-ac91-916e6b2b04d7 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c487dba-c445-4e33-8390-0e3b677644ce · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zheng et al
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01dc6a9e-3429-4a80-ae43-94fbd1f422be · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86795f54-c045-4a17-81a9-74e30fcc2ffc · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61763fde-4f72-4bea-a209-5e3cf9f54dc0 · outbound
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd860ed7-1571-4add-9e0d-9db5135b3653 · inbound
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7bd56d7-f80a-4bfa-be9e-c15e06bc3b09 · inbound
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9447640-1d57-4f04-bb6a-4e4ce0214ec9 · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86bb03c9-9819-40e2-95cb-6390ff59ae8e · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e54f074-d4a3-4308-8237-8717c9608525 · inbound
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b0a32282-5a01-4201-ade7-b59c79931a23 · inbound
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cacb1980-62da-4e63-b519-bb6e4d88023a · inbound
MAVEN: Improving Generalization in Agentic Tool Calling On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9caf502-53cd-465a-a0b2-1d1934c03339 · inbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.