Pith. sign in

Paper Citation Record · LEDGER

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2509.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04013 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:02.319589Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:08.844007Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:07.998622Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 128e5b78-77f7-44b3-8ba5-540d8af3ab63 · outbound

This paper cites Program Synthesis with Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.079925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.079925Z digest=sha256:30f4122cbead1213fb3539bb95aca32e98951ae957165634a060a3bd575d1876

Observation 032f5bc1-5851-425b-9f95-bae93273db73 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.085778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.085778Z digest=sha256:e6d692950fb6849cc65eb353c82971361e0a8bfcb51cbebbe8cb4df8db989c9f

Observation 5e28eac2-5474-4fa9-9b64-922e1d3182a5 · outbound

This paper cites Bailey, N.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Bailey, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.281185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.090860Z digest=sha256:1d5773c7bf43bef7456d6a1b562d646a9d28fb2c47b14fbfb35aff7c35fe8ff8

Observation 766cc208-33f1-4b9c-adb1-e312d203b2bd · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.267003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.095785Z digest=sha256:18fc16289840eabf0daeb6b4b4c414e3532fd42aee524c33b136d684919eec00

Observation a3b2f874-c6b9-4767-9bdd-7e6b6e164b3a · outbound

This paper cites Burnell et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Burnell et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.252544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.101619Z digest=sha256:4efbd9fa3026503ae4f52de49207138ab3215f0dea998956b5a04ab816e78632

Observation 5c5a3774-d159-438d-8764-c88fab8a9d99 · outbound

This paper cites Carterette, J.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.237912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.106197Z digest=sha256:2ef976dca9902099108e76c13c942c240af7bb3f24ccf8c6bc72825d4b9c61f9

Observation 555fd79b-d022-455b-9f0d-cdfb09ccc4a0 · outbound

This paper cites Carterette, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Carterette, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.223737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.111083Z digest=sha256:a27bcd0373576a01eb63482a48bb6bf83606bc87f5f71148c564132c047435c0

Observation 0952470b-d78b-4187-b02c-03e45a157225 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.115716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.115716Z digest=sha256:908d84eb8252559bd22b992ceea64b63931df0180c8674669f436dd9d60ee4f4

Observation 73abcfd4-d6c3-4275-beb0-e0e23c633649 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.209510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.120756Z digest=sha256:0a9181190abe603fa3cb7c13f4e17f25be7f5e942cd2ea39e75624e404027a91

Observation aa1b7741-f64a-4045-92a0-ab886997637b · outbound

This paper cites Clark, K.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Clark, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.194167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.125172Z digest=sha256:82851c0cd34308d0cf10ba81367e1e8a7dc9cf935e6005824c722adf033737b7

Observation cad664e0-b2a2-46bc-b0f5-99639c46deb1 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.129784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.129784Z digest=sha256:a0bca55ea9b995b1d7b2d12f47a180f21d80a939821ab989cfcd3a5c0e316fe2

Observation 97370bb6-f807-4ab7-86d1-458051a8a23f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.134983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.134983Z digest=sha256:c50911a154057bbee51b99343dd26dbd5160ae8115078a6adc97d28ff0f8e74f

Observation 3e54834d-57eb-4a74-9042-3cf4dbc2279b · outbound

This paper cites Frohberg and F.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Frohberg and F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.180081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.140039Z digest=sha256:98c3e453a16dc082ad81eefdca99b3ad85bfd3d21675a35b390076684e9a0ebd

Observation 2b89a67e-665c-4c54-b7c8-78dc26687eee · outbound

This paper cites Guiver, S.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Guiver, S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.166258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.144420Z digest=sha256:e9aaad70a62ca560fbbbeb9ed07a72ba72c8fe86d5fda18d484e3d3b4dd931b5

Observation 1ae1f1f2-014a-4992-96a9-b8f3e2c3a9ab · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.152235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.149440Z digest=sha256:f4dcec6e07dceed6bce2c5320093c55016064cd9e71fe309409c8d2048289dd8

Observation 094ceee2-e613-465d-9876-fff9a937eda6 · outbound

This paper cites Hendrycks, C.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Hendrycks, C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.137847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.153996Z digest=sha256:b0b8661734f8c43f0c2834c2aa5b04b70398e5a01867746306414a952fbbef35

Observation 04a26570-1182-4b11-bdd9-7fe9326ba056 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.158497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.158497Z digest=sha256:91e9590e6be8c6fbfd0879a84556a79f79e3270cf1897ddf932993f135989138

Observation 1701abe3-569f-4e95-ba65-170ea1c8c103 · outbound

This paper cites Kim et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kim et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.124043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.163069Z digest=sha256:4f6fa1591a56bea3d1544154e02a83c6f786bffecdd73fed049d9bb5fe018c9e

Observation f47fa918-c496-460b-bb29-1aa8ff18da21 · outbound

This paper cites Kojima, S.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Kojima, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.110964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.167804Z digest=sha256:ae0df3357fea6cce1c787b768c3e732bd28dab1bd0465bf63fe2afe392396d3d

Observation d0e77fad-3de6-4f99-b2c2-a5e45e072299 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.096351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.172823Z digest=sha256:beaa8523869b90c8f5f02a0cd4803c530433a6a7db5bcc905ed023aa7cf6579e

Observation 25a96312-2b0b-475c-91b2-649f60b911e9 · outbound

This paper cites Evaluating the Robustness of Analogical Reasoning in Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Evaluating the Robustness of Analogical Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.177268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.177268Z digest=sha256:95d09b846e22bc38da2711c69b63d03f801d0d8e92c79a113a515f06271dd751

Observation 887ea66b-3234-48da-a262-04da90ae8d1f · outbound

This paper cites CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.182126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.182126Z digest=sha256:5336d5c7f4f4409ad119c5dde47b636a8644e3a703ae102b5292321832ffa2da

Observation c0f0c23a-571e-4cbd-904c-acbbbaa255fd · outbound

This paper cites Lunardi, D.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Lunardi, D

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.081688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.186580Z digest=sha256:686808ed01d63a429d872c699a3780dac6cece07e387830d7e81e0508c170ae7

Observation 96ccef58-4bca-4a9b-808f-ac8f72575dd8 · outbound

This paper cites Mihaylov, P.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mihaylov, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.066683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.191311Z digest=sha256:742f691dce2dcd7828109b53b772d78e3467526834c9799a42eb17051df310cb

Observation b1d0c527-9588-48b4-8bb9-4dbb8515e46e · outbound

This paper cites Mitchell.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.052407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.196651Z digest=sha256:dd9b4cc483afa73c92c620190a2c8fa93152423b357e0d4802597c60608353db

Observation 226d59b7-7fb1-4c96-9e39-54c25d67787e · outbound

This paper cites Mitchell.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Mitchell

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.038500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.201503Z digest=sha256:a65388f4f5378f9e613563c2cb81f091ff86f9d91ef6efc43d475cc3547bee6e

Observation 33a6f282-cfff-4201-bc73-309166eb850f · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.206347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.206347Z digest=sha256:0df8e9963bc56da6a2b62b41235f93ac9eeac739954cbacae99a04d7634acde8

Observation fb30fad6-2087-4804-86dd-a88a403aa5df · outbound

This paper cites Ouyang, J.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Ouyang, J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:03.024374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.211401Z digest=sha256:714d84d811f53bfaed120d983c92b2e992566c0b5f39e7502e9debf639962164

Observation cd6581ea-42a0-4bde-8aa1-16a34cf5048f · outbound

This paper cites Variations in Relevance Judgments and the Shelf Life of Test Collections.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Variations in Relevance Judgments and the Shelf Life of Test Collections

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:31:02.633586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.215986Z digest=sha256:1c950bfae639cb254d83ac532078a381e22d7c5a97f45c5a61d93d232e805fa9

Observation 1928cce6-36f7-44eb-9b8a-6650c08abc75 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:03.010228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.220869Z digest=sha256:c58c7542e62c9534795796e81a2ea43647e67e13624ada1a0471fcf84070e67f

Observation 4a6e8c2e-a0c5-4393-8137-486c69e9eee8 · outbound

This paper cites Reuel-Lamparth, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Reuel-Lamparth, A

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.994446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.225854Z digest=sha256:01dc616eb0c0254be246a65b5ba35534c451ef7352342c579992fab76f8ee85e

Observation 9de30eba-ef39-4bd6-85ed-2ca5ebc5ff69 · outbound

This paper cites Sakaguchi, R.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sakaguchi, R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.979800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.230285Z digest=sha256:aa75b56cbb65be9cadc20c49a95c051d2e49a9c6e4865eb60c48f3cc7d56f4fb

Observation 8644bea2-f5d8-45ff-8543-8e3c73b0cf41 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.234717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.234717Z digest=sha256:104772be51c557930d67858111281e71c96ef0501b355a366308c9f1137e63cd

Observation f473b916-536e-47ec-a003-aa68c32eebdf · outbound

This paper cites Sanderson.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sanderson

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.964910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.238982Z digest=sha256:3c6cf3e6f79953d1fc87efa28b35fdb650b1066be7063968c123eeada2169cd4

Observation c31df391-90e6-4fb8-a8b9-f824cca61ea2 · outbound

This paper cites Sclar, Y.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sclar, Y

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.950379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.243182Z digest=sha256:45d037cb88a6c36e2d1c628b46e739e2dd66e75bcedabc572655e7cb1363a0a0

Observation b87919ff-4415-45b9-9be5-d6d8977c26cc · outbound

This paper cites Sparck Jones and C.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Sparck Jones and C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.935585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.248104Z digest=sha256:3e6f76e25af21c54b2629ea5915cc2ba04f1d0d8f8bba618bf674e992139b7a8

Observation b27103e1-f069-4ad6-83b9-9fbcef21a8a4 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.253189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.253189Z digest=sha256:fde2baf40ceca212c6b0f3c62343fb3cd7d00ad414014f546e307205470c3389

Observation 49213a02-9fdd-4668-a7b4-068e583792e8 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.920333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.257873Z digest=sha256:c75240e8ac89527bd2bd5d683af2dcbe9d5bc96be99d35b42ea3f1d2f1ef034c

Observation 3053029f-53cd-4dbf-b20d-ef2e25be86d0 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.903240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.262534Z digest=sha256:b828d04b8836d6af0a50d5b06efa5878c05dad65f4a54e0b4b54b5133c98d024

Observation 0ccf09db-f904-40b5-b722-508a892310dc · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.888353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.266985Z digest=sha256:0e2cf63bba98203f2e50db2d659cdbc8ddaf9dc57cfdeea6e3b093c6c97a67c4

Observation 1d473cba-0ea7-4dcb-88e8-b00755e160ae · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Finetuned Language Models Are Zero-Shot Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.271888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.271888Z digest=sha256:d22159c0e7e26e353bebbf0473a0fb4601cdbd1306a33b0416fe5a5527c795e3

Observation d2b7c9c3-2c39-49ea-bc90-c7c1ebba685c · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.873133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.276957Z digest=sha256:d64ca25923ced623694192e227b6bee1bbc1011327feaccca26119814a5caf99

Observation 5a247aea-7324-4277-b3d6-86a6d3e39e04 · outbound

This paper cites Emergent Abilities of Large Language Models.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Emergent Abilities of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.282896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.282896Z digest=sha256:a16a83b70160f0bf82b27ce1ac9b62df35c7251bade9f9cb48d15aff5c682b05

Observation 826811b5-17ba-4d78-a2d8-1eebc9fcec3b · outbound

This paper cites Welbl, N.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Welbl, N

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.856887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.287920Z digest=sha256:20421de4f3270f79cf8722a3fc93b87bc1350141f79cea5a765c86db3877a799

Observation bd8639d5-3661-49c9-9ae8-aeddd79afaae · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.842118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.292423Z digest=sha256:bc4219c0ec0e4fd5e9061ccc75967fa616375f78a118b227901080dd2da97455

Observation e282fa3b-bb90-4bd8-b2d2-6c7cae47a607 · outbound

This paper cites Zellers, A.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zellers, A

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.826115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.296921Z digest=sha256:a69d6ec23a4ed78789826622280bf5a2a6e5e66df93971affe7c306508f0f4eb

Observation bc8a2d74-d951-4f55-ac91-916e6b2b04d7 · outbound

This paper cites an unresolved cited work.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:31:02.811178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.301739Z digest=sha256:19df4afb414e671aa15effa355f9849f31f3bc1249b3457770595e826b467cb7

Observation 4c487dba-c445-4e33-8390-0e3b677644ce · outbound

This paper cites Zheng et al.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs Zheng et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:02.795648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:31:02.306286Z digest=sha256:0d962417795462a2bbcdda99029fb45538971fdc9969b88b31a86ec454343efc

Observation 01dc6a9e-3429-4a80-ae43-94fbd1f422be · outbound

This paper cites QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.310499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.310499Z digest=sha256:f1049b5931cb4595e8fb26e169f5532cf52c35672e59de2516639ac204faa1e1

Observation 86795f54-c045-4a17-81a9-74e30fcc2ffc · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.314947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.314947Z digest=sha256:5c757aa93aef22c1e9115c426d18a0b5976656d6dfa0c94816335664d2242976

Observation 61763fde-4f72-4bea-a209-5e3cf9f54dc0 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

On Robustness and Reliability of Benchmark-Based Evaluation of LLMs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:02.319589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:02.319589Z digest=sha256:81e6ed03370a700ec4c7478e38f065a21fc4f661788aa867ab21e32d22f4380a

Pith citing papers

Observation cd860ed7-1571-4add-9e0d-9db5135b3653 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:23:39.942350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:2d8972fec1fe9be8948ff811d7227c7cdf9b0121de597e0d3b1c74b0291b55ed

Observation a7bd56d7-f80a-4bfa-be9e-c15e06bc3b09 · inbound

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario cites this paper.

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:21:25.388199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T11:27:03.620795Z digest=sha256:19e3f5b7508dbccdb1d36820685a7ff6ae39893ccb8fe8549ed846cc27d9f55c

Observation b9447640-1d57-4f04-bb6a-4e4ce0214ec9 · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.932740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:1d26dbe74a5d2980d07103399ae919ba86fa81271c6d235223a135b425075d21

Observation 86bb03c9-9819-40e2-95cb-6390ff59ae8e · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.000489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:8a8a04a5397e7e9ce79af7b2c5a8d39b4e7f09fca0c96e71fb861c3589038c54

Observation 8e54f074-d4a3-4308-8237-8717c9608525 · inbound

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning cites this paper.

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.325279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:52:46.575173Z digest=sha256:0e2b86384d4f6ca64b1ad2000355a4d0353430dcd8e2ea304abe43026fb01667

Observation b0a32282-5a01-4201-ade7-b59c79931a23 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.568253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:e6bfba565060f63a7505c3ae9a8012af5ad9e8493c989df571b81f96ce0655cf

Observation cacb1980-62da-4e63-b519-bb6e4d88023a · inbound

MAVEN: Improving Generalization in Agentic Tool Calling cites this paper.

MAVEN: Improving Generalization in Agentic Tool Calling On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.135066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:6fbb3796e9801b66de7430233f8bdad1d04082621a9896b6bd0f0aee7cf582de

Observation b9caf502-53cd-465a-a0b2-1d1934c03339 · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.844007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.844007Z digest=sha256:76d986d6c899131467b227044c1d707fe7265bbac0cec553b3b2efc49451e8b9