Pith. sign in

Paper Citation Record · LEDGER

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2608.06202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06202 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved44
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26221a13-0976-4b7c-8ff5-babd1a283942 · outbound

This paper cites 2026 , note=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2026 , note=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.564761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.564761Z digest=sha256:2c3d7592652e3b68dbbf62a3225694483dc4966f73da6c1b3e928957e323ae94

Observation 622a2c7a-90e9-427b-bd74-9178d043a17c · outbound

This paper cites Findings of the.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.990676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.623981Z digest=sha256:3b0156ab7eb58c6056125e5cd3b7eff81ae7510caf7b004515a5c02146455938

Observation 462ca983-a498-4f75-b75c-16294ff53ff8 · outbound

This paper cites Proceedings of the 62nd.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.837480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.700906Z digest=sha256:26a2d601b58843a79b59812a6ce6256f11c676e3820825eccddaf8d51f3189b7

Observation dee58e1d-f0a9-4e4e-8c85-9406f8f71f35 · outbound

This paper cites Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.677459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.813047Z digest=sha256:94b451a7bcfe8e80b3d7928a980d7556ec2d49020df6eefe2ffe9ce717654a13

Observation 1d6aa5bc-2a07-47f9-8374-28c6e9e277b0 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.930407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.930407Z digest=sha256:f99d83fdb3a6e4374a053c1d8dcd41911cbf6d445fbc3cd47369d7398a1393e3

Observation 71b7280e-9268-4942-ae0b-d9bd554abc3b · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.547709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.026529Z digest=sha256:c20a20940a0240c64bd8ba8782f9a0aa8025af3da54a2f30570558da4f6a9331

Observation 60b402ce-04b2-4459-b5f6-e38590f0a660 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.091734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.091734Z digest=sha256:b57bbcec8e2f0211fe3d8eaa8f86b9dc80a93d8a3017d44ff04bee9e4792c002

Observation f4f4de7f-479c-4700-a2e9-6569fef31de5 · outbound

This paper cites European Semantic Web Conference , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) European Semantic Web Conference , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.332701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.198014Z digest=sha256:0db0e1c89eeaa3c1c1dfb4265d69f0cb26a689a61596cf1060d63af8b7e060bf

Observation 8c2e8861-ca7c-4252-972b-efc4ed7283d9 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions of the Association for Computational Linguistics , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.295139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.295139Z digest=sha256:8655c0a971a8cd7834b596799354027b8f57dc1e8edb9474258aecf0496ac786

Observation 94a50989-6931-4d0e-a9fa-d8b75f58de0a · outbound

This paper cites Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.383467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.383467Z digest=sha256:9ab24caa09a636edfc537c6bb3804c2b9aac2ee550ebe8aafe8739cddceec2a1

Observation 68e4a836-098f-4d8f-b437-4c400f68a726 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.454833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.454833Z digest=sha256:2449fd870a591501e05df137793075793e9f363e955146b6ea479f19b0420137

Observation 70309c71-8603-499e-99da-07f30a0a32c8 · outbound

This paper cites Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.985875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.497238Z digest=sha256:63903c00403716b058cf80af70628f6610999c7ee42cd74d1fc53d9197da5f6f

Observation cfb01546-9fc0-4fbd-97f4-7857da88a24e · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2020 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the association for computational linguistics: EMNLP 2020 , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.549221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.549221Z digest=sha256:3cb3554c44b89e4807000cfb462e4117d3bd2b9870d174f1e31ed601743d9c1c

Observation a7228e6f-6477-4bf1-98c1-836015154b6f · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.595044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.617451Z digest=sha256:fd1f96300aeb3c4902efd92d36f4be9629c6502c4f37b9f52dd4852a3898a23d

Observation 2d72a67a-258c-4fa6-8534-0578564f6b3c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.360053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.703087Z digest=sha256:2144160b3e8ccaf2e6616e16be708effee9243c319b9b908797fc5d810737359

Observation e8d60a6d-d40b-46f5-8eb6-d5ffcf91416b · outbound

This paper cites C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.795670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.795670Z digest=sha256:ac66980ec5674b7520bf457a8cf647ee443c2f3ab9478595d013b88c3e470543

Observation b47d6d3e-4761-4dc4-a5f8-a6b66cf1012b · outbound

This paper cites Transactions on Machine Learning Research , doi =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on Machine Learning Research , doi =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.241243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.882722Z digest=sha256:925baa39432d32a767bbaf64d1df838494a4516a2781b8168a9a2f33e20bbd72

Observation 8ef6a630-248e-4bea-aecf-078d979fbb64 · outbound

This paper cites ACM transactions on intelligent systems and technology , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ACM transactions on intelligent systems and technology , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.978703Z digest=sha256:7b6cdcd5577c7a43ca362c21ebfd9b896b449a95d28262678649e1dcfc6da775

Observation 7776f9c7-d798-446f-a516-ba187f2bbd87 · outbound

This paper cites Transactions on machine learning research , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on machine learning research , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.102903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.102903Z digest=sha256:c88c49d5c697d917b38b2b380694218415f94e6487612e4fd425c06e76595222

Observation 83be2d93-8c2c-4244-b90c-003d12d088db · outbound

This paper cites Measuring Massive Multitask Language Understanding.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.168261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.168261Z digest=sha256:d0903c5e895039d6f9cf3e5dfc28f9ee06a4c9b1c6f89837039d1001959aef70

Observation fe0054e4-ac38-4184-9bbe-44a120d787cb · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.260360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.260360Z digest=sha256:fbe0c91055df137473bcd4552c76715ee8ddc1e8764394b3aabfa31a78a0fd24

Observation 086b3cad-130d-45e1-8776-16f61cd85221 · outbound

This paper cites P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.369494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.369494Z digest=sha256:374c105726694d6f4ee335014a65b30ec8f6ff98c60174302e1ae5df2ba3cb8c

Observation f92216a6-268c-4b31-bd92-8fc6e0a4b4bb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.461520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.461520Z digest=sha256:2df42d22878ac01a77d0cf492f4d7f1138008db91ee635431668d2d1f3a9413a

Observation 1e24bf35-9c64-4a7b-aa78-09c43b5ead4d · outbound

This paper cites Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.861608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.550095Z digest=sha256:43d77b18732462ec544b9152ebe126b90f9bed9e2894990c12582fba3e5cffcc

Observation fe0a5809-6866-4f94-a79c-1da3448024a0 · outbound

This paper cites Understanding User Experience in Large Language Model Interactions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Understanding User Experience in Large Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.653851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.653851Z digest=sha256:eea99157f74aaa29f5f6592cbb249b54a23f42588c645f1813f5cc39ad5eed5f

Observation d46bb22a-b864-4e17-9d3f-89312334bcac · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.736893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.736893Z digest=sha256:161912070f9dd15c7c560a10066db1cdb571afafecfef9a051fa202ff3e95465

Observation a8b092d3-0c98-4656-96af-edda0a1640e5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.649164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.804536Z digest=sha256:60e289f76ab2ca5ee0c2ba915314d3c51e5ec67787f6b8e2919f978223589033

Observation 0c553b29-49b6-4375-aa24-bdf47c7ce632 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.856791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.856791Z digest=sha256:09e3c25b2ffb021c06651ff429e35682e191395d85979d4131cf0663ff98a4fc

Observation da414d8c-eecd-4044-b747-d975df546e44 · outbound

This paper cites arXiv preprint arXiv:2509.19364 , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) arXiv preprint arXiv:2509.19364 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.930656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.930656Z digest=sha256:70bc9370df0ffdbb99e134b018e521a6a0fcd65ed43b89fe829ae51d784a9306

Observation 2e3f3f2b-b54d-4663-a092-ccc87aea3b80 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:05.404673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.991547Z digest=sha256:3cdd1ebc9554e3a7e4406b801cfe6f635022e2c3f8b81d990d9f9643fe3f6bec

Observation 9a9935d0-f7b6-46d0-acd9-2b2a9276d937 · outbound

This paper cites 2023 , isbn =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2023 , isbn =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.071960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.071960Z digest=sha256:4e57f758880b4f059c69a1cc70e4a2f7a981ca907848e603f4c86b9a30788138

Observation 6161cf9b-f2c6-4a51-bb84-19e368a3ece4 · outbound

This paper cites On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.189426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.189426Z digest=sha256:250b07511c739a7667db1dc485fd34cbb1e1a09a5fa8a66f07d83c8b307a7ff7

Observation 9dccb30e-fa97-4f3c-9bda-fba0e3b7e762 · outbound

This paper cites Ask Again, Then Fail: Large Language Models' Vacillations in Judgment.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ask Again, Then Fail: Large Language Models' Vacillations in Judgment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.260166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.260166Z digest=sha256:2a203b07f4bf2f7b2303fc0407d1360252031e1437cd8faa12ab0a66906bc7fe

Observation c0b012f5-87a5-48b6-bcab-0b05fd694512 · outbound

This paper cites S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.316712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.316712Z digest=sha256:c5caee6567eee2199054543860f20cb6fcc9b5a2ddff260b2b6f2ba021f36a24

Observation b346a3ab-f126-4e6a-bc5c-5e514e1e25f6 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.409924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.409924Z digest=sha256:80bf67be914e20dc75e59ae347b9320f1ed0bfa806ed59470dedd65f0ef43a96

Observation 15f074f6-1ad5-452b-af55-9103cc4219d2 · outbound

This paper cites Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.503503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.503503Z digest=sha256:4df2c73fe204c5b44120ad0c6b2e7c44c8813209f2edfd533a3909e77b942a00

Observation 1313c07f-0d7c-4d4e-9d24-d81a7fb31217 · outbound

This paper cites LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.565332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.565332Z digest=sha256:3aa12edcd6a978c349788c546a756324ca366aa5bb95ef4954149cd4ea22d45d

Observation a458f07a-0ffa-4f31-865f-231be2c48040 · outbound

This paper cites Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.151838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.623847Z digest=sha256:f991cf29e9a8113cb475c33ce8cbb6093136a9ddf79abe5d0a535e493dc9e0bd

Observation 96133c27-3520-4a5a-9715-6591749ebe40 · outbound

This paper cites Practices for.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Practices for

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.934779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.700820Z digest=sha256:c03a2440f524fa5f70f4f124332b87aa6747ae757927cc7cc79d6d2afdf2a375

Observation 856bbd95-784c-4347-abaf-daea68706d34 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:45:04.655486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.816885Z digest=sha256:af12b3ae152e4e78a6e8f662c2b25eaa5c7597ba9d140b8bf8da574e8dc54cc1

Observation f43044d6-d349-4364-bbd8-fb8fb3bedbbb · outbound

This paper cites Aligning AI With Shared Human Values.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Aligning AI With Shared Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.935769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.935769Z digest=sha256:3e025ca9dc53f3cf0454094fe0e20a2430e4ff08b8ace8859e0cece28807362f

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:4c32239ba21fdb5502e764327c461c340e5abffceefebd40413df4a781e1a59f

Observation daa23e92-20a0-4295-a933-fcdb476df63d · outbound

This paper cites Science , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Science , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.125805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.125805Z digest=sha256:f04c21d5dfb2a92689a5d26a71a7584e03a2643cf97ab00a0c28802b2c51e2bd

Observation 0ca37549-c4c5-45e5-970c-15f40d1808ec · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.931158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.232606Z digest=sha256:9d1704e8c444d353810b7f70f7b578e73a8e821f18e6a75c38aa4621b80171e5

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · outbound

This paper cites The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:9258102338f2d8a87100ce95a71f5cf6ca1de459179ae7625ebe6ec81f9ae35f

Observation 05b8b964-fc11-4886-a522-ddbc7f20ed3f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.393890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.393890Z digest=sha256:400f70be8048898a3c9d0152ea9b2f0eac7150bc0812d8a619785fb339603787

Observation 09ebe1db-820d-48ac-9d25-7ef4f3d17a2b · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AI and the Everything in the Whole Wide World Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.451657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.451657Z digest=sha256:eb827ebfae7c480783ae243a15daa4855f6aeeffce1ffdffeafbe23f21a981b3

Observation c3f7131e-5e38-46df-9f2a-7ea4c19805f0 · outbound

This paper cites Proceedings of the 2025.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.498690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.498690Z digest=sha256:60270df3121bfc95874c796cb51fe455858164fa8e5bcce103e03c5f3f726465

Observation 2808b874-9e30-4988-997d-b30941a981eb · outbound

This paper cites Rethinking.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Rethinking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.593629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.593629Z digest=sha256:808aaff89a0d8f8a68483791b39310900e4449fb334dcc77d8b5b6b722c5f230

Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · outbound

This paper cites Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.707630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.707630Z digest=sha256:f3a571e63c54906b6b9785118708565f4b415af30ccd1a7ac4cb833a957f4a6c

Observation 2ff3cc6b-80e4-4297-a894-ab5196b95f4d · outbound

This paper cites Outsider.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Outsider

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.776696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.776696Z digest=sha256:ecf085e00b71c5d76369486888398277334f7446919f901d2afc4281623ca1a2

Observation 7a18bd0c-da49-4965-981a-f5c1d621b604 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sociotechnical Safety Evaluation of Generative AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.836468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.836468Z digest=sha256:6a81c12ce7d423ca2296d423c1aebf5434a8a5dc7d6b63e7b608d589850bccc3

Observation 3825143f-fd23-4b60-8b37-122987b4325c · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.911671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.911671Z digest=sha256:dfbc7d0f31296ec1af4b46cd656bca2b57102fa529f74f1104165d9e7a07522d

Observation 095f31e1-3024-4640-83ae-201405d6b4e2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.439156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.996295Z digest=sha256:c53e8feae41fb7b5a7e1cfc29ec0d107b3125fb28595ac5d68dff042d04c822d

Observation 19b54a70-5013-4aab-b741-c30f055a0581 · outbound

This paper cites Proceedings of the 29th international conference on computational linguistics , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 29th international conference on computational linguistics , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.212267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.106520Z digest=sha256:6e76afb8f805ef0a9890407ebbab71fa14385397ce8a1e5e4e03b7e9dfe98981

Observation b7ea2215-28e0-437e-82e8-bb84696ebfb7 · outbound

This paper cites International Conference on Learning Representations , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) International Conference on Learning Representations , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.193313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.193313Z digest=sha256:2ba3ed87b9a5f1078044171b4f81452019716b956fbc30a4e6a76faaa85e8717

Observation 4e9c0f34-3323-44d4-89bd-3ce922ffa349 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.278667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.278667Z digest=sha256:31a09d07c151d21e8be2f62bd21dd84806d47b75da8bfb1b1486a774c46f2120

Observation d7b4a3c1-3323-44c2-8021-4b673020cabd · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.345910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.345910Z digest=sha256:0a6547e8d9bbbf42ae87c0fe9bf6cbaa51eeb21b8bdacf17e3da65fadd8e9ba9

Observation 91789132-3150-4b20-8918-30ab698702b7 · outbound

This paper cites and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E

Reference 59

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.704440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.435665Z digest=sha256:30cdfcb1ccaa7b86949f848779997c4f6c5867d2af65e4dc26999e31173f4bc5

Observation 198004bf-5cd5-4244-8dbd-4afd50b28c74 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 60

Resolution
verified exact
doi, observed 2026-08-07T12:44:37.153884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.563039Z digest=sha256:2712e2bb6138af5fdd3c9cdef75a2d0fa21f2bd72db1d914481a83a61c7d6827

Observation 26b99a95-b552-483f-91de-7b8bd9d9cc69 · outbound

This paper cites Sacred or.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sacred or

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:03.955485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.672670Z digest=sha256:29c38ab0a18e236ce95b042a3f6ac9160dfeccafed4d3334acfb153b33571aad

Observation 5a6ad17b-8612-40c6-a6b1-a94cab07c62f · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.752949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.752949Z digest=sha256:0ed608867193c5614a68e3c3e38794e9b5268492f8e03d53fad49cade76ba7ee

Observation b1e89b97-aa78-4cdf-87ea-310d3df49387 · outbound

This paper cites and Metaxa, Dana.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Metaxa, Dana

Reference 63

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:45:03.305534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.796636Z digest=sha256:6a5b34305d51be3d09ff5e17d355acb0f011acd2a3689a3f0285a386a588d525

Observation af29f3f7-4e3a-4bc3-9c8a-b23f72efdb87 · outbound

This paper cites Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.850378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.850378Z digest=sha256:303e03e7d41bd2e58ddd27df317496bbcb12858351ae717985d1dfe6beca19b7

Observation 4b748991-47bb-4794-885b-a35805ba49ce · outbound

This paper cites In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.904062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.904062Z digest=sha256:8908203c639aafdcc8965e8326f6a8088ff460790d385dbc3ef55f35e24deb69

Pith citing papers

No inbound Pith citation observations are available.