Pith. sign in

Paper Citation Record · LEDGER

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2608.06202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06202 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved44
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26221a13-0976-4b7c-8ff5-babd1a283942 · outbound

This paper cites 2026 , note=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2026 , note=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.564761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.564761Z digest=sha256:4e3e9d9e7fd485b4804339f89f498ef1d91cc581bb0aa44097b6382d6a199649

Observation 622a2c7a-90e9-427b-bd74-9178d043a17c · outbound

This paper cites Findings of the.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.990676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.623981Z digest=sha256:9c7c268f7078cf09c091eed160c8f6e84bc87d8b994c64ff863ec6363b9be7db

Observation 462ca983-a498-4f75-b75c-16294ff53ff8 · outbound

This paper cites Proceedings of the 62nd.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.837480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.700906Z digest=sha256:ed52cc6f28bb8eaa78522b6a831c0b0c2d61e733db065597df7416a4b48d7019

Observation dee58e1d-f0a9-4e4e-8c85-9406f8f71f35 · outbound

This paper cites Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.677459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:31.813047Z digest=sha256:7e0942104eb37a5bcde6f4c77f9b45b49b54350e76fb8b4f9ef8ecf26373a489

Observation 1d6aa5bc-2a07-47f9-8374-28c6e9e277b0 · outbound

This paper cites JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:31.930407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:31.930407Z digest=sha256:cb32b8b1feb6275554e4d80a3fac83afb684ba80d580d5660bd886ad9141b8c7

Observation 71b7280e-9268-4942-ae0b-d9bd554abc3b · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.547709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.026529Z digest=sha256:def07599c5888eb65772a4a9c10d626c99a57c453cbb3fbbd25c069ea89dce70

Observation 60b402ce-04b2-4459-b5f6-e38590f0a660 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.091734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.091734Z digest=sha256:82249b996230a2dbcb630a878f04f8dda6732668b39f70428aacb39591282a2f

Observation f4f4de7f-479c-4700-a2e9-6569fef31de5 · outbound

This paper cites European Semantic Web Conference , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) European Semantic Web Conference , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:07.332701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.198014Z digest=sha256:18ec2f0f24b34ca2246e1a664a8c8dc070b133758a34957aadb33a9466d6c045

Observation 8c2e8861-ca7c-4252-972b-efc4ed7283d9 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions of the Association for Computational Linguistics , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.295139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.295139Z digest=sha256:ea2a938cb10d14be7ccc65ba2a9402610967e896965b6de3607293c4da8c9075

Observation 94a50989-6931-4d0e-a9fa-d8b75f58de0a · outbound

This paper cites Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.383467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.383467Z digest=sha256:569afb58ca16ec696448168c0205d4531609b086726eee3b5dbf8ccc1e10a8ef

Observation 68e4a836-098f-4d8f-b437-4c400f68a726 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.454833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.454833Z digest=sha256:7e7157c91ead34bcfe359f21374d1d912f3f03393a90e3a8c8ef87ce17ddedbb

Observation 70309c71-8603-499e-99da-07f30a0a32c8 · outbound

This paper cites Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.985875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.497238Z digest=sha256:c435ef87720eb677ca7a087231df1c795602b59056615bb966094cc1cbc15aa2

Observation cfb01546-9fc0-4fbd-97f4-7857da88a24e · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2020 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the association for computational linguistics: EMNLP 2020 , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.549221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.549221Z digest=sha256:4597c128a3c2fe491a082cf0489a6759047ddaf03005c9ff57c36c14607e4349

Observation a7228e6f-6477-4bf1-98c1-836015154b6f · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.595044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.617451Z digest=sha256:2621911e8c10fc224e4cec7f8ff4fdaad3b68ad73e506a07d08668b265d3ea17

Observation 2d72a67a-258c-4fa6-8534-0578564f6b3c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.360053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.703087Z digest=sha256:fa30de24401580d4ab323b1e39ccf582aad80c20c751413494d554d0c8e88bdd

Observation e8d60a6d-d40b-46f5-8eb6-d5ffcf91416b · outbound

This paper cites C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.795670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.795670Z digest=sha256:93af6dc87261cc1f4bd147bf8e717a07fe0f898e25655b9f0766f367472134e8

Observation b47d6d3e-4761-4dc4-a5f8-a6b66cf1012b · outbound

This paper cites Transactions on Machine Learning Research , doi =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on Machine Learning Research , doi =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:06.241243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:32.882722Z digest=sha256:9d8f14882245e5474cc07f1082acafa667f08c70f189ca933919d2175df9e1cc

Observation 8ef6a630-248e-4bea-aecf-078d979fbb64 · outbound

This paper cites ACM transactions on intelligent systems and technology , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ACM transactions on intelligent systems and technology , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:32.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:32.978703Z digest=sha256:6bdbc9b64e1a2ed9924475d68b5a8593e5f3d29fa493b434d25de9ad7330d4e6

Observation 7776f9c7-d798-446f-a516-ba187f2bbd87 · outbound

This paper cites Transactions on machine learning research , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on machine learning research , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.102903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.102903Z digest=sha256:366702787b3eaef3dbbaae66b4de46cd9548654f3fd6e901d0ed537af48744b4

Observation 83be2d93-8c2c-4244-b90c-003d12d088db · outbound

This paper cites Measuring Massive Multitask Language Understanding.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.168261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.168261Z digest=sha256:f959d17cec8683316f9b018fc9fbbf30bb827db16f62e5d9d60dc9894d71f776

Observation fe0054e4-ac38-4184-9bbe-44a120d787cb · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: NAACL 2024 , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.260360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.260360Z digest=sha256:9c9936b16d3f493e30a5d458b28c18f20a709556a1b0b0ae94dd81e0696c2dd8

Observation 086b3cad-130d-45e1-8776-16f61cd85221 · outbound

This paper cites P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.369494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.369494Z digest=sha256:554e6340d3adee123c067c93af2d58578a0381b44287283ab29192265f50af03

Observation f92216a6-268c-4b31-bd92-8fc6e0a4b4bb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.461520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.461520Z digest=sha256:85867e220ec0a3c135cb0e792db3cf5e17f4d60406c45f51e3f6c171dad7702f

Observation 1e24bf35-9c64-4a7b-aa78-09c43b5ead4d · outbound

This paper cites Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.861608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.550095Z digest=sha256:1dae186001c74992cc8e0c1752635bae6601d7bc8b0accf8654bbf4e09149fa7

Observation fe0a5809-6866-4f94-a79c-1da3448024a0 · outbound

This paper cites Understanding User Experience in Large Language Model Interactions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Understanding User Experience in Large Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.653851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.653851Z digest=sha256:d62c28ca99260e4f9e2bb6e725c3f0f261a0b4c16b70247f0715d7697b3b78ed

Observation d46bb22a-b864-4e17-9d3f-89312334bcac · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.736893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.736893Z digest=sha256:e63a4236a1128d74dabbe6bfa84fcfc8fb1c25b2f7feade5d54a9167f1028dbe

Observation a8b092d3-0c98-4656-96af-edda0a1640e5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.649164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.804536Z digest=sha256:e5822d45e75e0740a87143371edd77e512671992b03fa0b7091038030f35de16

Observation 0c553b29-49b6-4375-aa24-bdf47c7ce632 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.856791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.856791Z digest=sha256:3dd482f73b13b27924fecb8adaee4149301e8c053f3843fc9c07330c58dc7a69

Observation da414d8c-eecd-4044-b747-d975df546e44 · outbound

This paper cites arXiv preprint arXiv:2509.19364 , year=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) arXiv preprint arXiv:2509.19364 , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:33.930656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:33.930656Z digest=sha256:6b8dd4328d231c35aa0fdfe2cf90cc979e07b96bf7c3ebbbc79dee06dd2d1650

Observation 2e3f3f2b-b54d-4663-a092-ccc87aea3b80 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:05.404673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:33.991547Z digest=sha256:156c2024734139e21c60d455744475f86a28924b015b5c649121c96aba2ada7a

Observation 9a9935d0-f7b6-46d0-acd9-2b2a9276d937 · outbound

This paper cites 2023 , isbn =.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2023 , isbn =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.071960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.071960Z digest=sha256:e7ceabe36c6cae23d7041b98a762cda355e30882337a2918b898b9b6867baec4

Observation 6161cf9b-f2c6-4a51-bb84-19e368a3ece4 · outbound

This paper cites On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.189426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.189426Z digest=sha256:28d5056ce33170a3b88ad8656709e5e55a8453084b9c14162ae7432a9b0b767c

Observation 9dccb30e-fa97-4f3c-9bda-fba0e3b7e762 · outbound

This paper cites Ask Again, Then Fail: Large Language Models' Vacillations in Judgment.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ask Again, Then Fail: Large Language Models' Vacillations in Judgment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.260166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.260166Z digest=sha256:6b2799ef626e7bbcb122166a7fdbc2a2068f0a7722e35d8cb96254d0670935ff

Observation c0b012f5-87a5-48b6-bcab-0b05fd694512 · outbound

This paper cites S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.316712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.316712Z digest=sha256:c0eec8d85314f411d6b8f98c374f069d5bfe7b46542b5e7b5fa1efe23a909024

Observation b346a3ab-f126-4e6a-bc5c-5e514e1e25f6 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.409924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.409924Z digest=sha256:d00b8121fc10b4be512b76a15777b4f659cc5e30c740f335902de4140040253c

Observation 15f074f6-1ad5-452b-af55-9103cc4219d2 · outbound

This paper cites Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.503503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.503503Z digest=sha256:edbeb192d3919387328d7a2172e8e024409bddd0bd9da6552321e76d842461f3

Observation 1313c07f-0d7c-4d4e-9d24-d81a7fb31217 · outbound

This paper cites LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.565332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.565332Z digest=sha256:ee9cf4dd9b7a75b1d2220462cfa2c4fc32802ce5daf025b1efe4aa42ddf922cd

Observation a458f07a-0ffa-4f31-865f-231be2c48040 · outbound

This paper cites Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:05.151838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.623847Z digest=sha256:0eb3a0afe67975b3eec73a328740747c7ddd82a365b7f2f92b620fa220fb7c05

Observation 96133c27-3520-4a5a-9715-6591749ebe40 · outbound

This paper cites Practices for.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Practices for

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.934779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.700820Z digest=sha256:3c2761d487dc635d2fc4a94905517896443d3ca9f5ffebfa3227e314166da093

Observation 856bbd95-784c-4347-abaf-daea68706d34 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 40

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:45:04.655486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:34.816885Z digest=sha256:1d8d4b2a4b7c9f99f4435e05293a96034fd2885e4b9f4cd3bddb4208a47a552f

Observation f43044d6-d349-4364-bbd8-fb8fb3bedbbb · outbound

This paper cites Aligning AI With Shared Human Values.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Aligning AI With Shared Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:34.935769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:34.935769Z digest=sha256:9792306cbfb4a5d6b0c2166a5d5a11716eae8a0b79fbd1270a41fad5291384c5

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:38c56ea564ee88ed7887c36331a591fd2f0b59b9dcb8c05788f9043eb31267ef

Observation daa23e92-20a0-4295-a933-fcdb476df63d · outbound

This paper cites Science , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Science , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.125805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.125805Z digest=sha256:a04cb95cb4bca59fd1afdd97753a2cf02cc0ef58b7e796f37a5192516df47ee4

Observation 0ca37549-c4c5-45e5-970c-15f40d1808ec · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.931158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.232606Z digest=sha256:c48745a05a1391947fa12beb35031c4c746b3a84f6f8fc9374fa83816d6acff4

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · outbound

This paper cites The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:66fb4ff7c518ecc2faadf203cfbe1b4718109461309b3274e401f490f6a16db8

Observation 05b8b964-fc11-4886-a522-ddbc7f20ed3f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.393890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.393890Z digest=sha256:fd1882a17d24961bc0cf79d71eddac68b914297ed265b3a8bfaa319d276c397a

Observation 09ebe1db-820d-48ac-9d25-7ef4f3d17a2b · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AI and the Everything in the Whole Wide World Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.451657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.451657Z digest=sha256:d3fe4be53b50dd80450003a495f8cebd141fa0ae615f8fcbe1e08147b402dec6

Observation c3f7131e-5e38-46df-9f2a-7ea4c19805f0 · outbound

This paper cites Proceedings of the 2025.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.498690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.498690Z digest=sha256:09b97730e341bec94752503de16e178aa1606a187d8bad831c46d78f48a650c4

Observation 2808b874-9e30-4988-997d-b30941a981eb · outbound

This paper cites Rethinking.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Rethinking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.593629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.593629Z digest=sha256:8ddd30f0811b6f086c3570adb0dd40bb7b40e1cfacb1e3b286d3e24278e4985f

Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · outbound

This paper cites Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.707630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.707630Z digest=sha256:d4e2325d6679ae9e53d7bd427bc74dd6f2367567141c87205a6d1b4c09f21614

Observation 2ff3cc6b-80e4-4297-a894-ab5196b95f4d · outbound

This paper cites Outsider.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Outsider

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.776696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.776696Z digest=sha256:d43f3fc30007b88ddbc17e9c176494b10a4d2f704c15422bcc4d593298a565ea

Observation 7a18bd0c-da49-4965-981a-f5c1d621b604 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sociotechnical Safety Evaluation of Generative AI Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.836468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.836468Z digest=sha256:9cd3e8f0cf0220d5637f6136dc6dcf7b246416bf95b68f781a9cc14b8ad74e9b

Observation 3825143f-fd23-4b60-8b37-122987b4325c · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.911671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.911671Z digest=sha256:69403938b7fc0d3ab721bc3a0c7c5f8b0425a056e958a952d97f11f4ddf7fcd2

Observation 095f31e1-3024-4640-83ae-201405d6b4e2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.439156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:35.996295Z digest=sha256:e42a583e9088df41dfbac1b1e3b4475df4e34cd4b08a257634d5c03667650514

Observation 19b54a70-5013-4aab-b741-c30f055a0581 · outbound

This paper cites Proceedings of the 29th international conference on computational linguistics , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 29th international conference on computational linguistics , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:04.212267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.106520Z digest=sha256:4070e10731d755ce73732e77c66b087f0bd591d29aa28ca20bfd773a6ad945f9

Observation b7ea2215-28e0-437e-82e8-bb84696ebfb7 · outbound

This paper cites International Conference on Learning Representations , volume=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) International Conference on Learning Representations , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.193313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.193313Z digest=sha256:7ac8a7ba13b7cd1e2560cbff9d3c977f958824de0f629380ba0438259115553c

Observation 4e9c0f34-3323-44d4-89bd-3ce922ffa349 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.278667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.278667Z digest=sha256:63a14c08059c5533464bb461a82b184118f6608d715619b3c1149fa75a120f95

Observation d7b4a3c1-3323-44c2-8021-4b673020cabd · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.345910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.345910Z digest=sha256:f851a4d668eb28ec904bf73f8c90a1c2c130fa352fbc5ee9451f1d4272dd1b3a

Observation 91789132-3150-4b20-8918-30ab698702b7 · outbound

This paper cites and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E

Reference 59

Resolution
verified exact
doi, observed 2026-08-07T12:45:02.704440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.435665Z digest=sha256:dcf4aecb8182de83e365d4494de8624ff83ca019c49417c859e4bc99e0b80041

Observation 198004bf-5cd5-4244-8dbd-4afd50b28c74 · outbound

This paper cites an unresolved cited work.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work

Reference 60

Resolution
verified exact
doi, observed 2026-08-07T12:44:37.153884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.563039Z digest=sha256:f5a31ed3ccc43931db3de7ad53953d8f3942bc1cf384526ac4e61921c7da54f3

Observation 26b99a95-b552-483f-91de-7b8bd9d9cc69 · outbound

This paper cites Sacred or.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sacred or

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:03.955485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.672670Z digest=sha256:d051992a8ffaea3bd6f2a0d96a11418c5983eb0f58b8823c3245b3b477ad4e84

Observation 5a6ad17b-8612-40c6-a6b1-a94cab07c62f · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.752949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.752949Z digest=sha256:62398ed89cd43ee221ddae52cccaf9561b097ccccc70ffa11668766f9b88f658

Observation b1e89b97-aa78-4cdf-87ea-310d3df49387 · outbound

This paper cites and Metaxa, Dana.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Metaxa, Dana

Reference 63

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:45:03.305534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:44:36.796636Z digest=sha256:649b82406a68f79241158318a3e1fc9c5e2d271195494b4a791c08d334e16105

Observation af29f3f7-4e3a-4bc3-9c8a-b23f72efdb87 · outbound

This paper cites Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.850378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.850378Z digest=sha256:f1843278391c2e530b45a99019f1e736ddc474c8d9dc3e8b9b4e953ce64a9044

Observation 4b748991-47bb-4794-885b-a35805ba49ce · outbound

This paper cites In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:36.904062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:36.904062Z digest=sha256:7459d944d1fc69c8141ffeae20ab6f251c2643093bbbf1286981ff533f2a5af5

Pith citing papers

No inbound Pith citation observations are available.