Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2608.03340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03340 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:45:25.334279Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa26d19c-07a0-4ccb-8212-206ecd3ff29c · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:33.624847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.449417Z digest=sha256:dae1aef122c149f7480868833ce9c2c5ebe65487d467b2f1dd39f9045219e2a2

Observation afb8098b-667d-4ded-baee-5e717cb87a17 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:33.337602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.550110Z digest=sha256:0e34d094d32a381193895a85ed9e547f0e1ab2e7488701d4e5316df991de63d5

Observation 2a135fc8-f5ec-4625-b6f3-7c38ef41eae2 · outbound

This paper cites 2026 , month = apr, url =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , month = apr, url =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:33.203841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.658346Z digest=sha256:63bdc0fb8dd08ccc1b96168d84a321d681b62582e4504cf39387396afe92ebb3

Observation 2e939414-952f-4a61-a5f7-ae748decb158 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:20.763940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:20.763940Z digest=sha256:83c62bd623c79bd1e271e3735d23b274743277ac6ac1713bccee014925776840

Observation de45cbaf-98e6-4ad1-b2ca-d258a1fb85c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Advances in Neural Information Processing Systems , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:20.845031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:20.845031Z digest=sha256:da833641ad3dde9294bbbdb59ae2fd02ebdba0ec50906f1724a15c6b41d4c546

Observation 5282a702-3aa6-4dd0-b274-c032aa99ae9c · outbound

This paper cites 2026 , month = apr, url =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , month = apr, url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:33.008976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.922632Z digest=sha256:cf81a72116ea116fd373baf95ca8a99761e5cdf687280a5fc708058dd0847f9a

Observation 43ec1675-4301-409e-8cb1-e5d1bfc2bc60 · outbound

This paper cites Or Not? , author=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Or Not? , author=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.790256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:20.995809Z digest=sha256:ff7efd2051dcae5d60e6b739b48c1bfe0851729c6aaeb932c24248c3666494ca

Observation 01fef925-c661-4b57-90f2-304a50d79323 · outbound

This paper cites How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG

Reference 8

Resolution
verified exact
doi, observed 2026-08-05T20:45:25.518534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.108941Z digest=sha256:509b44dce178bced78f06969b383b8cbe94bbcd604983e870c21045f0770fc4b

Observation 127e310a-6821-4188-8f0c-d2facd005688 · outbound

This paper cites Benchmark\^.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Benchmark\^

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.604562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.186853Z digest=sha256:989db8a4c66578a44d2f4c1f723cb7f9da9f4dd996e39195421ca1e070b34f1f

Observation 5397c555-4833-4e7e-ac7a-dcebfe4fb56a · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.424635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.263902Z digest=sha256:b551fc816788a7208eb4a4eadca91542579cac804f39375721b1ed8771d54206

Observation db2dd29a-4624-4216-882a-828003b9b3a0 · outbound

This paper cites arXiv preprint arXiv:2602.10657 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2602.10657 , year=

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-05T20:45:26.143824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.316558Z digest=sha256:27f8a3c1adab846bc50ef98417a95f1aeff2d1abf4ccdd18d2caa7c2b0330194

Observation 0160079d-44a8-449a-b2be-c9a8af6fa2d3 · outbound

This paper cites BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T20:45:25.924443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.383348Z digest=sha256:986da0cdc06c0b1e62bda799b9ba944a4f2ffb03d8b5c8be6948ceabf8ff2db9

Observation 56379f78-0bb1-4456-978f-6501e01a31ab · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.241564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.444789Z digest=sha256:a8a6f1ba4ba9a128b5f90e61f19e95fb566c51757e3c190cce800fd23a8a098f

Observation ca020e86-6664-42d7-9928-97f3e0eab768 · outbound

This paper cites My answer is C.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks My answer is C

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:32.058435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.505260Z digest=sha256:2ecd4d2ecc79f936882ba90454d5740ffc028b68253816c765d766405067e375

Observation 733952c6-b42d-4a14-80ea-6fc49a159c37 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.905048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.603442Z digest=sha256:890fd5826ca83c3822aae9cb14b2cf5891e64b84215905ef0e2b18ca00a040e0

Observation 9e96d133-b9c0-4a7d-b1d1-6637c9cbba91 · outbound

This paper cites PNAS nexus , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks PNAS nexus , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.763326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.676758Z digest=sha256:8e2ca3d02d6e351172743b1f89ea6fcecb000ea8f53dc3247de820b3274a2380

Observation 88e24ca8-4eed-4a48-a8ae-7ffbf82131b4 · outbound

This paper cites Scientific Reports , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Scientific Reports , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.537238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.724056Z digest=sha256:dfb0e3055179a32c2d7d29d69bc68ebb36f0b0b4febd494b185bf488e60cf60b

Observation 229af391-81c8-4443-b82d-06b4295f31fc · outbound

This paper cites Open-World Evaluations for Measuring Frontier AI Capabilities.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Open-World Evaluations for Measuring Frontier AI Capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:21.806249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:21.806249Z digest=sha256:0d08a521570aed7d4cd6454cc4bf724acb535c91080038fdba1de3cc02eaad95

Observation a9ae06f8-1a30-4c7d-a166-8bd9fa0cd825 · outbound

This paper cites arXiv preprint arXiv:2502.14359 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2502.14359 , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:21.881351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:21.881351Z digest=sha256:110a23e442f94cf7c2cad55ca0f3b869957b53ac7182943e79a163e97e3dba7e

Observation ab190fd9-26ea-4b1e-9164-de9af70698f2 · outbound

This paper cites Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:31.406129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:21.988606Z digest=sha256:b5e6b069278e3efba927a52459560b1d09b9624d3d24bc3fa72783f32f3d6d93

Observation 06981e3a-8a59-40fd-8f30-9fd20198ea41 · outbound

This paper cites Communications of the ACM , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Communications of the ACM , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.050562Z digest=sha256:825604f0ca101bed1f205c5b79a81b0e182d0fae8fefeaa2abcdd9dad0f610cb

Observation db2dff29-6f6d-4a1d-87b7-d3f9368acb1a · outbound

This paper cites , author=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks , author=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.129241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.129241Z digest=sha256:a130bbeb6222e79d3345ce2fac768ff5926e9a1dd88d0fb512c3f76fc931f764

Observation 52719735-10b8-45f1-868b-0fcd1b6670fe · outbound

This paper cites Communications of the ACM , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Communications of the ACM , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.203610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.203610Z digest=sha256:5d90449e538dd00aa75095f22abcdc0b80f345f3b56322253629a4a09c960f50

Observation 4ec8f9f7-b2bc-4807-8087-fa7e1402fe32 · outbound

This paper cites Proceedings of the 57th annual meeting of the association for computational linguistics , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.268373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.268373Z digest=sha256:578a42c101460cdab3c977abf56212bbf784c8e2d6165ed840c27cfee6c6d856

Observation 7eac5cff-975f-4b49-ba7c-e35f7aca30af · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.341993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.341993Z digest=sha256:c5a41d032ed4fef7eb72eb71e3fa2ae6de05968acb9aa58d053e862a19965146

Observation 2294bfab-459e-4110-ac5c-71ea8238e437 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.417617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.417617Z digest=sha256:3b7eff3d2621798b91b9c79e9ff7132602d84faadb8c355f55de8d97a756c26c

Observation 0421babf-0f9c-4a70-9d31-da33b8c21e3d · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:31.127482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.512162Z digest=sha256:fb0add7a2b10258b5f43eeb86a879249b1a49141961168e2321813791e376a83

Observation f75f59ed-422a-4043-ad8a-b7b3b8e06dee · outbound

This paper cites Proceedings of the National Academy of Sciences , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the National Academy of Sciences , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.595110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.595110Z digest=sha256:006802aa4dcb7a20e0a5dc0c50ba6ebac3f24e4a59a696cefd1edeffe92915f7

Observation 48271d95-d7f6-4b66-b52c-923a698f89b8 · outbound

This paper cites Nature human behaviour , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Nature human behaviour , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.891981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.693317Z digest=sha256:9e6dfc7dd5293180f5f08b9fdd05c0a748c32fa86726b048b1424b8e14fce4fa

Observation 58550614-5a60-412d-9396-d7a0fe37dd61 · outbound

This paper cites Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.682265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.773396Z digest=sha256:bd5f79cf5f16ffc2a827eae788e8eab2bd95063cd9ef0e2a6cd8d213380441a7

Observation 5002aff7-5756-4581-85ed-c1810b55e424 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:30.347897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:22.867977Z digest=sha256:a71f0eebe368521d6f860f1ca51869481c6ac4b75f9611fd7fc4cfd3a8683f31

Observation c9e7186c-6e95-4e8f-8c91-66001519c0ee · outbound

This paper cites Speech acts , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Speech acts , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:22.910903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:22.910903Z digest=sha256:e6817e3c5a76deae669f3fc8cb0cff01baad98a6718b096b8636c7bed173be2b

Observation de5ab4dc-791a-4f89-8024-4038d17feb4c · outbound

This paper cites Linguistics and philosophy , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Linguistics and philosophy , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:30.011257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.019090Z digest=sha256:5c9d34a598a650102879b885509494e4c699a4930ad810343b6d44af42675a54

Observation 1bf4a69f-2297-4daf-b649-28a88359af36 · outbound

This paper cites The Stanford Encyclopedia of Philosophy , editor=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks The Stanford Encyclopedia of Philosophy , editor=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:29.675791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.114665Z digest=sha256:259090597795159701fe2c5d7ecf1e1f35ae78feec598ffa656529ab1c7bf075

Observation 240b3dd8-878e-4115-b473-bc08330fae02 · outbound

This paper cites Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:29.424806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.181949Z digest=sha256:aac1142737a61ab87ddf617a76b09dcdd2451e3fb914fa0bceb680d197efc916

Observation 91ed73d1-4662-4d14-98b0-fad4713952e3 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:29.151850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.249941Z digest=sha256:13304ed2189c40b5fc37e23dce6e6241a2ad8465387bd5129fec4110cc4c08ae

Observation f61ccf09-b85f-4386-8058-b9b243eefe7b · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.970581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.336492Z digest=sha256:ea5244ca8135c6f42965b0a9fd6fab6c29e5ec82230929c9e8d67abafa456dfc

Observation fd920a19-0991-495e-a77d-fd2c7c3dbaf1 · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:45:28.818855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.447424Z digest=sha256:0f03c4412af628dd28be0913c0db76f9504acf5234a852fb93aec5ebe44c6219

Observation 0078b89e-1746-4d7e-bf5f-3c29e74fab64 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.675032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.555239Z digest=sha256:63f7029361317622ea215cc7bbeea9a41effa040c91a4aedc15c609b7d244d0e

Observation 6873148c-fe7f-467b-a6f8-8f5722991a1b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.497198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.618974Z digest=sha256:9a73dacb51f9f8ab908119ab0445ca53efd926b674b1971d32c43b395f8bdab6

Observation 12c4f787-7c6d-4bbe-84f2-2285d4aa93d6 · outbound

This paper cites arXiv preprint arXiv:2602.17594 , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks arXiv preprint arXiv:2602.17594 , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.717026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.717026Z digest=sha256:170492a58cbfc00133060e00095a035d7ec8dd66295d0ef266654d535e2bb1c1

Observation bc39cb14-d273-4ff2-9027-d4f9655b9c73 · outbound

This paper cites International conference on machine learning , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks International conference on machine learning , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.795013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.795013Z digest=sha256:37eba433e61d066ece5f31393700299f23f62e3efa20010830055a4b71d5674c

Observation 84b1cb74-6818-459b-be54-daa536b16bac · outbound

This paper cites an unresolved cited work.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:23.882802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:23.882802Z digest=sha256:ca40fd775cf7c0979e167aae7472c144cfc19ef033087678057c41b30b852314

Observation 2c96cfc5-15b8-4a7f-b6c3-a30548af8fcb · outbound

This paper cites IEEE Access , year=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks IEEE Access , year=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.310393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:23.930324Z digest=sha256:97d1fa5a2e5cf7b55e551b0696b43f7db7b13403faf56515add8b02aad54d973

Observation 854b1981-6016-41b7-9e49-1ee231d6d212 · outbound

This paper cites Proceedings of the 29th Conference on Computational Natural Language Learning , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 29th Conference on Computational Natural Language Learning , pages=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:28.160734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.008348Z digest=sha256:8ffafbb7cb730999101bbae05de638eea51839393e0577bd612100b26cd6e4cc

Observation a38a459a-0d2f-4b6e-be60-1c74acd89b48 · outbound

This paper cites What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.096011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.096011Z digest=sha256:e82f11bb9418b474eab64c1a7d3e4add57eaaeb814f8c19dc062592f669d8de0

Observation b1a956fd-bbb9-4e7b-9c02-8668d5417fcf · outbound

This paper cites Findings of the Association for Computational Linguistics: EACL 2026 , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Findings of the Association for Computational Linguistics: EACL 2026 , pages=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.948372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.184144Z digest=sha256:c8bfe6f3f2bf47740fffa27a998b61bb91d17ce23cc66fa0041a678cc50fff24

Observation 311eb501-d071-4230-9228-4c7e3f35a4eb · outbound

This paper cites RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.247556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.247556Z digest=sha256:55801d1d992afb2eb1d7d1323d925005486a78861c393652d449c845ff6ad25a

Observation 09932863-acc0-41ed-b01c-36e43225482e · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.328597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.328597Z digest=sha256:154e267a40f9eae91b2e0283465e9ab7d4a35bef32636548a4590fd86db8dbee

Observation 1497d779-b546-44d7-b2dd-ab0c09bd559d · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:45:24.389948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:45:24.389948Z digest=sha256:3273e9be96bf9ccbded1ad5214aa24152d7465174ddf0d1099a461bfa5ad3502

Observation 4f64fa96-73a0-4e9a-a269-273326984b0c · outbound

This paper cites Innovative Journal of Applied Science , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Innovative Journal of Applied Science , pages=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.808382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.512569Z digest=sha256:21e75d2d7b5e6bc712d1a1603ece6f6233a3bc4d3374c4e1664a32ada80d62e5

Observation dc36a8b4-d02a-4906-9f36-a9568a5d8666 · outbound

This paper cites Expert Systems , volume =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Expert Systems , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.626372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.612540Z digest=sha256:bcae71c3d3d38f053fd68a3c7a2fcd93fc3e4d79615f7841922d03d52122d59e

Observation cfb00901-91b5-4d91-89b7-987a2f2881da · outbound

This paper cites 2026 , eprint =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks 2026 , eprint =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.427352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.704523Z digest=sha256:d426350c323c74157813199bacaa6fe8c67bc3a4ee5b6db6fbc31f97444c51e4

Observation 08b8dbf1-42f4-4ea3-bb39-3c57b724d0f0 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.245639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.801821Z digest=sha256:9eb4e270ce11ea6d729717d9d44123af9cf837626d139c8903e7ba92c5f555df

Observation 2839b1e6-23ef-42fb-bc6d-273e6745d3a4 · outbound

This paper cites Journal of the Royal Statistical Society: Series B (Methodological) , volume =.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Journal of the Royal Statistical Society: Series B (Methodological) , volume =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:27.045407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.926045Z digest=sha256:21f7b3551972ae58b6e313e2f9a2988c75fe8e614d1408c523ddf722206670d4

Observation d02853ed-6643-44ce-84a4-2b90765cb88d · outbound

This paper cites AI Magazine , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks AI Magazine , volume=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.893930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:24.998130Z digest=sha256:7a4941e6895d32365213184f76a88ef53efd6e8dc0969333de9d251b7506df44

Observation 4e4206c7-0ba1-4179-a766-04f0138960dc · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.734269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.082085Z digest=sha256:78f80703008d7dd032281d76363bc9e7c5d3ac389fd2bcf478dbca750c4ee61b

Observation e6832bf7-f5df-4d04-b6e8-73c470f1627f · outbound

This paper cites Proceedings of the 9th Widening NLP Workshop , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 9th Widening NLP Workshop , pages=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.588331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.171035Z digest=sha256:69d1307c80f369d6865b27d1cd51974f299217ba1968d470cb9e724d836e0355

Observation b9f74602-3860-44b0-9da1-e520bc2815a5 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Transactions of the Association for Computational Linguistics , volume=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.430611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.259248Z digest=sha256:086855d8c23f5723212af1d15ad2363ce578c6516779c386dd92f43c66c9905a

Observation 93664a2b-5556-4d70-8de5-bb9cd8492617 · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:45:26.280577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T20:45:25.334279Z digest=sha256:df411097dd4df0a62657deb228be8979ee5683a3b8035c673381d37a3f6f7f3c

Pith citing papers

No inbound Pith citation observations are available.