Pith. sign in

Paper Citation Record · LEDGER

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

As of 11 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00482 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.955891Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5b9bc80-2b88-42ea-94a2-c81776cfe024 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:09.977304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:09.977304Z digest=sha256:ed468663c418f4ee796e2db3df359b6f0f0aa482efaf88ac16db48561581d3af

Observation 2efa816f-2096-493c-9719-742628342339 · outbound

This paper cites CaLMQA: Exploring culturally specific long-form question answering across 23 languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CaLMQA: Exploring culturally specific long-form question answering across 23 languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.073701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.073701Z digest=sha256:e53c906a44d39ee0be34b183041062be3ea296e099784752dc8970ab83509935

Observation 427b5538-3fc3-4592-a80a-46dc79315ec1 · outbound

This paper cites Program Synthesis with Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.157068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.157068Z digest=sha256:2bd8377f21e52e67a71841644ab421f9dccf678da6a33b61ebbb49938a877188

Observation 02652e39-b34d-480a-8d59-119b86945385 · outbound

This paper cites Axolotl: Scalable fine-tuning framework for llms.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Axolotl: Scalable fine-tuning framework for llms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.225114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.225114Z digest=sha256:ef837e6f9dd76ab89cad709f0960b1021ef37d4c6af4ec660a31aad921c81011

Observation 70cf9db4-a008-4061-b223-5319f8215fd2 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.311875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.311875Z digest=sha256:349c27ea3b300c53e5a238a2b8ca9681697f6f92034d06fd6233df6e31e2ce52

Observation 7c7944e0-cd66-499e-acff-4d0ee4f76357 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.404642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.404642Z digest=sha256:a95529e251c5e2376222617179b4c045e2978191267d1c1938dfbab596d5a325

Observation f7e73d9e-db43-402d-8436-39b403c66e3a · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.925972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:10.504819Z digest=sha256:70998513c902e72a8e9ae78228857cfaeb0ab434914ad38b52ed5a6a78fce196

Observation 6deff245-bd09-4939-867c-5908395d103e · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.604298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.604298Z digest=sha256:cad103e7fd2365fbf492611a129364dfd09c95ffba1eccc2545a50340c0910ff

Observation 793c3c87-6d02-4ab3-aa89-120a638244cf · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.688904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.688904Z digest=sha256:0b00623c2690ec1ceb2dcbc7a1be4b87a72d89c01650ccd92081c848840aac9d

Observation 2828d71c-cc47-4304-8dec-f632bc8c7838 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.759064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.759064Z digest=sha256:ce3daa652d417991c602e4a13466588dfa22f97f67777478d4355394156d02fe

Observation 072cf44b-c212-4d71-b7d3-1275b99ec56c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.838533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.838533Z digest=sha256:bcf0ab182c2e2971328fc1d6130a0a9f0fe7e5d7de076cbb336421ab16f50a25

Observation 144866f4-6643-48b7-b0de-4ca0fff94548 · outbound

This paper cites SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.653087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:10.913333Z digest=sha256:ddf7db113ad3353a1c37c40dcf888a350faef04fb23f48cc502af14db1647371

Observation c622398f-affd-4769-8e73-e3cb93c2f173 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.040928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.040928Z digest=sha256:95985f080657eb8dcda768fb03e86922341f45433fea4f15dc683263a570808d

Observation 8c027e46-3d15-4ef3-8b6c-d084e209aa6b · outbound

This paper cites The Llama 3 Herd of Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.143287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.143287Z digest=sha256:ec60f01374d6ecd9d441df265715532cd9df78b1c6679a9d2ec16ec90871641a

Observation d306e540-5cc0-4a1f-b69b-a6d272c47262 · outbound

This paper cites NativQA: Multilingual Culturally-Aligned Natural Query for LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.380004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:11.225092Z digest=sha256:e20c17b2288147eb0b978077942e727bc7e661a10a470adea93c58c19a8e26e6

Observation a19cbe23-d5a2-425a-beb6-719f4e207aa6 · outbound

This paper cites Measuring massive multitask language understanding.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.314107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.314107Z digest=sha256:7e2990d9c9ef2e04bc156c6a6b30ccb2e265f2292b18dc7fa170fa1d90f1b89c

Observation 5c2d87d0-10fe-4bec-862b-7d175d674e06 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring mathematical problem solving with the math dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.398525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.398525Z digest=sha256:4d25efc12c7030f15f6ea30fedf08558fc1d2533f7e9356058597688e82fca16

Observation e2c7b2e9-2678-4c5d-807b-1346ce605d34 · outbound

This paper cites MedQA-SWE - a clinical question & answer dataset for Swedish.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MedQA-SWE - a clinical question & answer dataset for Swedish

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.428624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:11.541072Z digest=sha256:833dfae06da220015088daf5d91314704c20de6b42ea2265079855d16a6ef56c

Observation 91da4474-d693-434e-be15-8ab7ed1a45b5 · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.651962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.651962Z digest=sha256:e7f1c79222ee6281910f1956ee75116d051953685f9486691253359a5838d79d

Observation b0849f0c-eafa-4615-97f2-6f26cd597f58 · outbound

This paper cites MoralBench: Moral Evaluation of LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MoralBench: Moral Evaluation of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.740741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.740741Z digest=sha256:89ad547d1798cfa15d4c683e1ceadd5e0e30394e974aae41fb36c12b383190d6

Observation 384af186-ac2b-4525-9d81-b3fc01d6e567 · outbound

This paper cites KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.176490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:11.856041Z digest=sha256:6f3f1b53ea3e94c4f6763f4b114b09cf25285c6fefb15c161fdf32f6a0b0a2e8

Observation b5c7092a-64d8-4ce1-afb6-19c1b56f2281 · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dynabench: Rethinking benchmarking in NLP

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.926812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:11.960334Z digest=sha256:11b1c3877d9662f9d2d3f623f651bc7b0a95c2d6c71be4c106e072ebae14cd3b

Observation 7c1e2b82-d1b1-49cf-8110-794e6fa5f3f2 · outbound

This paper cites CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.623537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:12.098448Z digest=sha256:ad8cf459fce62ca9d6e1adcb0e258935934f3ebb56c950f0eb0ba5a485c8a988

Observation 1db570e7-318b-486f-88f3-dddbcd66b252 · outbound

This paper cites Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.371801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:12.199734Z digest=sha256:e4ec0baadc9ac04c471577e7b34652e198b9ec1f6c70ba6950eb1408e9c5ef54

Observation 8c49a033-ff53-4aec-b58c-66fc59956c6d · outbound

This paper cites Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.304393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.304393Z digest=sha256:e5f5587b6186f17de52884ad90362ac5fee0368bc9058f1d76105cf1c3e55fea

Observation 5623cccb-be36-42ef-b4fe-07c8e3111493 · outbound

This paper cites The NarrativeQA reading comprehension challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The NarrativeQA reading comprehension challenge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.027075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:12.414934Z digest=sha256:74ba238ba5710da81d2885ff72e208281dcc581f9fad4ee906d329f982cca1ea

Observation 7b21840f-e276-48e2-99dc-7b9e718c39b8 · outbound

This paper cites KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.530994Z digest=sha256:0117ed2c92aea246f5b3c93fc72cc64903e3df5fb2a12784e6a5f0c62cbd363c

Observation 51d8d5d0-d31e-4d17-9557-1320c5baa8c3 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.613938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.613938Z digest=sha256:4c7fb63cf84fed4aa69f4aac43b13107ad7f2273ccf3109fdcab2df9bdc0a767

Observation 6c3d54d9-ab05-4870-a72c-8534c0f1cf09 · outbound

This paper cites KoSBI: A dataset for mitigating social bias risks towards safer large language model applications.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoSBI: A dataset for mitigating social bias risks towards safer large language model applications

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.686665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:12.725096Z digest=sha256:53fdca41ace6a7f2742b0fce2cae755e6d8f20827c8fe9a73102681cc42fafd9

Observation aa081316-6729-4805-b54c-00c92fb35d73 · outbound

This paper cites KorNAT: LLM alignment benchmark for Korean social values and common knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorNAT: LLM alignment benchmark for Korean social values and common knowledge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.394750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:12.848365Z digest=sha256:d31df4b8c3f94c1978207b1bdb0d5c2405fe7661fb5ebbf483f4deeae8ec6b29

Observation 6b5fb4e9-c1fb-4fef-982b-9a5323918bb6 · outbound

This paper cites LegalAgentBench: Evaluating LLM Agents in Legal Domain.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.933539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.933539Z digest=sha256:61d7fd211b5cc38451b4f10f1ce6540f0bf7c5830e37cecfea17e5af6aefbda2

Observation 67f7a762-27b5-4fb1-b424-061269a1fc1d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:26.087765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.031385Z digest=sha256:5c7ed900306181dcfbba937be55b6dec8062b82ccddab6a522820b72441b1b59

Observation 0d952236-2223-4d8c-ad29-5afd2ade9bda · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation TruthfulQA: Measuring how models mimic human falsehoods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.117735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.117735Z digest=sha256:cf3e58cc9855b4d3e49989a635d9fbc060b222f9b018290b9f48c2cda27f4a98

Observation 9543c30d-1994-4996-b889-20b8a5679c43 · outbound

This paper cites Benchmark data repositories for better benchmarking.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark data repositories for better benchmarking

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.844219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.193135Z digest=sha256:904d12902a2b741bb298415b773f0212b245910b1c680b66ed00f239eae3fcf3

Observation 8efbfdfb-c22c-4d53-8960-c00821637f53 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:25.663250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.299417Z digest=sha256:94311b620e0d79c26769d730afeff00161a02df9e8d858009c2bd5713841e9c4

Observation 3e283af4-44c8-41ed-937e-5f530e085286 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.416530Z digest=sha256:c8f84c79462571e7e2783e4cc76660b8e9d751bc488c47df40e7fc4685ce3439

Observation ab79f9d2-19ed-49ae-92f0-8831166b35e3 · outbound

This paper cites FActScore: Fine-grained atomic evaluation of factual precision in long form text generation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FActScore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.078670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.497724Z digest=sha256:45523461cc5e57d7e6f2f74ddfdb0c04ee474d52e3952953bfe00bb6558e80f4

Observation e86d7949-42bc-49ed-886f-eb7ae1559373 · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.563546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.563546Z digest=sha256:50ff7057335c02ff90d60f58506f47317a8d04dee862c5b99d6a048b4bb4ee30

Observation 970e71b2-80f9-4235-8a20-6bc47dc0dba2 · outbound

This paper cites BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.796108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.674542Z digest=sha256:6dd6f0c490bc3b4c25c345176835b8e34f608ca90f0f3e82613176d1c24c2487

Observation b746cf2b-6230-403e-8f8f-4ca4330ee10b · outbound

This paper cites Extracting cultural commonsense knowledge at scale.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Extracting cultural commonsense knowledge at scale

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.597620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.809086Z digest=sha256:397e916bc1907883186ba450b40473449a19995e0c0f47b03fb1c8260fe0efbf

Observation dc5d50d0-ff3c-4ea3-9cf0-2eff798a6137 · outbound

This paper cites MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.233222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:13.880832Z digest=sha256:7edb741d2a59a105ebddd346574acc70adec56c361e760a5fdd45b20e64d6705

Observation a167dabf-4945-4a79-84f1-5217c4f6ea1f · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.986625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.986625Z digest=sha256:c9ac0dfa32c30fb247d28a2fb9660f2241b93d40ba09b22492217b437090e22d

Observation c92b6544-3460-43fc-b53b-4942d32311cf · outbound

This paper cites BBQ: A hand-built bias benchmark for question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BBQ: A hand-built bias benchmark for question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.031957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.078274Z digest=sha256:36aaa7a899537984affb0f338904968b7fb926fcc7b556cb1b1f84ced476d62a

Observation b115779b-e43d-467f-9b5d-3848fd76cf1a · outbound

This paper cites Survey of Cultural Awareness in Language Models: Text and Beyond.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Survey of Cultural Awareness in Language Models: Text and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.144191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.144191Z digest=sha256:5e85facfb91e6b56958dcd85e34084e1d6b15a6aa68f5b7d1829026c65da1e14

Observation 434e5316-32d9-4fce-937c-0ebbae8b572c · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Zero: Memory optimiza- tions toward training trillion parameter models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.722060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.231714Z digest=sha256:33085429b1d3b10020ad70168229602bd0eb5b2b72e614c20db66bd24c89dd97

Observation 431e4c29-6e41-4595-958d-42b3722aa736 · outbound

This paper cites NormAd: A framework for measuring the cultural adaptability of large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NormAd: A framework for measuring the cultural adaptability of large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.401621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335556Z digest=sha256:fd08d7c7111682060628b877d12be8ee51f731d1869a182ff7942e40ab2cbfee

Observation 5fc97c41-d2a5-466c-9cd3-876a0506eda3 · outbound

This paper cites DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.085948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.581006Z digest=sha256:d4bbe1f1a6f9a5d9a3c695662d40405cda9a2f4c9303db3a5b9c24693e7c4b4f

Observation a931e440-3305-4057-84b3-49e2394adb4d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:22.795446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691459Z digest=sha256:038f533a1d603de4aacb858db46c53cffc7830681bb6d1718700de47526aa058

Observation 29c42e3f-8d2a-4402-9d66-c368ccbf24cb · outbound

This paper cites Kochenderfer.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Kochenderfer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.550842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.833051Z digest=sha256:8a757179ca6e30b1bb14447969cc39e4c0e9653ef5fd4c1848d7e972e323e22d

Observation 32256750-95a7-49dc-bd84-fbcc2c0f5f5e · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.Commun.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation WinoGrande: an adversarial winograd schema challenge at scale.Commun

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.369793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:14.924761Z digest=sha256:262469229e5b8c6e6f2a07d1a779d34ce0373d2e97261e539a4a5ccca0c6bf09

Observation cc6416ba-d929-4711-95aa-cd19b259b7d7 · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Social IQa: Commonsense reasoning about social interactions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.146589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015768Z digest=sha256:c537dbbb0d2919330c48dadf6444e934c1606c0bea835e54808f7b62793fd164

Observation 597877f1-8911-4dba-b89d-f29fbcd96adb · outbound

This paper cites Benchmarks as microscopes: A call for model metrology.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmarks as microscopes: A call for model metrology

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.940061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.107192Z digest=sha256:8d9bb3523312db47210dc5db62fc260f4f9be03092cc33fe0d136c44ebf8e210

Observation 7db6d897-08bd-47d6-854d-be04c1ce887e · outbound

This paper cites Multi-fact: Assessing factuality of multilingual llms using factscore, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-fact: Assessing factuality of multilingual llms using factscore, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.694305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.229945Z digest=sha256:b2453bff32173e4a9bbe6055d91891d74e50a960abd499187cafe4d611d9138d

Observation 1d3478be-d6f6-4cab-b9db-7783e1a7ad1d · outbound

This paper cites YourBench: Easy Custom Evaluation Sets for Everyone.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation YourBench: Easy Custom Evaluation Sets for Everyone

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.321570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.321570Z digest=sha256:4cf24a341748ff457092241da65375878c6ea3e885e39c53f1f3b6b58d8b75a0

Observation ac95e546-01e9-4a9e-b94c-0804e1c866ba · outbound

This paper cites CultureBank: An online community-driven knowledge base towards culturally aware language technologies.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CultureBank: An online community-driven knowledge base towards culturally aware language technologies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.384899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440338Z digest=sha256:b64a84e35a9b4cd2f1276015e8cd496a4713efd32e679a73c7687b6e6f4a00be

Observation 2632e0cc-e035-437f-9565-3dfdb648bb34 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.542569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.542569Z digest=sha256:8ee195b38810ebf81f797f9208a4f8db52651faf963c371d086b7339daaec488

Observation 26e37416-51ce-415c-9642-e621f807dab6 · outbound

This paper cites KRX bench: Automating financial benchmark creation via large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KRX bench: Automating financial benchmark creation via large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.170385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.639270Z digest=sha256:7e21deb47b73171663e06bf0b7ce78c62e5a45f86a2f55e05e4c7b1b041ff35f

Observation ad0d7d1a-0d36-4340-9651-9a858d447735 · outbound

This paper cites Beyond Classification: Financial Reasoning in State-of-the-Art Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Beyond Classification: Financial Reasoning in State-of-the-Art Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.739015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.739015Z digest=sha256:2c378dd2ed1e5742134de539758677c01db194724ab576cc7942dcf9a40db646

Observation 3201edf1-4f3e-4d99-a74a-b41de35e9f1c · outbound

This paper cites Multi-step reasoning in Korean and the emergent mirage.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-step reasoning in Korean and the emergent mirage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.945080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828572Z digest=sha256:87f2e557f7446693260e25089195b177562b2d51769a8fda6ff7943f406fa8a4

Observation 624cdaba-8ce7-4ba1-b8f4-cfe3cd287df1 · outbound

This paper cites KMMLU: Measuring massive multitask language understanding in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KMMLU: Measuring massive multitask language understanding in Korean

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.774825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:15.911316Z digest=sha256:a2343c1f4647bf24be171d7d2b35d76e8ce475aa81afb981ceab482e79182694

Observation 062acff7-3666-48da-83a3-d3376048856b · outbound

This paper cites HAE-RAE bench: Evaluation of Korean knowledge in language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation HAE-RAE bench: Evaluation of Korean knowledge in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.551437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.003057Z digest=sha256:2fd51150e7b3cf8846105e8f1c02fc135faced791594698480fb61a11a16c17a

Observation 51cdb724-54f3-47c5-9fc7-280e106df693 · outbound

This paper cites Challenging BIG-bench tasks and whether chain-of-thought can solve them.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Challenging BIG-bench tasks and whether chain-of-thought can solve them

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.366551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099651Z digest=sha256:ae19e2ce568caf9465c1c9474d44a952693a8e41bb769893734fe03ce8bfe27d

Observation f7a0cb8e-86be-4b03-ac03-4c9eb57a57ad · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.178310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.178310Z digest=sha256:be44bc3621d6bfbacf5e9167d6a9c49695fa29c77a818bd78ec507cb85d02409

Observation 432bf232-706d-4036-b311-207966c0a5a0 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.285247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.285247Z digest=sha256:415b16340df89c476a5252bf513b5c37e8fd5bf2a55c707720828848ae6c833c

Observation 00d892b9-90c2-4f7b-a54b-7846b57cd4c5 · outbound

This paper cites Gemma 3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Gemma 3 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.366173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.366173Z digest=sha256:23185052dc6a47bc326ed981835e3a17da63412ae04bd778b81a695b6fddaf04

Observation 4b644f29-8f98-4946-a501-c1d937aec4bd · outbound

This paper cites Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.206356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.462836Z digest=sha256:746fa99ef670c0e4e69a0a468380fbc5ee931f7a62f3581f087bf04f54a4472c

Observation 6098c32e-3b67-48eb-92a2-d20244f6f03a · outbound

This paper cites SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.993735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586923Z digest=sha256:dd4f4d3daf32c6b499efaf3849c0b6b33c36ffe5c0899f23f646ffe4930e503e

Observation 1fd0d21e-54f8-4e60-83f6-9b48a90b3a76 · outbound

This paper cites KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.759287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.759287Z digest=sha256:eaf5c86631910801c2b167f8484e3e64b7267d2be853c7bb48402df84f59fe07

Observation 49db165a-1c43-4554-97df-544abd4e6381 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.840752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857171Z digest=sha256:1b8f3a5799bbd7e69fcb1c8a2e18950381d74fdb17f210e5b089360c60a381f2

Observation af13bd18-0319-4209-a00e-594f1aa283f0 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.656148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:16.951540Z digest=sha256:ae70f43018e3d0ad9289e4afa1f72be3fa8414c4d353dbfdda82def353ff0b4e

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:5360243bcb426381ee61fa6c256696750b90fa15f84b001033125fd03868334b

Observation a3199245-1f8c-4d78-958b-bed6e0247619 · outbound

This paper cites Qwen3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen3 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.152295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.152295Z digest=sha256:6c936c926141c8fb4e922c0ed7e59300203de4726e01f7ee455e801157e928c4

Observation e2d4076c-8373-42b7-bbb3-a99321d88f7a · outbound

This paper cites Qwen2.5 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen2.5 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.279546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.279546Z digest=sha256:335bac384b58f7cb4f7d39f0cc8c35890bbcd35e6d632dd9d45ed65439c01a01

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:ee7ac22c8c208619af993ab07b3f27bfdea1f044d5119fbabe696aa91140fafc

Observation bc172362-c9b7-4ef7-9b97-e23956a44a2b · outbound

This paper cites FLASK: Fine-grained language model eval- uation based on alignment skill sets.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FLASK: Fine-grained language model eval- uation based on alignment skill sets

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.440683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447929Z digest=sha256:61f1ae061f4ad0a1df296b78a8a91a6f62d1ca80d8ef3064eec7ee57fc030185

Observation 7da61289-4ca5-42d3-86b8-83542a2efdc6 · outbound

This paper cites GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.253520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503711Z digest=sha256:f3c599a664bfb22f8ac8fdc2afc8a2e3f81ccab90aaa8952118894abfdf0af8b

Observation 725c2996-fa56-420c-b4f4-d17e36ae7af6 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.614168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.614168Z digest=sha256:fe5128acb07361b667d64d8096c476a53abd8a9c4e76215d33faf7225e7c93e6

Observation ad13bf2f-b95b-429a-96e4-7f7b91ef36e0 · outbound

This paper cites Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.085125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692121Z digest=sha256:bd62a4cdad93f8cf0e483b9b4039abe2865fa3f4658a41141d62e146d7117f8b

Observation 29a84389-3535-41e3-85c8-51ccf4318ed9 · outbound

This paper cites Task me anything.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Task me anything

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.975615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749553Z digest=sha256:1525b53f1777fdc69831677712b393728b520beb59bdea817110d5670828324e

Observation 86d268ff-e827-4c90-a2e4-463393160202 · outbound

This paper cites Users can interactively explore the overall data distribution they are interested in.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Users can interactively explore the overall data distribution they are interested in

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.847818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836203Z digest=sha256:c887bed804583ae930023b0120640ef2c9ffdb43530b9cd4dbfc0fca665fafda

Observation e0de9c67-b6cb-47e7-acf9-528c7ec99ed7 · outbound

This paper cites By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.739965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.905218Z digest=sha256:36b8cb1212275cd6753d5e03287227c5cba8679edb4deb342ccdc3b697bd86c9

Observation 7f2adf41-922e-4c81-9f4c-6dc5c22c83cd · outbound

This paper cites Is the Earth flat?.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Is the Earth flat?

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.606763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:08:17.955891Z digest=sha256:c73fee1956618ed5abb77cc256188122e7606fed4734c766159e67f268dce8d8

Observation 88c7fafc-7f12-49be-82bc-bc53c5530e79 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.678445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.678445Z digest=sha256:7104f1af34efed4e9f74d3f5c56866d68704fb81875a1b6889548bdc65a71ff9

Pith citing papers

No inbound Pith citation observations are available.