Pith. sign in

Paper Citation Record · LEDGER

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

As of 11 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00482.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00482 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.955891Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5b9bc80-2b88-42ea-94a2-c81776cfe024 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:09.977304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:09.977304Z digest=sha256:ed468663c418f4ee796e2db3df359b6f0f0aa482efaf88ac16db48561581d3af

Observation 2efa816f-2096-493c-9719-742628342339 · outbound

This paper cites CaLMQA: Exploring culturally specific long-form question answering across 23 languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CaLMQA: Exploring culturally specific long-form question answering across 23 languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.073701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.073701Z digest=sha256:e53c906a44d39ee0be34b183041062be3ea296e099784752dc8970ab83509935

Observation 427b5538-3fc3-4592-a80a-46dc79315ec1 · outbound

This paper cites Program Synthesis with Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.157068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.157068Z digest=sha256:2bd8377f21e52e67a71841644ab421f9dccf678da6a33b61ebbb49938a877188

Observation 02652e39-b34d-480a-8d59-119b86945385 · outbound

This paper cites Axolotl: Scalable fine-tuning framework for llms.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Axolotl: Scalable fine-tuning framework for llms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.225114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.225114Z digest=sha256:ef837e6f9dd76ab89cad709f0960b1021ef37d4c6af4ec660a31aad921c81011

Observation 70cf9db4-a008-4061-b223-5319f8215fd2 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation PIQA: Reasoning about physical commonsense in natural language.Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.311875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.311875Z digest=sha256:349c27ea3b300c53e5a238a2b8ca9681697f6f92034d06fd6233df6e31e2ce52

Observation 7c7944e0-cd66-499e-acff-4d0ee4f76357 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.404642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.404642Z digest=sha256:a95529e251c5e2376222617179b4c045e2978191267d1c1938dfbab596d5a325

Observation f7e73d9e-db43-402d-8436-39b403c66e3a · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.925972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:10.504819Z digest=sha256:523dea08a599f3c0d4b8f06e87f014a023d8d13f088cdf2f324a5d1206d0b07c

Observation 6deff245-bd09-4939-867c-5908395d103e · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.604298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.604298Z digest=sha256:cad103e7fd2365fbf492611a129364dfd09c95ffba1eccc2545a50340c0910ff

Observation 793c3c87-6d02-4ab3-aa89-120a638244cf · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.688904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.688904Z digest=sha256:0b00623c2690ec1ceb2dcbc7a1be4b87a72d89c01650ccd92081c848840aac9d

Observation 2828d71c-cc47-4304-8dec-f632bc8c7838 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.759064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.759064Z digest=sha256:ce3daa652d417991c602e4a13466588dfa22f97f67777478d4355394156d02fe

Observation 072cf44b-c212-4d71-b7d3-1275b99ec56c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:10.838533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:10.838533Z digest=sha256:bcf0ab182c2e2971328fc1d6130a0a9f0fe7e5d7de076cbb336421ab16f50a25

Observation 144866f4-6643-48b7-b0de-4ca0fff94548 · outbound

This paper cites SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SciEx: Benchmarking large language models on scientific exams with human expert grading and automatic grading

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.653087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:10.913333Z digest=sha256:892dda5cf977680f71f27742a2e76f3b16311622c354ef4f83e808e9603deb67

Observation c622398f-affd-4769-8e73-e3cb93c2f173 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.040928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.040928Z digest=sha256:95985f080657eb8dcda768fb03e86922341f45433fea4f15dc683263a570808d

Observation 8c027e46-3d15-4ef3-8b6c-d084e209aa6b · outbound

This paper cites The Llama 3 Herd of Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.143287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.143287Z digest=sha256:ec60f01374d6ecd9d441df265715532cd9df78b1c6679a9d2ec16ec90871641a

Observation d306e540-5cc0-4a1f-b69b-a6d272c47262 · outbound

This paper cites NativQA: Multilingual Culturally-Aligned Natural Query for LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.380004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:11.225092Z digest=sha256:d302e92b6948e1253799ca57a90f256b2b1b5719a67be1d2064f9448f2175918

Observation a19cbe23-d5a2-425a-beb6-719f4e207aa6 · outbound

This paper cites Measuring massive multitask language understanding.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.314107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.314107Z digest=sha256:7e2990d9c9ef2e04bc156c6a6b30ccb2e265f2292b18dc7fa170fa1d90f1b89c

Observation 5c2d87d0-10fe-4bec-862b-7d175d674e06 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Measuring mathematical problem solving with the math dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.398525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.398525Z digest=sha256:4d25efc12c7030f15f6ea30fedf08558fc1d2533f7e9356058597688e82fca16

Observation e2c7b2e9-2678-4c5d-807b-1346ce605d34 · outbound

This paper cites MedQA-SWE - a clinical question & answer dataset for Swedish.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MedQA-SWE - a clinical question & answer dataset for Swedish

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.428624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:11.541072Z digest=sha256:8f440526a9c9154de805906038f07b71edd4fd579d4ce8fce364c57df7c87897

Observation 91da4474-d693-434e-be15-8ab7ed1a45b5 · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.651962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.651962Z digest=sha256:e7f1c79222ee6281910f1956ee75116d051953685f9486691253359a5838d79d

Observation b0849f0c-eafa-4615-97f2-6f26cd597f58 · outbound

This paper cites MoralBench: Moral Evaluation of LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MoralBench: Moral Evaluation of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:11.740741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:11.740741Z digest=sha256:89ad547d1798cfa15d4c683e1ceadd5e0e30394e974aae41fb36c12b383190d6

Observation 384af186-ac2b-4525-9d81-b3fc01d6e567 · outbound

This paper cites KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoBBQ: Korean bias benchmark for question answering.Transactions of the Association for Computational Linguistics, 12:507–524, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:28.176490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:11.856041Z digest=sha256:b8dfb935497a090b6d42a1cdf7e30f09afbfa36972cf3a8ec590d31232489b19

Observation b5c7092a-64d8-4ce1-afb6-19c1b56f2281 · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dynabench: Rethinking benchmarking in NLP

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.926812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:11.960334Z digest=sha256:e91a927eb6622cec579d68f7991a4d08a3573f0298c5812bd205b51b3e526e0c

Observation 7c1e2b82-d1b1-49cf-8110-794e6fa5f3f2 · outbound

This paper cites CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.623537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:12.098448Z digest=sha256:c00ca8156d99faa4230d393a41f532b6b0d3028bebc23c1a8a24cf58fb331417

Observation 1db570e7-318b-486f-88f3-dddbcd66b252 · outbound

This paper cites Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Developing a pragmatic benchmark for assessing Korean legal language understanding in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.371801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:12.199734Z digest=sha256:c969f2174e04802cc555b4ed87b2241be8b3a2ea35948ecaae8bf2aeb6a9bb52

Observation 8c49a033-ff53-4aec-b58c-66fc59956c6d · outbound

This paper cites Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.304393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.304393Z digest=sha256:e5f5587b6186f17de52884ad90362ac5fee0368bc9058f1d76105cf1c3e55fea

Observation 5623cccb-be36-42ef-b4fe-07c8e3111493 · outbound

This paper cites The NarrativeQA reading comprehension challenge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation The NarrativeQA reading comprehension challenge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:27.027075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:12.414934Z digest=sha256:1a62bdc851c8b50cd878b858869adab4316a9b03b42ed9bfcf345730c686b4f9

Observation 7b21840f-e276-48e2-99dc-7b9e718c39b8 · outbound

This paper cites KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.530994Z digest=sha256:0117ed2c92aea246f5b3c93fc72cc64903e3df5fb2a12784e6a5f0c62cbd363c

Observation 51d8d5d0-d31e-4d17-9557-1320c5baa8c3 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.613938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.613938Z digest=sha256:4c7fb63cf84fed4aa69f4aac43b13107ad7f2273ccf3109fdcab2df9bdc0a767

Observation 6c3d54d9-ab05-4870-a72c-8534c0f1cf09 · outbound

This paper cites KoSBI: A dataset for mitigating social bias risks towards safer large language model applications.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KoSBI: A dataset for mitigating social bias risks towards safer large language model applications

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.686665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:12.725096Z digest=sha256:5a1e734900428d52d6dddb0dd4bad21c03a12e9b775b2b1922ec4a20a82ec238

Observation aa081316-6729-4805-b54c-00c92fb35d73 · outbound

This paper cites KorNAT: LLM alignment benchmark for Korean social values and common knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KorNAT: LLM alignment benchmark for Korean social values and common knowledge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:26.394750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:12.848365Z digest=sha256:c8b53d13bdaa52562192a6b2df702c635b229c6b4c40716e808594c1141e040a

Observation 6b5fb4e9-c1fb-4fef-982b-9a5323918bb6 · outbound

This paper cites LegalAgentBench: Evaluating LLM Agents in Legal Domain.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation LegalAgentBench: Evaluating LLM Agents in Legal Domain

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:12.933539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:12.933539Z digest=sha256:2efeea6c21e692dd716d913dc3654eb182af756e663b82a89c4c4b241e86e021

Observation 67f7a762-27b5-4fb1-b424-061269a1fc1d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:26.087765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.031385Z digest=sha256:f0f697d7a5e0a90f4c056eb2f991da4eb269af6ae740b0bbeafd2b5b98cc4551

Observation 0d952236-2223-4d8c-ad29-5afd2ade9bda · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation TruthfulQA: Measuring how models mimic human falsehoods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.117735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.117735Z digest=sha256:cf3e58cc9855b4d3e49989a635d9fbc060b222f9b018290b9f48c2cda27f4a98

Observation 9543c30d-1994-4996-b889-20b8a5679c43 · outbound

This paper cites Benchmark data repositories for better benchmarking.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark data repositories for better benchmarking

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.844219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.193135Z digest=sha256:6111b6c5dd62e3276ff0afe611d7819a5a1bf934c8844aa9a53153769aa219fa

Observation 8efbfdfb-c22c-4d53-8960-c00821637f53 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:25.663250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.299417Z digest=sha256:8d3bb46194185f11abe1e565efcda7791b2b9b3ee244d0083d90992d5b8de2fb

Observation 3e283af4-44c8-41ed-937e-5f530e085286 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.416530Z digest=sha256:125fa635e9bfd4c6c19da12fcd8b040ef16484ed2769dc9c14de51b770d393fc

Observation ab79f9d2-19ed-49ae-92f0-8831166b35e3 · outbound

This paper cites FActScore: Fine-grained atomic evaluation of factual precision in long form text generation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FActScore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:25.078670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.497724Z digest=sha256:f37b7f25db5830f1196b52a7b7bedbe4c0988b23d9b4cc6e20cfca61987228dd

Observation e86d7949-42bc-49ed-886f-eb7ae1559373 · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.563546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.563546Z digest=sha256:50ff7057335c02ff90d60f58506f47317a8d04dee862c5b99d6a048b4bb4ee30

Observation 970e71b2-80f9-4235-8a20-6bc47dc0dba2 · outbound

This paper cites BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BLEnD: A benchmark for llms on everyday knowledge in diverse cultures and languages

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.796108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.674542Z digest=sha256:ce061ebb5ffe218cbd141602153f379a393d395d5e532f42d5e4134d2af174ce

Observation b746cf2b-6230-403e-8f8f-4ca4330ee10b · outbound

This paper cites Extracting cultural commonsense knowledge at scale.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Extracting cultural commonsense knowledge at scale

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.597620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.809086Z digest=sha256:b782cdac564b0fa4fac6f76c755e0c2f9e2c67f3dc486d00950bde239c93e221

Observation dc5d50d0-ff3c-4ea3-9cf0-2eff798a6137 · outbound

This paper cites MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MixEval: Deriving wisdom of the crowd from LLM benchmark mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.233222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:13.880832Z digest=sha256:51adbdeabb6d5be5661fb7d1263e5273c330f333fe367fd41e51976468ef28ca

Observation a167dabf-4945-4a79-84f1-5217c4f6ea1f · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:13.986625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:13.986625Z digest=sha256:c9ac0dfa32c30fb247d28a2fb9660f2241b93d40ba09b22492217b437090e22d

Observation c92b6544-3460-43fc-b53b-4942d32311cf · outbound

This paper cites BBQ: A hand-built bias benchmark for question answering.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation BBQ: A hand-built bias benchmark for question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.031957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.078274Z digest=sha256:ddeacb79f7e63f6f9fd29308e77d8d982f3de1179cb60a33ceffed697fd76e29

Observation b115779b-e43d-467f-9b5d-3848fd76cf1a · outbound

This paper cites Survey of Cultural Awareness in Language Models: Text and Beyond.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Survey of Cultural Awareness in Language Models: Text and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.144191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.144191Z digest=sha256:5e85facfb91e6b56958dcd85e34084e1d6b15a6aa68f5b7d1829026c65da1e14

Observation 434e5316-32d9-4fce-937c-0ebbae8b572c · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Zero: Memory optimiza- tions toward training trillion parameter models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.722060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.231714Z digest=sha256:3456cd0d077cee60fbc2395d79920a6fde1a22b902197402d89f95fb07a597ee

Observation 431e4c29-6e41-4595-958d-42b3722aa736 · outbound

This paper cites NormAd: A framework for measuring the cultural adaptability of large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation NormAd: A framework for measuring the cultural adaptability of large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.401621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335556Z digest=sha256:bf484dcf8c74d26b0049652e1594e070f137c1dd50073ac9d9fed11a8d6822d7

Observation 5fc97c41-d2a5-466c-9cd3-876a0506eda3 · outbound

This paper cites DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.085948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.581006Z digest=sha256:613c4d667dd44ab6fa9bbbb2debef7872a2c6fae5e5a0e56a97cd4cad4e61538

Observation a931e440-3305-4057-84b3-49e2394adb4d · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:08:22.795446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691459Z digest=sha256:a74c708d75bc2cf70cdcc9431f1e95964881b237f2709072a8dba0a025083ef6

Observation 29c42e3f-8d2a-4402-9d66-c368ccbf24cb · outbound

This paper cites Kochenderfer.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Kochenderfer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.550842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.833051Z digest=sha256:9b329cc33f6929656678e4beef57b793c80fc8faa6be8a24f4296992a50f138e

Observation 32256750-95a7-49dc-bd84-fbcc2c0f5f5e · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.Commun.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation WinoGrande: an adversarial winograd schema challenge at scale.Commun

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.369793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:14.924761Z digest=sha256:a400c00a68616e49644ef214a16b3918b4e8036cf3c51c02ec1d7cff80960559

Observation cc6416ba-d929-4711-95aa-cd19b259b7d7 · outbound

This paper cites Social IQa: Commonsense reasoning about social interactions.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Social IQa: Commonsense reasoning about social interactions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.146589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015768Z digest=sha256:3de638ef34702c162ee3fa6b2744173084e4b1817108e5d10bb7764848830860

Observation 597877f1-8911-4dba-b89d-f29fbcd96adb · outbound

This paper cites Benchmarks as microscopes: A call for model metrology.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmarks as microscopes: A call for model metrology

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.940061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.107192Z digest=sha256:faa4cee33e136ea1b8c89a5b1bcbfdb2343c8d9b24c684270bda49447b4d7303

Observation 7db6d897-08bd-47d6-854d-be04c1ce887e · outbound

This paper cites Multi-fact: Assessing factuality of multilingual llms using factscore, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-fact: Assessing factuality of multilingual llms using factscore, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.694305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.229945Z digest=sha256:882ce9a18f25b1460a98319d2681adc7831582e3a9510e57a409c5f9fec9c7c1

Observation 1d3478be-d6f6-4cab-b9db-7783e1a7ad1d · outbound

This paper cites YourBench: Easy Custom Evaluation Sets for Everyone.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation YourBench: Easy Custom Evaluation Sets for Everyone

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.321570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.321570Z digest=sha256:6e9f678ae73f6acc958543bcca1b88f0b08de1310bae3b4908acc909e7138049

Observation ac95e546-01e9-4a9e-b94c-0804e1c866ba · outbound

This paper cites CultureBank: An online community-driven knowledge base towards culturally aware language technologies.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CultureBank: An online community-driven knowledge base towards culturally aware language technologies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.384899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440338Z digest=sha256:6ebf512516031846e3407c1b10c96eb5c623f5b5c6a14b73c30ea5c6961eb781

Observation 2632e0cc-e035-437f-9565-3dfdb648bb34 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.542569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.542569Z digest=sha256:e842eb4d9ca098677ee327ae3d3a4fcf730ee770246c1c69047b4a0c0c3cc41c

Observation 26e37416-51ce-415c-9642-e621f807dab6 · outbound

This paper cites KRX bench: Automating financial benchmark creation via large language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KRX bench: Automating financial benchmark creation via large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.170385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.639270Z digest=sha256:b4073f783ac50bf9d87f2002a928b1038cd3d5ed0878771c9f052ac2d3e5ee97

Observation ad0d7d1a-0d36-4340-9651-9a858d447735 · outbound

This paper cites Beyond Classification: Financial Reasoning in State-of-the-Art Language Models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Beyond Classification: Financial Reasoning in State-of-the-Art Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.739015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.739015Z digest=sha256:2c378dd2ed1e5742134de539758677c01db194724ab576cc7942dcf9a40db646

Observation 3201edf1-4f3e-4d99-a74a-b41de35e9f1c · outbound

This paper cites Multi-step reasoning in Korean and the emergent mirage.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Multi-step reasoning in Korean and the emergent mirage

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.945080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828572Z digest=sha256:1a074c13f31d5f37374b58e2f8864b875dd78f1b043ecc98d1179592630bda92

Observation 624cdaba-8ce7-4ba1-b8f4-cfe3cd287df1 · outbound

This paper cites KMMLU: Measuring massive multitask language understanding in Korean.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KMMLU: Measuring massive multitask language understanding in Korean

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.774825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:15.911316Z digest=sha256:4748df31f92444c6bcd16ba252e66210d57f6132de9592dae630f30e1998b132

Observation 062acff7-3666-48da-83a3-d3376048856b · outbound

This paper cites HAE-RAE bench: Evaluation of Korean knowledge in language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation HAE-RAE bench: Evaluation of Korean knowledge in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.551437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.003057Z digest=sha256:0f9b6ccab82e379515276806b34e534749a4f78fcdfe813b8244e5d64322ab21

Observation 51cdb724-54f3-47c5-9fc7-280e106df693 · outbound

This paper cites Challenging BIG-bench tasks and whether chain-of-thought can solve them.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Challenging BIG-bench tasks and whether chain-of-thought can solve them

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.366551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099651Z digest=sha256:817a131fa3cf1c392cc96d8ad9b75cc42b307c8eeedc21d1710ac6cd3f12cae5

Observation f7a0cb8e-86be-4b03-ac03-4c9eb57a57ad · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.178310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.178310Z digest=sha256:be44bc3621d6bfbacf5e9167d6a9c49695fa29c77a818bd78ec507cb85d02409

Observation 432bf232-706d-4036-b311-207966c0a5a0 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.285247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.285247Z digest=sha256:415b16340df89c476a5252bf513b5c37e8fd5bf2a55c707720828848ae6c833c

Observation 00d892b9-90c2-4f7b-a54b-7846b57cd4c5 · outbound

This paper cites Gemma 3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Gemma 3 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.366173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.366173Z digest=sha256:23185052dc6a47bc326ed981835e3a17da63412ae04bd778b81a695b6fddaf04

Observation 4b644f29-8f98-4946-a501-c1d937aec4bd · outbound

This paper cites Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Benchmark suites instead of leaderboards for evaluating AI fairness.Patterns, 5(11):101080, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.206356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.462836Z digest=sha256:e97a4957de5bf2d37eb348d3d779ba39eb5adbbeb44c7fec0455b631bff91d4d

Observation 6098c32e-3b67-48eb-92a2-d20244f6f03a · outbound

This paper cites SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation SeaEval for multilingual foundation models: From cross-lingual alignment to cultural reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.993735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586923Z digest=sha256:dd8d53f63efe94322fde9638d029f2d460deb776c6bebb3ed4a95d9dd72f18ff

Observation 1fd0d21e-54f8-4e60-83f6-9b48a90b3a76 · outbound

This paper cites KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.759287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.759287Z digest=sha256:eaf5c86631910801c2b167f8484e3e64b7267d2be853c7bb48402df84f59fe07

Observation 49db165a-1c43-4554-97df-544abd4e6381 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.840752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857171Z digest=sha256:da909ea48e2d1d4356d81f47b42ffffd36b1c2f539d90b09647d2a633250a5ab

Observation af13bd18-0319-4209-a00e-594f1aa283f0 · outbound

This paper cites MMLU-Pro: A more robust and challenging multi-task language understanding benchmark.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro: A more robust and challenging multi-task language understanding benchmark

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.656148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:16.951540Z digest=sha256:8e4fe702faeed31aaf8006d0abc05b80367311a1e9bea2adac6048adba39f46c

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:5360243bcb426381ee61fa6c256696750b90fa15f84b001033125fd03868334b

Observation a3199245-1f8c-4d78-958b-bed6e0247619 · outbound

This paper cites Qwen3 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen3 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.152295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.152295Z digest=sha256:6c936c926141c8fb4e922c0ed7e59300203de4726e01f7ee455e801157e928c4

Observation e2d4076c-8373-42b7-bbb3-a99321d88f7a · outbound

This paper cites Qwen2.5 Technical Report.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Qwen2.5 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.279546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.279546Z digest=sha256:335bac384b58f7cb4f7d39f0cc8c35890bbcd35e6d632dd9d45ed65439c01a01

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:ee7ac22c8c208619af993ab07b3f27bfdea1f044d5119fbabe696aa91140fafc

Observation bc172362-c9b7-4ef7-9b97-e23956a44a2b · outbound

This paper cites FLASK: Fine-grained language model eval- uation based on alignment skill sets.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation FLASK: Fine-grained language model eval- uation based on alignment skill sets

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.440683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447929Z digest=sha256:d84c3993530442b556800a7d635c31a58bb48d2ae73c5194162e1d0ed104f7a0

Observation 7da61289-4ca5-42d3-86b8-83542a2efdc6 · outbound

This paper cites GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation GeoM- LAMA: Geo-diverse commonsense probing on multilingual pre-trained language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.253520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503711Z digest=sha256:a5e943ba5bfee5efd962fbe9b591034f65364f3b09caa5e1992e9dd27a8487ec

Observation 725c2996-fa56-420c-b4f4-d17e36ae7af6 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.614168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.614168Z digest=sha256:fe5128acb07361b667d64d8096c476a53abd8a9c4e76215d33faf7225e7c93e6

Observation ad13bf2f-b95b-429a-96e4-7f7b91ef36e0 · outbound

This paper cites Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Maruf Hossain, Guang- Jie Ren, Kate Soule, Yifan Mai, and Yada Zhu

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.085125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692121Z digest=sha256:3132d4b4c76feb1bd3ee6e4b9667c11c58c9e11a372669df529fff1b7c4516d6

Observation 29a84389-3535-41e3-85c8-51ccf4318ed9 · outbound

This paper cites Task me anything.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Task me anything

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.975615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749553Z digest=sha256:bed680f89e40de52144506bbaadfcc3159cbf1402a649adc0f1b54f8d03acbd7

Observation 86d268ff-e827-4c90-a2e4-463393160202 · outbound

This paper cites Users can interactively explore the overall data distribution they are interested in.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Users can interactively explore the overall data distribution they are interested in

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.847818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836203Z digest=sha256:923286bc0424d4b9036f996401800f8dff829e62e4d5b97cfb3bfc2d5af48139

Observation e0de9c67-b6cb-47e7-acf9-528c7ec99ed7 · outbound

This paper cites By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation By reviewing samples, users can verify whether the dataset matches their needs and explore datasets suitable for their purposes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.739965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.905218Z digest=sha256:04324f4769f9a72f419b510556de2b694e48c003367c9d9f2445590b22950f75

Observation 7f2adf41-922e-4c81-9f4c-6dc5c22c83cd · outbound

This paper cites Is the Earth flat?.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Is the Earth flat?

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.606763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:08:17.955891Z digest=sha256:d78990cbee78655d8b50929c0f9382135f95218eca1837271a2e04cf2039ab1f

Observation 88c7fafc-7f12-49be-82bc-bc53c5530e79 · outbound

This paper cites an unresolved cited work.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.678445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.678445Z digest=sha256:7104f1af34efed4e9f74d3f5c56866d68704fb81875a1b6889548bdc65a71ff9

Pith citing papers

No inbound Pith citation observations are available.