Pith. sign in

Paper Citation Record · LEDGER

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.06745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06745 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.901468Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T21:09:30.462709Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T21:12:08.583301Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9971c8b-1223-4725-a55e-fad83c9db937 · outbound

This paper cites llava-7b 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities llava-7b 3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.264742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.733490Z digest=sha256:edd3c399a76dfb5fee3bcb5034deeec1e1155c351d8d910283ac86d333cabee6

Observation 24ae2197-60bb-4d1a-96e5-c76da6fe78af · outbound

This paper cites A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.111732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.894669Z digest=sha256:a735cdffd1b03c38ac26aca56d18140654d6c8af78a7ab07fc89bb56d22c66fe

Observation fdcd4fad-ad4c-4133-9df6-9b3e23eefa6b · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.100731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.898103Z digest=sha256:f7eab52016cc473b45fa87d47d88ff9ef3002ec3f7a4d9fd746e286cdd509827

Observation 03900b93-7987-4443-9004-1852de35bb7c · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.090126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.901468Z digest=sha256:851469ee6c5fa1b76f298492c6afac01d1339c813e532a21444791d0133512d9

Observation a73e2cf4-f347-4d47-9b62-aabc3282662f · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.693587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.693587Z digest=sha256:18c3cda9dd8a8d0aeccca0195d46780942eb87f0c9ae751b884988771f842a45

Observation 61dd9087-020f-4142-9b45-bfd8f4805819 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities tinyBenchmarks: evaluating LLMs with fewer examples

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.697517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.697517Z digest=sha256:b0474bed3e7a09cd0f79137777451121d4d980677ec99e2c1e5d97892edf75bd

Observation e589acc7-12b6-4795-a779-6ffc28164486 · outbound

This paper cites Efficient Lifelong Model Evaluation in an Era of Rapid Progress.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Efficient Lifelong Model Evaluation in an Era of Rapid Progress

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.008433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.701768Z digest=sha256:4184e973d1dde8fd9d6462998ee628ab77a67c2ad48fd1b78b886afa2bc7f9e7

Observation c1183266-9857-418a-a685-8f5b28542b99 · outbound

This paper cites When is it Better to Compare than to Score?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities When is it Better to Compare than to Score?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.709449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.709449Z digest=sha256:1f4e2446ba4e92e29173aaed1741e27d1dc3bd4bdeae420b0c5427d699c58c63

Observation 0a0c7732-bf17-4a9b-a8d5-76da4ebe76bf · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.717055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.717055Z digest=sha256:6a842ff0ac190e3f2f08d9b25a083265f66288f64e4735a0abe53a8a658c97d6

Observation 1e01ed4f-49e3-4e02-820c-b462d300196d · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.721018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.721018Z digest=sha256:0450fbdc84ac433a9124bc2e0413387880306c0893ca4bc66d4dc930c09dfeff

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · outbound

This paper cites Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:d3e611b1a18724e0f39d563d601b576a265bf4f36d08732691b99e4118dcfc19

Observation f212106c-a053-4858-99fc-9fc7dff01d6e · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-11T19:24:17.729268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.729268Z digest=sha256:91cfab12a04833d59a925527d059912ea21eef2e3de63a9721536f8d52627524

Observation 6051a52e-b1f5-4209-ba17-827c9e4d50ec · outbound

This paper cites internlm-xcomposer 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities internlm-xcomposer 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.255171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.737268Z digest=sha256:6952ec1d7b1512007ea6d3b8464586a7b4f676264fb18035ca7d0515ec0e99fa

Observation a0ea9fba-5106-4252-beda-39ee94da3414 · outbound

This paper cites idefics_80b_instruct3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics_80b_instruct3

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.245573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.740911Z digest=sha256:1bb06ca1a40ed4e85d59cf87f1170f02de9764b657ad6405e0525d3ccbd3c982

Observation cd287553-a19d-42ad-bac6-52eef6bc42d5 · outbound

This paper cites idefics-80b3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics-80b3

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.235013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.744497Z digest=sha256:1b159d5d1f9e37df7ad6946afb6c04c982cbfb497e31ca2e6fcf2946d9a3606b

Observation dde5673d-dbd5-4dcc-b881-a5277231c45c · outbound

This paper cites claude3_sonnet3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities claude3_sonnet3

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.223053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.748186Z digest=sha256:ad846ccfe46bbd668c911ec100718dc7e10b9d74c5cafaedfde3abc5c36c8746

Observation 3a770b11-7318-4ed6-a2a8-462d5ce5f172 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.212567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.752064Z digest=sha256:59a4ffd4b692f7e218acf5dddc8afd8edb4bd4e42f63d0b2b12771dce9ef0119

Observation 51a05230-5484-440f-94a9-8464647a9396 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.190924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.760166Z digest=sha256:91deedef61f1153f0f21f29e7b7a50209ae0a841083948081c000c237741bfca

Observation 48f12b08-3461-42bd-8a28-ac6ab2969a12 · outbound

This paper cites Who is the Byronic hero?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Who is the Byronic hero?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.180056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.763565Z digest=sha256:515957a3274a12253ded28a3b36883c7ac945d8d88c9a909f80be7f56c424ca2

Observation d3ce9320-7062-4fa4-a463-8a6d5331f4dc · outbound

This paper cites palmyra-vision-3 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities palmyra-vision-3 3

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.169222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.766901Z digest=sha256:8276c23a94f7751b44799d91f75bc977f3643af225c09ea8f0cfcff939e61cc3

Observation 6e757455-8564-43db-b8f6-733c7f5dec94 · outbound

This paper cites _ bought a….

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities _ bought a…

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.158438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.770148Z digest=sha256:ec45910a573a98ff45b8c11900539019cce3be86eb8138cd8a1f9746094f335f

Observation 2209fb4d-30d5-4d2c-86d2-de1240b84379 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.201532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.773522Z digest=sha256:115ac18f6968179c0a08f0241a393ddb57be72214b31be8ba7194b4f334ee272

Observation 39b74299-4a1e-4453-b2f9-95d8a1c956b2 · outbound

This paper cites decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.147195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.776707Z digest=sha256:49846773eefcbb310d777c5f0b51634a53baaad53a19c333289094a7ade714ab

Observation aa7140e3-5b19-462e-b605-662464a3b8c0 · outbound

This paper cites high-quality.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities high-quality

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.135343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.883115Z digest=sha256:75ab4831cf5442dbc485c387409ebffc08c786c4df844729698bc4882a663679

Observation 7d46f22b-a73f-4e5a-891b-c4695d6d0f8f · outbound

This paper cites These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.123424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.891223Z digest=sha256:1488bed7c6118950db0776cf400320632114b83e0f855a7f6522bcebabf93ae9

Observation 7b64b2cd-5d17-4a7c-9983-a194ef74fa47 · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.713429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.713429Z digest=sha256:71966b34e4dd9a5fdf125b0432ef6dc5e1f358fabe11340fc1777aad781b1c74

Observation 50460282-d1c5-4bbb-818c-f5cb97887a06 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.689075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.689075Z digest=sha256:a119915bdb2d78b4216f7ebb1c601e28a10a9d306a0aefcef6a31c4466e48f8c

Observation ab076295-4cb9-4034-a2f6-215f70230b11 · outbound

This paper cites Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.676182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.676182Z digest=sha256:75375e7c2719615e56c935d5848bdb2312c8bc2e886e1c6f788304254a64bc7d

Observation 4ae36895-c721-44f3-bfd2-e0190f64c863 · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Data Contamination Report from the 2024 CONDA Shared Task

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.705665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.705665Z digest=sha256:587e2b739c20f1268de355b81c689881abf2e3ce15befbe7427cf662e704e53c

Observation a964547f-d5ea-47ed-bef3-e0f263c30130 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.680849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.680849Z digest=sha256:0f8d51b40d209528ba4536dc91e850d6227cc175bee77c2a4e55d8a5b55f6577

Observation c8bc9018-c127-4fa4-aca2-b78c1748846f · outbound

This paper cites Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.057411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:24:17.684622Z digest=sha256:38e7d183070681032105b78d655c562da1bbcbb6ee91f28c55683e725bb80de8

Pith citing papers

Observation 82bba862-11c9-4694-8553-0fa2e4bc86c0 · inbound

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model cites this paper.

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:12:08.585633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T21:09:30.462709Z digest=sha256:d3ee6c86799f7b17b60c6984b8594011c636695037df0e58e8c8edc00a5b337a