Pith. sign in

Paper Citation Record · LEDGER

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.06745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06745 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.901468Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T21:09:30.462709Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T21:12:08.583301Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9971c8b-1223-4725-a55e-fad83c9db937 · outbound

This paper cites llava-7b 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities llava-7b 3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.264742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.733490Z digest=sha256:d2f50f41c5a36d3887910446ae5a353873a4fa36d790835ba71ff8733737e5d1

Observation 24ae2197-60bb-4d1a-96e5-c76da6fe78af · outbound

This paper cites A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.111732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.894669Z digest=sha256:2769c8ee2dbacdbb70d1cb39ce9205a5d317627bd84d89a43ee8424d8da09965

Observation fdcd4fad-ad4c-4133-9df6-9b3e23eefa6b · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.100731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.898103Z digest=sha256:3b4c1008cbac03d82990bc1f0956ebfe25361939ce2b33447073a1f2dd9bca44

Observation 03900b93-7987-4443-9004-1852de35bb7c · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.090126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.901468Z digest=sha256:f471cb1d86f561fff71214c9aea179e53d0c6efa37e8ceacdfdc8195f9a251d5

Observation a73e2cf4-f347-4d47-9b62-aabc3282662f · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.693587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.693587Z digest=sha256:51c05c86f59764cd0c66019b2c0b925733e86c858d0c1ff296503db45aaa91cc

Observation 61dd9087-020f-4142-9b45-bfd8f4805819 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities tinyBenchmarks: evaluating LLMs with fewer examples

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.697517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.697517Z digest=sha256:b0474bed3e7a09cd0f79137777451121d4d980677ec99e2c1e5d97892edf75bd

Observation e589acc7-12b6-4795-a779-6ffc28164486 · outbound

This paper cites Efficient Lifelong Model Evaluation in an Era of Rapid Progress.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Efficient Lifelong Model Evaluation in an Era of Rapid Progress

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.008433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.701768Z digest=sha256:65d9a8bb3967646e53b5d7a20a21652fbee49b5df10dad68d29997acebf4303d

Observation c1183266-9857-418a-a685-8f5b28542b99 · outbound

This paper cites When is it Better to Compare than to Score?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities When is it Better to Compare than to Score?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.709449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.709449Z digest=sha256:1f4e2446ba4e92e29173aaed1741e27d1dc3bd4bdeae420b0c5427d699c58c63

Observation 0a0c7732-bf17-4a9b-a8d5-76da4ebe76bf · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.717055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.717055Z digest=sha256:6a842ff0ac190e3f2f08d9b25a083265f66288f64e4735a0abe53a8a658c97d6

Observation 1e01ed4f-49e3-4e02-820c-b462d300196d · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.721018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.721018Z digest=sha256:0450fbdc84ac433a9124bc2e0413387880306c0893ca4bc66d4dc930c09dfeff

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · outbound

This paper cites Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:d3e611b1a18724e0f39d563d601b576a265bf4f36d08732691b99e4118dcfc19

Observation f212106c-a053-4858-99fc-9fc7dff01d6e · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-11T19:24:17.729268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.729268Z digest=sha256:91cfab12a04833d59a925527d059912ea21eef2e3de63a9721536f8d52627524

Observation 6051a52e-b1f5-4209-ba17-827c9e4d50ec · outbound

This paper cites internlm-xcomposer 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities internlm-xcomposer 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.255171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.737268Z digest=sha256:d37dc63e4564db20f22ae1743604fa5945db7502d29d5cec3037d25593535b8c

Observation a0ea9fba-5106-4252-beda-39ee94da3414 · outbound

This paper cites idefics_80b_instruct3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics_80b_instruct3

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.245573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.740911Z digest=sha256:8bf6a80d87bad7353325a8247f691fbc986681f7c67bb8aa1ae872bd533ce215

Observation cd287553-a19d-42ad-bac6-52eef6bc42d5 · outbound

This paper cites idefics-80b3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics-80b3

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.235013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.744497Z digest=sha256:1238a8e53a24956c3f735e7fb8e6756b88ba9b796614d65496fcef12d9e50abf

Observation dde5673d-dbd5-4dcc-b881-a5277231c45c · outbound

This paper cites claude3_sonnet3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities claude3_sonnet3

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.223053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.748186Z digest=sha256:ad34727c7734b5c94c6f65ac63897dfb87dc2839bc0778d80e403539dfdbee60

Observation 3a770b11-7318-4ed6-a2a8-462d5ce5f172 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.212567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.752064Z digest=sha256:c231d24e3e6824c40b9b7874bbc6722f22b4edd60e91c94bc426d38b65a49b67

Observation 51a05230-5484-440f-94a9-8464647a9396 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.190924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.760166Z digest=sha256:c4493ab193eac448e28918a77af4654d5dd0b710cf0588a733748a3f6e1841ec

Observation 48f12b08-3461-42bd-8a28-ac6ab2969a12 · outbound

This paper cites Who is the Byronic hero?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Who is the Byronic hero?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.180056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.763565Z digest=sha256:eaef7ed77ec84180dd1caa3776eece11e61b351a855cb1b240d3ebab4e400b94

Observation d3ce9320-7062-4fa4-a463-8a6d5331f4dc · outbound

This paper cites palmyra-vision-3 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities palmyra-vision-3 3

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.169222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.766901Z digest=sha256:705882e580a9963b7265c3a28b18d7242e087489ff78a28b6d51ab2af225c0f4

Observation 6e757455-8564-43db-b8f6-733c7f5dec94 · outbound

This paper cites _ bought a….

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities _ bought a…

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.158438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.770148Z digest=sha256:c0de8ed5d36be6c4ff5c65a149cef9bf71a3fa5bb50686d3a95316912dc9d6f6

Observation 2209fb4d-30d5-4d2c-86d2-de1240b84379 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.201532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.773522Z digest=sha256:0f37cb7771aa60a2c9d0fb812b570857fe8641fac7514a803ff63999dfa87ce9

Observation 39b74299-4a1e-4453-b2f9-95d8a1c956b2 · outbound

This paper cites decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.147195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.776707Z digest=sha256:69c2095a650683f04cfe30d8629efb0e291cd48cf69a24ff3014b20b5c871359

Observation aa7140e3-5b19-462e-b605-662464a3b8c0 · outbound

This paper cites high-quality.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities high-quality

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.135343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.883115Z digest=sha256:945b7bb5dc568e376eca0d78ab486b61740af751d8649c75f137892ede8031c2

Observation 7d46f22b-a73f-4e5a-891b-c4695d6d0f8f · outbound

This paper cites These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.123424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.891223Z digest=sha256:cca883797ab870be96799c81305d0741ef3f7432624f0127393c1b66d239467d

Observation 7b64b2cd-5d17-4a7c-9983-a194ef74fa47 · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.713429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.713429Z digest=sha256:71966b34e4dd9a5fdf125b0432ef6dc5e1f358fabe11340fc1777aad781b1c74

Observation 50460282-d1c5-4bbb-818c-f5cb97887a06 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.689075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.689075Z digest=sha256:a119915bdb2d78b4216f7ebb1c601e28a10a9d306a0aefcef6a31c4466e48f8c

Observation ab076295-4cb9-4034-a2f6-215f70230b11 · outbound

This paper cites Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.676182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.676182Z digest=sha256:75375e7c2719615e56c935d5848bdb2312c8bc2e886e1c6f788304254a64bc7d

Observation 4ae36895-c721-44f3-bfd2-e0190f64c863 · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Data Contamination Report from the 2024 CONDA Shared Task

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.705665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.705665Z digest=sha256:587e2b739c20f1268de355b81c689881abf2e3ce15befbe7427cf662e704e53c

Observation a964547f-d5ea-47ed-bef3-e0f263c30130 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.680849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.680849Z digest=sha256:0f8d51b40d209528ba4536dc91e850d6227cc175bee77c2a4e55d8a5b55f6577

Observation c8bc9018-c127-4fa4-aca2-b78c1748846f · outbound

This paper cites Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.057411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:17.684622Z digest=sha256:c0fb8561a70894a41eeff57b90e3ef0f49bb56bfbe59d7159b6baea8e50f78fa

Pith citing papers

Observation 82bba862-11c9-4694-8553-0fa2e4bc86c0 · inbound

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model cites this paper.

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:12:08.585633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T21:09:30.462709Z digest=sha256:4d6e9f7ccf73194725c06ca3dc58e9713e3719e7dbc86e1b42d366538e2bfc72