Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.901468Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.06745.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.901468Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T21:09:30.462709Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T21:12:08.583301Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9971c8b-1223-4725-a55e-fad83c9db937 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities llava-7b 3
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24ae2197-60bb-4d1a-96e5-c76da6fe78af · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fdcd4fad-ad4c-4133-9df6-9b3e23eefa6b · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 03900b93-7987-4443-9004-1852de35bb7c · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a73e2cf4-f347-4d47-9b62-aabc3282662f · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61dd9087-020f-4142-9b45-bfd8f4805819 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities tinyBenchmarks: evaluating LLMs with fewer examples
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e589acc7-12b6-4795-a779-6ffc28164486 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Efficient Lifelong Model Evaluation in an Era of Rapid Progress
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c1183266-9857-418a-a685-8f5b28542b99 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities When is it Better to Compare than to Score?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0c7732-bf17-4a9b-a8d5-76da4ebe76bf · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e01ed4f-49e3-4e02-820c-b462d300196d · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f212106c-a053-4858-99fc-9fc7dff01d6e · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6051a52e-b1f5-4209-ba17-827c9e4d50ec · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities internlm-xcomposer 3
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a0ea9fba-5106-4252-beda-39ee94da3414 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics_80b_instruct3
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd287553-a19d-42ad-bac6-52eef6bc42d5 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics-80b3
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dde5673d-dbd5-4dcc-b881-a5277231c45c · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities claude3_sonnet3
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a770b11-7318-4ed6-a2a8-462d5ce5f172 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51a05230-5484-440f-94a9-8464647a9396 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 48f12b08-3461-42bd-8a28-ac6ab2969a12 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Who is the Byronic hero?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3ce9320-7062-4fa4-a463-8a6d5331f4dc · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities palmyra-vision-3 3
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6e757455-8564-43db-b8f6-733c7f5dec94 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities _ bought a…
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2209fb4d-30d5-4d2c-86d2-de1240b84379 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 39b74299-4a1e-4453-b2f9-95d8a1c956b2 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aa7140e3-5b19-462e-b605-662464a3b8c0 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities high-quality
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d46f22b-a73f-4e5a-891b-c4695d6d0f8f · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7b64b2cd-5d17-4a7c-9983-a194ef74fa47 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50460282-d1c5-4bbb-818c-f5cb97887a06 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab076295-4cb9-4034-a2f6-215f70230b11 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae36895-c721-44f3-bfd2-e0190f64c863 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Data Contamination Report from the 2024 CONDA Shared Task
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a964547f-d5ea-47ed-bef3-e0f263c30130 · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8bc9018-c127-4fa4-aca2-b78c1748846f · outbound
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 82bba862-11c9-4694-8553-0fa2e4bc86c0 · inbound
Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.