Pith. sign in

Paper Citation Record · LEDGER

Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.07545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07545 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:30:40.772943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:17:57.221151Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cbee2a8d-6a73-4c03-bd79-2bfedb1e788f · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:14.000184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:93063a8b864dd407e1e6c31a973be8866ccad9288dcc0bf87b90c21eb8602aa6

Observation c79cca48-3ef8-47d3-b589-8fb08bb09684 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.386774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:16f86f364b20e76552113e0ba50a4258cc229d374eaf512415295c50f643e3f7

Observation 54984122-afbb-486b-a687-06d47e20df1d · inbound

CARROT: A Cost Aware Rate Optimal Router cites this paper.

CARROT: A Cost Aware Rate Optimal Router Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T05:30:40.772943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:30:40.772943Z digest=sha256:879ba9fd785afc46c3658c47e5dd6f36311d5826f78f9aad0ca7628ee280cc16

Observation d031f4f1-005b-4baf-8be1-021ff3e3a4fa · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.030772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.030772Z digest=sha256:edefc21dbe35e82cae9cef63bf850df082407cb436925eace8404acda1896cd6

Observation 47cbe7f8-93d0-48a8-82ed-ba153745657c · inbound

DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation cites this paper.

DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:57.059533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:57.059533Z digest=sha256:c95716aac02732064b0b0a81a75f146a843c41a0100b9db63b61708c3f1f4278

Observation 2a45c7a3-b8cd-497e-8811-e2f50673d835 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:40.029820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:40.029820Z digest=sha256:cda22b5212faad0ea6619683213e64e190c1f37cccde37a830dc81374ea54c94

Observation 8fe4b995-1c8b-407f-b102-12b70a3744c7 · inbound

SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version cites this paper.

SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:17.759571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:57:17.759571Z digest=sha256:2093bc8de6455ad1f78d3ee3ff5f11d3bd8097d001d9c219b848b7f25d2795ab

Observation 7727e9bb-03d0-4a6f-908e-27a44bdca87b · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.530217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:3d6d8767de66f2be813b9b3294e677e722808bdc8237e1dde118d244b2073181

Observation 437054b5-47cb-4991-b811-c6b8ac66e482 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:00:51.716995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:e2499e7e69bc73e8131ed96fc9b861816be4e798c0aa9acef0300e97a9760a31

Observation 9abee72f-ed8b-4cc7-bed1-4b05ec0a9710 · inbound

KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration cites this paper.

KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:05.759485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:05.759485Z digest=sha256:38958604fafded347b527da6b097b27945b7ef60b3468f27c959bf1f3b0add24

Observation b9726ac4-c6f7-42ef-9a95-0595c74e9aae · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:9bdcb8d331e89519e1285f6d4fd06501ed8b74274275b1fefdefdc0e54ec5296

Observation 5a9037bc-3af2-4654-baba-8e4918c7f508 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:08:25.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T23:07:38.352198Z digest=sha256:d226e265c5818960502a6f52f6889d14475fc3916e84e1b79c70f59ef109cb4d

Observation 9cd40c27-7468-4b36-8464-767b41362652 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:10.686785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:05:10.686785Z digest=sha256:f933afa92bc73d6e78d9e4f4e34529783ad6978d8012cf89f56fdefa734a4881

Observation 7d54dfb1-535a-410c-bd23-840e74d9d914 · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:17:57.222547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:9ed2e05a503264ec347196c4c66e24d3d5bc78d0590809890ecc0ff323c68356