Pith. sign in

Paper Citation Record · LEDGER

Efficient Benchmarking of Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2308.11696.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11696 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:40:23.474799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:37:14.599642Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aec9314c-16e6-4227-94d2-049a26902099 · inbound

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models cites this paper.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models Efficient Benchmarking of Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.131464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:40a5da048937ba23226faa06dd00d551133560c53f4dec2daaa342ef91e37e7f

Observation b4d52ffe-f59f-4b81-baae-5b941819db4e · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Benchmarking of Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:23.474799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:23.474799Z digest=sha256:6565fb66f8f66ace57b1ac0178a4867702eb4018ade42998e4fe26b61c097526

Observation c8e958b9-687b-4b4d-9dad-93e531d2c948 · inbound

Metritocracy: Representative Metrics for Lite Benchmarks cites this paper.

Metritocracy: Representative Metrics for Lite Benchmarks Efficient Benchmarking of Language Models

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:34.399951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:50:34.399951Z digest=sha256:579d603707d1e3c6d55390256519565d708866cd8d3450955b680d298da663aa

Observation a20a6729-266d-4ba0-812b-bb669541417c · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Efficient Benchmarking of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:01.663487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:01.663487Z digest=sha256:1320e2a39e6c01d813615b3b60ecd96e6da7ae87174080e2a27062f08bd08fa4

Observation 24dd9ceb-dede-4a83-8364-593207157b50 · inbound

Query-efficient model evaluation using cached responses cites this paper.

Query-efficient model evaluation using cached responses Efficient Benchmarking of Language Models

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:56:00.055951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T00:57:26.494031Z digest=sha256:8329867634b5c481465af81036224e13bb5c3e21c6a69dc7049498591d3c5aae

Observation deaf8605-9742-4671-a893-d242a5ca1f18 · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Efficient Benchmarking of Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:52:49.477606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:389b326fac1e8cfa4dd607f07629a5bb5f67e6a3c06d439e5cb523b4ac7f1cf4

Observation 197d6ccd-ded6-4336-bc8d-c193e0de6b86 · inbound

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs cites this paper.

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs Efficient Benchmarking of Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.486175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T16:56:29.536912Z digest=sha256:d6e115e7fdbaf478ebe1e8e741852ea6b89ebd7dcef0167990ab93aebc94d80b

Observation 7a49a4ea-4add-4ebd-b9be-34f0e11c3acd · inbound

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces cites this paper.

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces Efficient Benchmarking of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.601224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:57:20.848868Z digest=sha256:624be086d30fa614494cc2d2967626dc8cb209588f91c5494d70ccfbcdb893b4