Pith. sign in

Paper Citation Record · LEDGER

A Survey on Large Language Model Benchmarks

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2508.15361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15361 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:13:24.680011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:09:36.772485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7d9b07c0-a263-4ad8-bc88-041e23cf45a8 · inbound

QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation cites this paper.

QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation A Survey on Large Language Model Benchmarks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:24.680011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:13:24.680011Z digest=sha256:6cbceb1f612b6cccaefd22d9f74e36af017f62fad47289d00dc83e297a73df0a

Observation 1906bceb-aa64-460b-b17d-08968dbc2cad · inbound

Geometric Metrics and LLMs: What They Measure and When They Work cites this paper.

Geometric Metrics and LLMs: What They Measure and When They Work A Survey on Large Language Model Benchmarks

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:42.882488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:47:42.882488Z digest=sha256:44def57ea33073af0d533be9cbd04d77fcdf940bf7d4933c98c11608b7a983f7

Observation 07fe1144-ddf9-4df4-b08b-c07c736bcc2d · inbound

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports cites this paper.

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports A Survey on Large Language Model Benchmarks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:32:11.218878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:32:11.218878Z digest=sha256:34f908494dd4707b284faa7d9bfa1fcc7c617edf12fe660c3269ca803971dc38

Observation 5f724800-82da-4ebd-abcf-5d6369148920 · inbound

Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization cites this paper.

Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization A Survey on Large Language Model Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:20:46.009392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:18:52.540487Z digest=sha256:a317c67840973ffc30e851d161846b3457aa44e10af31b9fedd8c99ff43247c6

Observation db88210b-8749-4036-9c9c-4d316f149e1e · inbound

"LLM Agent Performance" Is Not a Single Evaluation Target cites this paper.

"LLM Agent Performance" Is Not a Single Evaluation Target A Survey on Large Language Model Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T05:04:13.908579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:04:13.908579Z digest=sha256:79c4e46cf20c2382eae72f59667ca809b58b0ec39493407e6f2ec2ef1de349de

Observation be0d1a4e-6ed5-4262-b13f-5a01d2087c8a · inbound

When control meets large language models: From words to dynamics cites this paper.

When control meets large language models: From words to dynamics A Survey on Large Language Model Benchmarks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:54:13.236200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T14:52:44.632671Z digest=sha256:36f31f9df53946bb2b61d0b6bb39d4d5c3c425599d889eeab58e488336119eae

Observation 21e40198-d230-4747-be18-015923311a6e · inbound

Designing for Error Recovery in Human-Robot Interaction cites this paper.

Designing for Error Recovery in Human-Robot Interaction A Survey on Large Language Model Benchmarks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:01.242242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:09:13.981398Z digest=sha256:1de2c11a55809dff7adb226e20ce608184141cc4a687de6574161756e09161d0

Observation c65da8d6-9ea8-4c10-819d-5b3c4d5b5b60 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints A Survey on Large Language Model Benchmarks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.428535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:8cd493b85184dcd77c17e19f65775acb543ab2dad24d0d0fdb4c383e4bad96d6

Observation 348809dc-1ec0-4909-ab9c-5f398ad4fd73 · inbound

Efficient Ensemble Selection from Binary and Pairwise Feedback cites this paper.

Efficient Ensemble Selection from Binary and Pairwise Feedback A Survey on Large Language Model Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.350817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:34:15.455035Z digest=sha256:7cab212b8349ca9731f7351548308baad1ff99ab6b6116ab0c65ba3c474bf684

Observation 902288b6-666d-4eca-9fbe-18693aa5ac5a · inbound

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels cites this paper.

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels A Survey on Large Language Model Benchmarks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:57:42.491127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T17:55:35.764347Z digest=sha256:0c6af5cce58cd58cd7d2a05355aa4ee3b9570d10ba3fde8e3db64f2b7c20be59

Observation 12885f74-9760-4eeb-b52d-540af269d061 · inbound

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling cites this paper.

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling A Survey on Large Language Model Benchmarks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:40.197295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T16:33:30.850113Z digest=sha256:159e779c03cb3d8d5895f307819c5ec53b2352f795e5089320baff54b3c85dc3

Observation a896fbe4-8c19-4610-b21a-9f6d2fe46397 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness A Survey on Large Language Model Benchmarks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.397631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:e6a88db7dc417464d8ae6088bfcb9dc85e56283dc2919c5828c57bf81e60b5e1

Observation ca851cbb-deb2-42f9-b633-41d537c38d3c · inbound

Can LLM Teams Play What? Where? When? cites this paper.

Can LLM Teams Play What? Where? When? A Survey on Large Language Model Benchmarks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.389918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T07:50:35.575641Z digest=sha256:633d81dc6e10f5a06499eb43843349513d44a03cbd3be34c54d059b7e74d2484

Observation 8a497ae1-a56e-4b3c-a908-df4ec41bc80c · inbound

MAVEN: Improving Generalization in Agentic Tool Calling cites this paper.

MAVEN: Improving Generalization in Agentic Tool Calling A Survey on Large Language Model Benchmarks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.138266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:c3de47dbe6a21cee86ac42a3bf21e9be85f7bb28fae4c21efc7281da8f96f2cb

Observation b4546810-6750-4ab7-ab34-ec27b1103b63 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A Survey on Large Language Model Benchmarks

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.617219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:0a1547f6e8cbfa02311788147bb6259f09aa4f49e7850ec4978c67370bbd71da

Observation 2dab6c41-64f5-493a-9cd3-c7222a0710bf · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.774324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:1cfce3ab2291573f2fcf52496d8198772593d8f8655fdce54c00565c217a931c

Observation 346564ca-59ff-4047-9b06-31b8fc02ed89 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.464446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.464446Z digest=sha256:ace80dc265687c07331492b9d1d78b8f6f15feab743a7158177e4da6948efa00

Observation d5ae7f1e-aa0e-4b00-962f-d06f9c6a82b3 · inbound

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese cites this paper.

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese A Survey on Large Language Model Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:58.465844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T13:03:58.617626Z digest=sha256:4e229d62d350705d7b6f2c49ef381c92e452733edd94e4491dfc5410bf0c6001