Pith. sign in

Paper Citation Record · LEDGER

A Survey on Large Language Model Benchmarks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2508.15361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15361 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:32:11.218878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:09:36.772485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07fe1144-ddf9-4df4-b08b-c07c736bcc2d · inbound

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports cites this paper.

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports A Survey on Large Language Model Benchmarks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:32:11.218878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:32:11.218878Z digest=sha256:9bfb493942bfec5962d78f8796b11ff76af9f0119ac85ca34dcf392c508a44dc

Observation 5f724800-82da-4ebd-abcf-5d6369148920 · inbound

Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization cites this paper.

Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization A Survey on Large Language Model Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:20:46.009392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:18:52.540487Z digest=sha256:639d475c1fd1612b23b2b5647bb0ab765dedb0739a77baa853511b0b3676dcc0

Observation db88210b-8749-4036-9c9c-4d316f149e1e · inbound

The Necessity of a Unified Framework for LLM-Based Agent Evaluation cites this paper.

The Necessity of a Unified Framework for LLM-Based Agent Evaluation A Survey on Large Language Model Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T05:04:13.908579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:04:13.908579Z digest=sha256:ab886ebc1490e9e24b64292455542232f354d5e921d2d89a5a11979136038eff

Observation be0d1a4e-6ed5-4262-b13f-5a01d2087c8a · inbound

When control meets large language models: From words to dynamics cites this paper.

When control meets large language models: From words to dynamics A Survey on Large Language Model Benchmarks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:54:13.236200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:52:44.632671Z digest=sha256:f89953f5ae0ff29c2150ba1ae772283ff9c8af6a02e9fe7c0c95830ea89201be

Observation 21e40198-d230-4747-be18-015923311a6e · inbound

Designing for Error Recovery in Human-Robot Interaction cites this paper.

Designing for Error Recovery in Human-Robot Interaction A Survey on Large Language Model Benchmarks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:01.242242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:09:13.981398Z digest=sha256:586788c8027d6976931504b4a2d3534b963e213241643c66b358cf85832ee9b6

Observation c65da8d6-9ea8-4c10-819d-5b3c4d5b5b60 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints A Survey on Large Language Model Benchmarks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.428535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:ca59cfa1ff5137341935664cf6489740d79572bcbf4254f086ee7cee3515ef9d

Observation 348809dc-1ec0-4909-ab9c-5f398ad4fd73 · inbound

Efficient Ensemble Selection from Binary and Pairwise Feedback cites this paper.

Efficient Ensemble Selection from Binary and Pairwise Feedback A Survey on Large Language Model Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.350817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:34:15.455035Z digest=sha256:4fa8593978f59ca7b319a245ea92cc22d736ab3b3bf8529bbc05c12e8fbf3417

Observation 902288b6-666d-4eca-9fbe-18693aa5ac5a · inbound

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels cites this paper.

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels A Survey on Large Language Model Benchmarks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:57:42.491127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:55:35.764347Z digest=sha256:467773f95b6a22b301abe6bc946f1ed5234c689710cf396fa9dc50903256b924

Observation 12885f74-9760-4eeb-b52d-540af269d061 · inbound

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling cites this paper.

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling A Survey on Large Language Model Benchmarks

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:40.197295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:33:30.850113Z digest=sha256:7f9ae5720d5c505da832877878d87fa5eaaa596ff5493d6ede79af9d59a6aa80

Observation a896fbe4-8c19-4610-b21a-9f6d2fe46397 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness A Survey on Large Language Model Benchmarks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.397631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:77fb29ab5aa38b81cd666d2ee279fd882e25f420031bde3343c94b8494677c27

Observation ca851cbb-deb2-42f9-b633-41d537c38d3c · inbound

Can LLM Teams Play What? Where? When? cites this paper.

Can LLM Teams Play What? Where? When? A Survey on Large Language Model Benchmarks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.389918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:50:35.575641Z digest=sha256:7f37187b383cc0b2d309ca6dae5ba4c8542b36af511cc301491bb85960de656e

Observation 8a497ae1-a56e-4b3c-a908-df4ec41bc80c · inbound

MAVEN: Improving Generalization in Agentic Tool Calling cites this paper.

MAVEN: Improving Generalization in Agentic Tool Calling A Survey on Large Language Model Benchmarks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.138266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:42:11.459778Z digest=sha256:d02ebefa05b7e2481b85b308527514694fb35c8b7db43a109e06c6e5019a827f

Observation b4546810-6750-4ab7-ab34-ec27b1103b63 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting A Survey on Large Language Model Benchmarks

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.617219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:989a12d4bf86a2a701b68e152de8342ada784d0b027cf67f93cf085350a04e20

Observation 2dab6c41-64f5-493a-9cd3-c7222a0710bf · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.774324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:c7da9b604457b04e88f004ace54b0ced1334c9402cf2f7dc6fc07da3e44f85d0

Observation 346564ca-59ff-4047-9b06-31b8fc02ed89 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.464446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.464446Z digest=sha256:39639d230fd07db25883666746723948cb357644abe496c4cfd591ec3e3162d1

Observation d5ae7f1e-aa0e-4b00-962f-d06f9c6a82b3 · inbound

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese cites this paper.

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese A Survey on Large Language Model Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:58.465844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T13:03:58.617626Z digest=sha256:b07c94dab8bc36020fd5220b9d4e96458077c530a68e2fc6516fb8c029270f26