Pith. sign in

Paper Citation Record · LEDGER

Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.19450.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19450 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:24:25.344894Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.903971Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da975d38-1082-4a7f-8cb6-a4ee7715d135 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.488804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:431bf3857e3f02edc8880378245ad1937da6c49ec1c9bbb96599678433a2a937

Observation eab653d2-d2a8-4d6f-a349-0f105c367b7c · inbound

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models cites this paper.

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:42:12.049878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T00:42:11.891829Z digest=sha256:6487e26160d2288c66f14bc8857e4873e1c93dfd4d3623b37a8633fddc44accd

Observation b0cc88f7-7a2c-4e29-b638-15b9039612f3 · inbound

Causality can systematically address the monsters under the bench(marks) cites this paper.

Causality can systematically address the monsters under the bench(marks) Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-08T20:24:25.344894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:24:25.344894Z digest=sha256:56327300f914dadef745cf18e15247e32d52269b07eb1f5805e0b7df2ababefe

Observation f329c6bc-ee4e-4f40-a39c-393130f64d25 · inbound

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark cites this paper.

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:07.629687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:07.629687Z digest=sha256:44e7d75ceae6a53aaf139eb401c137667fae328376bbe3eec07c28e684dd125f

Observation 3826905d-f638-4ed1-aa2b-97dc4f193be1 · inbound

Probing for Arithmetic Errors in Language Models cites this paper.

Probing for Arithmetic Errors in Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:59.669192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:58:59.669192Z digest=sha256:ecb4476dc1f3a19974d7c36ac36ac10a38dc1452504798997ddfdca3e75ebf6e

Observation fac53161-6d61-45da-b0ef-7080273f2620 · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.703761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:e79d3ffd546079775bfd5ceb4b4c370b2e341f6373bb794c0d93314ab5e3e8e1

Observation ac9ef06e-0d3a-4efc-a2b4-f1446c679a46 · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:00.854306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:c4e6ee20d7a149b044f838c082dba5f0d2f83b74cd161ae3dca97c831ffd98a8

Observation 36339f75-5eda-42d9-a948-c9fb76348a22 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.713534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:06:08.545334Z digest=sha256:8d1b9ca2cf71510a4e7038669ab07fd92cae55b29be732d87566416318b21c20

Observation 6374d269-33f9-4537-8891-e3976752897d · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.862792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:18:51.642405Z digest=sha256:174256a376ccf41fee4a8600e79042941b211e96d6f700a12ba7d27d4538dfe8

Observation 3e85165a-702b-4235-8e8b-4d0b8e12e8b4 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T17:25:30.907575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:25:30.907575Z digest=sha256:88bde24c08f5f00a298acd2dde77fb0d9b8286499bea9991b6fc5031a6c693c4

Observation e40167f0-3073-4d43-b3d3-dcb104f12179 · inbound

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling cites this paper.

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:40.203346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T16:33:30.850113Z digest=sha256:77ac3fb3385d9d714b0e462cbc7df5c2203d85d847db6246b0e6da15feed5975

Observation 730dc534-79cb-4f5f-94b6-fd545a8a024e · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.905611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:c1c1f861ad8f1d1e63b5c4aa247546158c73ebf2ef23b1798d1dfc35b8a848ce

Observation 9f4da5cd-79e4-435b-9d6a-4608fea40c7b · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.483003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.483003Z digest=sha256:e10ff7175151c5bddf1110de42bcd401732b1991b2b5ca650bd6cb5a44316346

Observation 6f402b3a-c51d-487d-ac1b-8d2bd36d255e · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.704626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:35:32.667865Z digest=sha256:b6115f911a1678556863da96bbb04b78decd6ed7a280c6e6cecb197f7764669c

Observation f775b9b4-0b01-49ef-b1c7-7cf72a1db984 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:35:02.975384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:35:02.975384Z digest=sha256:b5a830e96c88bc4184b9aed614eba201e53cf65e497b7bd680a132a144965817