Pith. sign in

Paper Citation Record · LEDGER

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2602.20629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20629 v3

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:18:04.034797Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:44:33.099337Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 823e46f8-6b12-4c13-8a04-f35076b474e2 · outbound

This paper cites write newline.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.028080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.028080Z digest=sha256:3259a5d4e917126953ef3c74bb2fa01009e1bb7a7a8caa0794b52742bd6fdaf2

Observation 71b434c8-f418-448f-a43d-cab9ca6efe4c · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.095450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.095450Z digest=sha256:7babb533a2303d9c574bf8274b7e463094c8235cbe3d2d14fa73c0a2e9c5c1e3

Observation a214a5ff-e6b1-4aae-9b70-6a1422b6c415 · outbound

This paper cites W., Keutzer, K., and Gholami, A.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs W., Keutzer, K., and Gholami, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.156429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.156429Z digest=sha256:74d850f64abc7d6bd6bcc53168af04b8f2960809ba2210981b02b30901500e74

Observation 4038bffb-dc9d-4c3e-9f82-90084f56cd57 · outbound

This paper cites Analysis: An Introduction.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Analysis: An Introduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.254788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.254788Z digest=sha256:e0e5deace23e341da56ba45c3432ee419f575e5c2d8627ee55ce2fea692acbdc

Observation d7b984f8-24e4-4a0d-a97e-d191e2fdd598 · outbound

This paper cites and Zastawniak, T.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs and Zastawniak, T

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.330783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.330783Z digest=sha256:c267cf6e31e8d2298f7db77d9a3591df0a3f2b5306a610aaee54ef382751a335

Observation 1ad1334e-4ff0-48fb-a4c4-b12cf92f7cb6 · outbound

This paper cites Stop summation: Min-form credit assignment is all process reward model needs for reasoning.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Stop summation: Min-form credit assignment is all process reward model needs for reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.400505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.400505Z digest=sha256:91013e790cc86f0ce58942907c5e915de0cccf362e7299c7345e2df0b82fea19

Observation a8748a5b-1243-4c3b-80d3-cea3d1d60de0 · outbound

This paper cites U-MATH : A university-level benchmark for evaluating mathematical skills in large language models.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs U-MATH : A university-level benchmark for evaluating mathematical skills in large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.476118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.476118Z digest=sha256:2d23f5739380b5853a47638dad581f24c30c9d40d9ec84829c8a3b7c1f4a43ac

Observation edb0e50d-7403-446c-8d1a-8a713b02f0aa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.574531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.574531Z digest=sha256:2e20bcfcf2b0bee2657db6cf6403ac2cb52894bd69cd3f7c481bd72fca7da07c

Observation f4cd5d79-af42-448c-802d-ec619a407558 · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.675070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.675070Z digest=sha256:c71fff21d8a89db14a85689d23f786756e8cb82b32901374eed4252f330f7406

Observation e848e829-fd52-4441-883a-00bcd7e234b8 · outbound

This paper cites The Lean theorem prover (system description).

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs The Lean theorem prover (system description)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.752228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.752228Z digest=sha256:826d86d3456f14f7bf52ed73a24584d41921b34b02645588c61e9207047d0fbd

Observation 7443fbf5-5637-40ca-a343-a8dcf864bbb7 · outbound

This paper cites D., Nikolova, K., Georgiev, N., Kalinkova, V., and Ismoldayev, M.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs D., Nikolova, K., Georgiev, N., Kalinkova, V., and Ismoldayev, M

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.859238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.859238Z digest=sha256:818c76978dc5f6209a872e2d99e7efcb926704f44755b39c65fdc2c57163abab

Observation b05a3079-dbfe-451a-822f-2fdf7a2807fd · outbound

This paper cites Graph Theory, volume 173 of Graduate Texts in Mathematics.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Graph Theory, volume 173 of Graduate Texts in Mathematics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:00.969443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:00.969443Z digest=sha256:f8bb6ee27f85c72d9ba74b342f5ed31a51b7c1a4a28fd884181293e6700fbfc7

Observation 8df2e43a-81ed-4f51-86c1-8d7ca34319c7 · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.068658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.068658Z digest=sha256:0ae389c2ca7753e7816c249f966b3f9a920f1fa49b57ac64796fe33309303434

Observation b086e352-bd30-462a-82bf-04eb6f206894 · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.136138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.136138Z digest=sha256:d882b5d38e3028f391adcc7810c5ba3011a2812bd47ec61aa0e4931dd3d22cef

Observation dd05b09b-89ea-4dc8-aef8-c0a5580faefc · outbound

This paper cites Algorithms.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.210349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.210349Z digest=sha256:6d31238b80d09cbd8bf53b8bc04a24c1565eb9d9b6ef9570cd0e000a9872e06d

Observation 526f11c0-2670-4579-98da-bf6a8271bf3f · outbound

This paper cites Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.280536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.280536Z digest=sha256:2055a8bf41271a1cee573e7cf7af0645010079ac85ab62a87974c82d9550205e

Observation 6e6d869f-eb4c-4f57-9385-cd585674287b · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.374648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.374648Z digest=sha256:9ddda2e9fcd862ad78d9c9ebb782c6ac53009e4cdfa5071ddb534aedaf2e8c8a

Observation d4d0de98-7073-42af-9e63-3e2584bff874 · outbound

This paper cites L., Knuth, D.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs L., Knuth, D

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.432635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.432635Z digest=sha256:c3a2d3468403f04837b176b035330c04fe61c6612c47f1d37e60e37dd47883b8

Observation adb61b36-c9fc-4952-a7b5-d5f6a1c73651 · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.526953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.526953Z digest=sha256:58df06243fe6f8945c058b6a3836fd49a2e4f75898d5c9ced5b66c0280b7558d

Observation bb9a1ae6-b37b-49df-a2e1-c7996c32c7a5 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Measuring mathematical problem solving with the MATH dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.600073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.600073Z digest=sha256:ae21a3d6ac7b9bdbb17f8b02a20264c3efd7011074175479d189ff7d2d57e743

Observation 896db4d2-7183-4b37-a7b4-ee5b1200ce00 · outbound

This paper cites A rosetta stone for AI benchmarks, 2025.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs A rosetta stone for AI benchmarks, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.646630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.646630Z digest=sha256:58e20896c963209e56d22a419c3634e6a60810df1fea2f46b89758a1b25579a9

Observation 6228f483-e8f7-43f5-8fc1-003298e782ba · outbound

This paper cites and Rosen, M.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs and Rosen, M

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.712711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.712711Z digest=sha256:8588f77b7d1a0ad6b8f43b116439b365bb354489baf3f62011e2bff1089cc7c4

Observation 878f1a1f-ee9a-4f1f-a2cd-ca9b894db637 · outbound

This paper cites Z., Sahai, S., and Leong, B.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Z., Sahai, S., and Leong, B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.801773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.801773Z digest=sha256:2653c54b1ad59310162d74129fd37ac797aea9a56b428ce5e75390b7cb7ee5b9

Observation fd6f9f29-9663-40f1-abd2-5f4527a2ecb9 · outbound

This paper cites Probability Theory: A Comprehensive Course.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Probability Theory: A Comprehensive Course

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.862136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.862136Z digest=sha256:7aec0cfd1c672cfcd4179ebdc559952e94fb38b3031c569f8740917119a73524

Observation 17c9f268-3cb3-46e8-9924-38777a67f496 · outbound

This paper cites No free labels: Limitations of LLM -as-a-judge without human grounding.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs No free labels: Limitations of LLM -as-a-judge without human grounding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.930387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.930387Z digest=sha256:71031bf100d4a5607b2887ab0ce25d655a1489767d53add1bd7b8b4f767b8eff

Observation 586f2b6c-f143-4e00-af12-5de1aae16d3d · outbound

This paper cites Grading scale impact on LLM -as-a-judge: Human- LLM alignment is highest on 0-5 grading scale.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Grading scale impact on LLM -as-a-judge: Human- LLM alignment is highest on 0-5 grading scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.025323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.025323Z digest=sha256:f4045fdccb57161d8e922bde31f8544cde6a0943da75fa9a04d7f85373452cef

Observation c0574767-f252-41f1-b5e5-dae9cccdeb21 · outbound

This paper cites S., Zhang, H., Zhuang, V., Zaharia, M., and Min, S.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs S., Zhang, H., Zhuang, V., Zaharia, M., and Min, S

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.092659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.092659Z digest=sha256:7035003db9b388f13217ffb20a8df29779fe0a9a9d60b22b6a4b8999d32ec801

Observation 2de0f02c-fc26-472a-9da4-f3ae07e2778d · outbound

This paper cites Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.170342Z digest=sha256:d5caf023c526daa02721f30426ae73842dccabf865c2837baa4b81096175238c

Observation 32010958-c6f4-4d70-9020-d7a2fee0095f · outbound

This paper cites RefGrader : Automated grading of mathematical competition proofs using agentic workflows.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs RefGrader : Automated grading of mathematical competition proofs using agentic workflows

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.267945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.267945Z digest=sha256:1751b4472aad28b8fe00ef7a8878c59397aa4ef80e8f36dd8a16b82448733a6e

Observation 5e484d45-2bac-4484-972d-fdcf928d5721 · outbound

This paper cites Scaling generative verifiers for natural language mathematical proof verification and selection.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Scaling generative verifiers for natural language mathematical proof verification and selection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.351440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.351440Z digest=sha256:301e2d878e5abc007907e5be4005861597c22fd9fc5d8169159a1c1e99901ce7

Observation ac8d2d33-3ecd-4f73-8cfb-37ddf7485811 · outbound

This paper cites and Ne s et r il, J.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs and Ne s et r il, J

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.458595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.458595Z digest=sha256:8d0c512c872fda385db7abfc354c9a3daf90d4ef0850325a9beddf2eb0bfab25

Observation c7f4a3a1-f9a2-46ee-b4ea-88292f2d1f25 · outbound

This paper cites and Raghavan, P.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs and Raghavan, P

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.528262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.528262Z digest=sha256:6a0d62a36fdf352b504828f82ec17baeeecb14cb72f6c961abf53d1d25fb32a5

Observation 08b015f0-5bdf-4e99-94ba-0b9b8f2794f6 · outbound

This paper cites S., and Montgomery, H.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs S., and Montgomery, H

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.597956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.597956Z digest=sha256:d0200b0fc3c4a384b96d5ce855a15d0aec90979f4b8bc6212f14416c628fcce4

Observation 0035540d-fb83-4987-bc0c-cb155b3a6f18 · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.701570Z digest=sha256:acd02beee68fe6a9fb1c9e8a1a217fe8b4a5c141de4e0adad25bf8dedc396d29

Observation f35dc2e2-e4ea-4f4b-8f33-db5468070c85 · outbound

This paper cites BrokenMath : A benchmark for sycophancy in theorem proving with LLMs.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs BrokenMath : A benchmark for sycophancy in theorem proving with LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.769988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.769988Z digest=sha256:ffe8da2b7ae5dd05e9fc59503521c8a4a1e9c37420197fbbbc006b85045a98b2

Observation c4d4fa90-c084-4bdb-9bf1-af126d323a7d · outbound

This paper cites ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.833491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.833491Z digest=sha256:0deda5376cf5b2e096822a3dff7d34b38241c4e6edd22a7317654bba522f8235

Observation b60c8a8d-2af9-4dc5-8ddc-dcdc603f59eb · outbound

This paper cites Yale-QEDBench : Code and dataset for quantifying the alignment gap in automated evaluation of university-level mathematical proofs.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Yale-QEDBench : Code and dataset for quantifying the alignment gap in automated evaluation of university-level mathematical proofs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.905609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.905609Z digest=sha256:3aab0c5175a5c5e47d873b14d1e32d6fcb392f63ae3e85d4a781ed325801f0fe

Observation d5ed46a8-b710-4928-b0ae-f517b72e0498 · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.979503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.979503Z digest=sha256:8430ac1ecef3492e9b1cdb0630fb1a8b501bb31c75e43ca2e391294ae345781f

Observation 1f8abaf7-22a3-46b4-8828-108f2c723af3 · outbound

This paper cites IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.059645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.059645Z digest=sha256:7c82c9c38abfe5514054fd4540b4a920e1cc725a92d754218e9848336425f442

Observation 913b81ff-ad3d-4393-9a90-73dccc64e506 · outbound

This paper cites Benchmarking Large Language Models for Math Reasoning Tasks.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Benchmarking Large Language Models for Math Reasoning Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.146717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.146717Z digest=sha256:7aa2f59c7f2fa3d9efef500c34c2c23e79abf5d3f8785b48bdf338f5719de524

Observation fb9209c6-3750-4404-a047-bd86abebee1e · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.245527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.245527Z digest=sha256:05d828e2c5a875a9bea6af443e5110939b9c58f259d675d5a2b64265b526495a

Observation 7924858e-51e8-4dc8-809b-76288cd38f51 · outbound

This paper cites A Computational Introduction to Number Theory and Algebra.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs A Computational Introduction to Number Theory and Algebra

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.340532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.340532Z digest=sha256:89c81819c671d76480e7012076d09cc97df5590bc44fffa3d9e9c1368af69ce7

Observation b5eae475-91bd-4692-a2ef-b93c6307278c · outbound

This paper cites an unresolved cited work.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.418451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.418451Z digest=sha256:67c6e4ce918161adcb48d52cb7ee952f355b1253b393697206cdc8901a6ceb14

Observation fd18ce12-afc2-48fc-9bab-6da92ae147aa · outbound

This paper cites From calculation to adjudication: Examining LLM judges on mathematical reasoning tasks.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs From calculation to adjudication: Examining LLM judges on mathematical reasoning tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.478809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.478809Z digest=sha256:5aa6d415b2e41b76550331af48c489f42cd6770502248e6307f2a466865f4781

Observation a7b8f2ba-a924-4d1f-a9c4-594111c1ffc3 · outbound

This paper cites Ordinary Differential Equations and Dynamical Systems, volume 140 of Graduate Studies in Mathematics.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Ordinary Differential Equations and Dynamical Systems, volume 140 of Graduate Studies in Mathematics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.593708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.593708Z digest=sha256:49b00e1fbc7fc6002e7638549e338c45f5aca4f051e58757fe005e39094106db

Observation 97704c96-41b7-4f0b-9b25-644efea1b5dc · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.669960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.669960Z digest=sha256:ec4115369e5c180f29540e9b1cb9e7d7ff662171c039562f37b38761b922aec2

Observation 7946e161-cc25-4a5b-885b-7f6d809d0f0d · outbound

This paper cites D., Leng, C., and Liu, F.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs D., Leng, C., and Liu, F

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.731342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.731342Z digest=sha256:bfdc75c19ec7c779ebd04820d321d5e4fe76a994d42675a02a275bddd005503c

Observation 0792537d-098d-437d-8ae2-b3128e394508 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.784064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.784064Z digest=sha256:6aace6823772fb751c92c69ae936540dcac3e558842b0626e9d29d28859a7308

Observation 0ac07f8c-b302-4334-9815-8be635115515 · outbound

This paper cites One Token to Fool LLM-as-a-Judge.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs One Token to Fool LLM-as-a-Judge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.858801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.858801Z digest=sha256:8e979bb634292198f946c2d6bc9aac0ad859730b241558e5fdd38e1c66b9aea2

Observation 7870a2a1-3c16-4b57-8bb4-c1f57e856d61 · outbound

This paper cites P., Zhang, H., Gonzalez, J.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs P., Zhang, H., Gonzalez, J

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.931054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.931054Z digest=sha256:a8c7ee8fabcff532ccd3f7679df68f8cbf60f1407a82c2fa0c3c1db183a86a89

Observation bdb3bfc1-bea4-4fd5-80f9-6a8738c50063 · outbound

This paper cites Unmasking reasoning processes: A process-aware benchmark for evaluating structural mathematical reasoning in LLMs.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Unmasking reasoning processes: A process-aware benchmark for evaluating structural mathematical reasoning in LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:04.034797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:04.034797Z digest=sha256:553413473fcf4ca54f01ac1b305572ee50011965c68526c3ed00671666c0faf4

Pith citing papers

Observation 492f8f4c-38d1-4d7c-a3ff-bb7b2e68b861 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:53.720593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:46:50.177357Z digest=sha256:a34a5b0af6a4ee068e7d80daa62d7410ab69d68b02e178bcbb91b55308ea78cd

Observation 680b7fbe-ccce-4dd1-8ee6-4b057db4ffa4 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:53.720593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T22:38:26.111517Z digest=sha256:e49fe61e4d912562291a728073665af0ee121975bbf9ad92998f07edafbbb026

Observation 18da3e37-f0bd-495f-b826-7288dbc0102c · inbound

A Scalable Approach to Evaluating Moral Sensitivity in LLMs cites this paper.

A Scalable Approach to Evaluating Moral Sensitivity in LLMs QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T05:44:33.099337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:44:33.099337Z digest=sha256:35661c6d05ca84884e44c3c3cfcb0ba15864793931955146d4d7b90fc58ebfbe