Pith. sign in

Paper Citation Record · LEDGER

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

As of 14 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 8 inbound Pith citation observations for arXiv:2509.26076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.26076 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:37:32.560955Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:18:03.059645Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:47:34.576528Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a0b2507-5364-45d0-b760-f35a7f8848f4 · outbound

This paper cites auto" max tokens=32000 reasoningtokens=31000 GPT-5gpt-5 reasoning effort=.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation auto" max tokens=32000 reasoningtokens=31000 GPT-5gpt-5 reasoning effort=

Reference 3

Resolution
malformed identifier
no resolver link, observed 2026-08-04T13:37:32.491188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.491188Z digest=sha256:b6eae3a64e046570e254a4665da211857afcae030882db6b035dcf0811517943

Observation 859ddf17-b539-4bd9-aaf0-85183824d9cb · outbound

This paper cites True”, “False.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation True”, “False

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.486780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.486780Z digest=sha256:b1d3bff92838053dfb787fe8f3820f8f1b76c839e17091afc37d004ef15f6f31

Observation 7c497c22-3734-4a45-91c4-60d6ae89c9cd · outbound

This paper cites Underlying system is ArchLinux with many standard open-source computer algebra systems (like GAP) pre-installed.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Underlying system is ArchLinux with many standard open-source computer algebra systems (like GAP) pre-installed

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.527322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.527322Z digest=sha256:bfba51bf5583d9ce0091a888578961b7ae82cfdfd6535adc324ad289a580a3bf

Observation 1793932a-450d-4f67-8f15-7aec79b04cdf · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.531671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.531671Z digest=sha256:bf1719e51d8f7571de79266c6cc1ef53788b8d209e4dfcf0a40010a4138aea2a

Observation ab2dc781-301c-4951-8797-eb55724f477c · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.536767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.536767Z digest=sha256:927826ce13c7baa72a749cb867d5237b579e241bbef092e946e7229076230534

Observation ed45ed3d-7901-4e1b-a2b2-daf658740686 · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.540729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.540729Z digest=sha256:052353a4cc2d298e88c22f88a8da19b13e888f4cf98dfc7de5a6b7e329bbf041

Observation 0b91531d-a425-4294-a01d-215f96e14b7a · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.544745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.544745Z digest=sha256:ec245c8dacb47a23621904ab79ce5ba99d1796d2d4af29cf6c630de09d28b907

Observation a13de7cd-8341-4d60-9d74-722221a2b506 · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.548634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.548634Z digest=sha256:5c4b08c86a789a7847e8e3428b469df9491c4113889b1f643982d08e850c2a4c

Observation c9b8500a-a5ff-4292-ad51-64bca15661ea · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.553254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.553254Z digest=sha256:c1994322b80d9ba19d7d7e07d77fbe053c71e612a416c7dc01a5739d4e8ea410

Observation e86454a3-6f74-44f0-86a9-23feabeadd4b · outbound

This paper cites an unresolved cited work.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.557052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.557052Z digest=sha256:dd3e9672a0355cb990be956323071e5a9c9eaca7214a5990636d86c394963018

Observation b0bebd92-ee2c-4141-ab8c-a4915d36f924 · outbound

This paper cites Automorphismsˆ2: {G.automorphism_number()ˆ2}.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation Automorphismsˆ2: {G.automorphism_number()ˆ2}

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.560955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.560955Z digest=sha256:18aa7e1a720df84c7194285d126918696d81433d8f1168ae6493102e96d274ab

Observation 0223ef71-1371-4e6e-be0e-a8e68dc2a275 · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T13:37:32.480041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:37:32.480041Z digest=sha256:323111fa7bb7a27104aa7c9d43e9d5ca1e95714471b6565f460879d86b8e0503

Pith citing papers

Observation 1f8abaf7-22a3-46b4-8828-108f2c723af3 · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.059645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.059645Z digest=sha256:ab6f76f08e500106eb35989159b3b1df8f8db0c8235a486a3c3ab63f0f51c737

Observation d7557716-135e-44c1-8d77-199091052b5e · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T02:19:32.479333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:53:12.708706Z digest=sha256:2b63f6711141e38d8a7134867eecd56f9f32d2ddeab5e22e849b16e59c8b6fa3

Observation f01528dd-f169-4bd6-a94c-d75730cf07d8 · inbound

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs cites this paper.

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T02:19:32.479333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T22:24:04.625823Z digest=sha256:9a76a59b4f591b8f8e65897ddeefbf0ffe61d7b36405a7a176af0811c080ca47

Observation ea917522-b376-4351-a2bc-241d004cb161 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-07-10T02:19:32.479333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:46:50.177357Z digest=sha256:28579f628517df0f1873dae35d8dc0da0076bcfe55345e0bb4f3614ba06ff086

Observation 8cad4ab0-97cb-45f3-aafb-bbed3f1ae5b1 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-07-10T02:19:32.479333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:38:26.111517Z digest=sha256:7e9ecfe5a4aafec58cffdd38272ca0aec3e0e82530e1f3e376c5d18b5ae5c38d

Observation b357d459-e2a5-4e25-a717-66f7bdb4f7cf · inbound

Failure Modes of Large Language Models on Research-Level Mathematics: A Taxonomy and an Empirical Characterisation cites this paper.

Failure Modes of Large Language Models on Research-Level Mathematics: A Taxonomy and an Empirical Characterisation IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T02:19:32.479333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T05:09:33.431246Z digest=sha256:d5b52c226e29647e97cdc2299c6dee0b0ac4030eb0a771d6548b8532313b9601

Observation 1f328be5-23e9-4f65-acc4-f99081908c3c · inbound

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics cites this paper.

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T20:47:34.577936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T20:39:42.009312Z digest=sha256:57115196d803aad856b403a5ee7ad0cf328ddb6c8a683216e309f1fdae58b6be

Observation 6b684b02-1c91-4de6-8725-708a62daab8b · inbound

Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory cites this paper.

Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T18:16:49.376545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:16:49.376545Z digest=sha256:49aa61c826ccc70b3831fd05b2f076accf042171f2e28ea55b84b7e9ba5b8130