Pith. sign in

Paper Citation Record · LEDGER

Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2503.21934.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21934 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:46.479498Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7186d4e3-47cb-4a87-8547-e1c4c2413e5e · inbound

Sustainability via LLM Right-sizing cites this paper.

Sustainability via LLM Right-sizing Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:25:03.677934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:24:58.690462Z digest=sha256:d4dcb3a839c653b321bf3164043107b9c2d7e4b043f82a7155048f9f56066d0c

Observation 2ef8e4b6-d6f6-4472-a70a-31803eb8d840 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.739163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:601bffeb6cdba314781f6afdddaeffbb4866a7699569ed671781021f04abaf90

Observation 4ca2fab4-917f-4e1c-ae7a-11af97afdb54 · inbound

MathArena: Evaluating LLMs on Uncontaminated Math Competitions cites this paper.

MathArena: Evaluating LLMs on Uncontaminated Math Competitions Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:10:14.866036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:10:14.812539Z digest=sha256:e7cd584245de0cf4733cb9aa76121e0974848e031f37536a59c70d3c8452638e

Observation dcbe21ce-9286-46e4-961d-606652a4cf72 · inbound

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems cites this paper.

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.479498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.479498Z digest=sha256:28cac71bc4b8e2658d329c75e3e5bfd5bda80ab87ec0e8153242429af026bed4

Observation b12c2113-dbc3-4f3b-bc7e-d081fa75d0cf · inbound

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning cites this paper.

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:49.653031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:49.653031Z digest=sha256:517812d8e29c326612ae4ace41002406ad114b637d978108cc09d3e9c13921de

Observation 8987591d-b5bf-43a3-a7f3-650d244d6e3a · inbound

Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models cites this paper.

Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:10.154073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:10.154073Z digest=sha256:7fddcb5fe8efaea6fc3976f7a627fc3c4fc178c3a671193da98db08ca6b69b1c

Observation 4447a8c4-0263-42a2-a4c4-cadde5a67f21 · inbound

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning cites this paper.

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:25.107443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:25.107443Z digest=sha256:71342ed7ecd739148ba75ae3191bb4e78d56e90242d3ff88d357d654dcf236bc

Observation 571dd12d-aa7e-403e-bd5d-1c666edd50f7 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.635014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.635014Z digest=sha256:fa62b0f3cfbfb143265f3b168dc7294978dda7a97707e7817fdbadeb53fd5ba1

Observation 0e0d4d6e-709f-4c24-9038-60039d47a15a · inbound

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving cites this paper.

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:30:44.646893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:30:44.646893Z digest=sha256:aaab919ba109f2d65d634890f7740a74fa5d4b5108ea157c46f37c8ae313440c

Observation bbf8efaf-4312-4d5e-984a-68622d4d12cc · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.057856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.057856Z digest=sha256:4acbc6a636be36925cf7fe3f385d4e6c77ded91292a744e62c0440ddce4139d2

Observation 5ca9addc-564d-4e46-9b2c-6d37d151fd36 · inbound

The Mathematician's Assistant: Integrating AI into Research Practice cites this paper.

The Mathematician's Assistant: Integrating AI into Research Practice Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:07.713740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:07.713740Z digest=sha256:3be5ddcf1cf9cc9ef2e06e66c7fc0d076166513b0d6deed8af6dbfb681220bf6

Observation 715d8491-2da2-4694-8254-8e41e9adfdcd · inbound

Can LLMs Generate and Solve Linguistic Olympiad Puzzles? cites this paper.

Can LLMs Generate and Solve Linguistic Olympiad Puzzles? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:44:24.438380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T22:41:03.883239Z digest=sha256:49f7ea02ea5f39393f0a506f284c52dfc95f1ffe9bedb2576104cc3c9ae2df6d

Observation 78430be8-a216-4003-9f15-1a71ccf2654b · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:46.119576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:46.119576Z digest=sha256:d87c86325774167a72d8ab01316afa444fc56473109e1fdca69f0f83eedd3e9f

Observation d590de59-60f5-4031-b38e-81c170d6b9f8 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:55.662016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:55.662016Z digest=sha256:95bce36d7c2773b8ae6d6eb583a53ad76f1e39afec9420e2456227398d059f74

Observation 162e4100-9614-4493-bb65-7b9fe7071677 · inbound

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? cites this paper.

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:35:36.417510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:34:03.866227Z digest=sha256:30f226ea2396a13dec98e38029ab4d6ccaf3f8dc5e6904334c986482b09e5391

Observation 1f532c31-8a77-484c-b75e-edff271502d6 · inbound

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? cites this paper.

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:14:48.500366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:14:48.500366Z digest=sha256:4d95697fe3be5f4b3c59f9c6edcf81906ac3872a2431d8d0d9b533b3f98e5a92

Observation eda4132e-139b-44ba-80ff-6ec655e83baa · inbound

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models cites this paper.

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:50:25.760811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:47:21.466578Z digest=sha256:3e9f54b251b28b87b2c950072a71ce67f22fb1df9d44a535144687eccbb37d2b

Observation 0035540d-fb83-4987-bc0c-cb155b3a6f18 · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.701570Z digest=sha256:b2fe400c728d769adebd52364d40988408fb1e12880df49ca84fd26626d0c283

Observation e0c75cff-7f64-4eac-920f-2f3dfe992c03 · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.895410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:8c27e552e312c8b541b2c90ee9c64d73a86760030e57d2d44d28b8182fc1c824

Observation 1129824b-89c7-44b2-82c9-e4c788d514e5 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:43.452726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:8c7b66fa1645d9ea21b50962d7073ba93e6b48cbfa452a1b64a9d998f00bcf03

Observation e45c6353-82eb-4015-b797-cac984145015 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:8e7f5893efbf323c7a73dcc8de7db0e6e50897c835ed46a549e0c4bcfa54e9a2

Observation c45d02b1-4c41-4f6c-85c3-de7936e8e9e8 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:56:24.283023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:46:50.177357Z digest=sha256:d97bc58c8c659cc82a464f2f5d0bb421e049be4681a1219fba859069a33d8b19

Observation 73d9f5fc-d150-4c4f-9873-3d0072549cdd · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:06.649026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:38:26.111517Z digest=sha256:9f3e122443c1812e94ca6f8e42f4f85541132a4741e4d38aa2ab1b7cd0004659

Observation 16bf2f75-0898-41af-8055-9dfe13d1d263 · inbound

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision cites this paper.

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T02:58:00.106348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T02:56:19.104905Z digest=sha256:46b038c0a5ef7f06d02676261363a1d4e33a868fd86eb0d4e241a97c9ac173b6

Observation 59f5b090-c75d-4a59-ac2b-895f994e0604 · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.943151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:dbd371935a62e8e3c1d25ac8288a661a5f8bc30a1b118b933bb0908f20e5f017

Observation 39d0ecde-988c-48a8-be4b-60096152efd9 · inbound

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs cites this paper.

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:06.574894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:58:06.574894Z digest=sha256:bd53b6f87477fac013b017a2256b4f9b26d7f60c2aea0b9548f1f1772e5e04b9

Observation bf24a7a1-9fe9-4d55-b960-91b809835627 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:25.854976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:25.854976Z digest=sha256:3660cc34f09680766f3bf593eb7414531fe13669c25776efaf61fb5223625fa1