Pith. sign in

Paper Citation Record · LEDGER

Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2503.21934.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21934 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.067334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5759f327-6af7-4236-a1dd-9b17df1ad56b · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.067334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.067334Z digest=sha256:adf96bcf382e93278b01dbd1ea45640414bc4ff948b58c6542a98cce0bbde20a

Observation 7186d4e3-47cb-4a87-8547-e1c4c2413e5e · inbound

Sustainability via LLM Right-sizing cites this paper.

Sustainability via LLM Right-sizing Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:25:03.677934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T19:24:58.690462Z digest=sha256:8d39878f8b1e8f5516e3cd4b54371c39e66983352fee1d71fb4a6ca8592435a0

Observation 2ef8e4b6-d6f6-4472-a70a-31803eb8d840 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.739163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:9107bb512c79a221f4812616f81eb00c9e6b6fc6e070d926c5a94055c4de1af7

Observation 4ca2fab4-917f-4e1c-ae7a-11af97afdb54 · inbound

MathArena: Evaluating LLMs on Uncontaminated Math Competitions cites this paper.

MathArena: Evaluating LLMs on Uncontaminated Math Competitions Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:10:14.866036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T00:10:14.812539Z digest=sha256:4a0bb91cb0c4dc647b49880e938066471208718771c933c611427daa9578c9e2

Observation dcbe21ce-9286-46e4-961d-606652a4cf72 · inbound

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems cites this paper.

Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.479498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.479498Z digest=sha256:505376e856d91757fdd03b9967dd6f3931187c4cdc494d4f6d4925d76640f3ed

Observation b12c2113-dbc3-4f3b-bc7e-d081fa75d0cf · inbound

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning cites this paper.

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:49.653031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:49.653031Z digest=sha256:ebf9a09692f8fe989bed08861512b61bd829ac98d7d47a0bd47141d473f93f86

Observation 8987591d-b5bf-43a3-a7f3-650d244d6e3a · inbound

Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models cites this paper.

Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:10.154073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:10.154073Z digest=sha256:1de61f68d2a5a8cfc2cf169cf6141a7d9a84141541e85797f74b604bbde0cf15

Observation 4447a8c4-0263-42a2-a4c4-cadde5a67f21 · inbound

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning cites this paper.

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:25.107443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:25.107443Z digest=sha256:34b29d6b5d4830dad936c9d4803f23eb8d23041d95eda74dc6c20f0dece8dd3d

Observation 571dd12d-aa7e-403e-bd5d-1c666edd50f7 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.635014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.635014Z digest=sha256:480f6917f90902de6bc373a731179fc58b9743efba66acccdff4717e6d9214af

Observation 0e0d4d6e-709f-4c24-9038-60039d47a15a · inbound

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving cites this paper.

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:30:44.646893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:30:44.646893Z digest=sha256:85ac4f350216dfa67a15bd1971bc0d46473957a69896c872d4b8076956a12767

Observation bbf8efaf-4312-4d5e-984a-68622d4d12cc · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.057856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.057856Z digest=sha256:5f34d97375604c072f9053876c0d61a80b6a46adedb5548af9aca898b68bd74b

Observation 5ca9addc-564d-4e46-9b2c-6d37d151fd36 · inbound

The Mathematician's Assistant: Integrating AI into Research Practice cites this paper.

The Mathematician's Assistant: Integrating AI into Research Practice Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:07.713740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:07.713740Z digest=sha256:f56358d76a21a9bd08f5c4bebb0084e6b7329d5781c9730ccf71e1b473f998fd

Observation 715d8491-2da2-4694-8254-8e41e9adfdcd · inbound

Can LLMs Generate and Solve Linguistic Olympiad Puzzles? cites this paper.

Can LLMs Generate and Solve Linguistic Olympiad Puzzles? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:44:24.438380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T22:41:03.883239Z digest=sha256:b7fe12b922ee205973ddd212815d2f70a79b45ef294cdfa912560b21542ba28e

Observation 78430be8-a216-4003-9f15-1a71ccf2654b · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:46.119576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:46.119576Z digest=sha256:dac75c2caeccc9833e9e91ab256b34c139f16f06d873302ddcf39a25cc278dbf

Observation d590de59-60f5-4031-b38e-81c170d6b9f8 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:55.662016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:55.662016Z digest=sha256:c8dcefbd1281cd87c6d67ef1653740a8cae2a0a82b51c091bcf70659f8dd21d4

Observation 162e4100-9614-4493-bb65-7b9fe7071677 · inbound

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? cites this paper.

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:35:36.417510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T01:34:03.866227Z digest=sha256:e63e0beda566d6b4cd24f7445940636428a255ac9a2d6e21bcb9ff33bef768f5

Observation 1f532c31-8a77-484c-b75e-edff271502d6 · inbound

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? cites this paper.

Large Lemma Miners: Can LLMs do Induction Proofs for Hardware? Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:14:48.500366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:14:48.500366Z digest=sha256:fde1c17ecfc4d571819642bfc1dc58daecdac7354be9344c88fd83385f360f90

Observation eda4132e-139b-44ba-80ff-6ec655e83baa · inbound

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models cites this paper.

From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:50:25.760811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T22:47:21.466578Z digest=sha256:d7790801dde68fa02013829e15b9e03c7d6fdb88ef674f474afc5afb47b7e5d9

Observation 0035540d-fb83-4987-bc0c-cb155b3a6f18 · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:02.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:02.701570Z digest=sha256:7913983cbe714d08392bf206aabea5247e189d2ba3331aa7d92b0b757811fd10

Observation e0c75cff-7f64-4eac-920f-2f3dfe992c03 · inbound

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence cites this paper.

BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.895410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T19:48:13.133733Z digest=sha256:08b4fb18051cb7d5f8148b34d2d5a2732da741f3ff9d492c93b917fb1cfa4b79

Observation 1129824b-89c7-44b2-82c9-e4c788d514e5 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:43.452726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:e3c8b21c3e660b8d617c0f47b05a2b5a20a012506e7209a8935d0376d4e2b494

Observation e45c6353-82eb-4015-b797-cac984145015 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:458f34631172811bba4497cb1af93f3d973597a35f44564c4bc5d8c09f806b72

Observation c45d02b1-4c41-4f6c-85c3-de7936e8e9e8 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:56:24.283023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:46:50.177357Z digest=sha256:54b04e28d83bd01106304b31c23e6786e5b32d8ca3763376bd5d35738cd4f25a

Observation 73d9f5fc-d150-4c4f-9873-3d0072549cdd · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:06.649026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T22:38:26.111517Z digest=sha256:f5a5d0d294c7c59c11a11dd258b267e145d342a47304c44d9ebdb812596164ca

Observation 16bf2f75-0898-41af-8055-9dfe13d1d263 · inbound

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision cites this paper.

STAR-P\'olyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T02:58:00.106348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T02:56:19.104905Z digest=sha256:4d01e5a7562578e2cdad38f5d58a27ed278ef325c99c819e49bfb5dc70fe7927

Observation 59f5b090-c75d-4a59-ac2b-895f994e0604 · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.943151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:5c6ea787f7e9adf77ffb6342aae3c01cca5c87d199973d959da3160f9d985dff

Observation 39d0ecde-988c-48a8-be4b-60096152efd9 · inbound

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs cites this paper.

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:06.574894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:58:06.574894Z digest=sha256:5f1c7a1d75fce93a5ee8deaecdc4e5286efcd1505b112a6f65b65e37f57f8463

Observation bf24a7a1-9fe9-4d55-b960-91b809835627 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:25.854976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:25.854976Z digest=sha256:d20bd885feab89d2cd02fddba82b9eb3e8093a82eef0b56d4460686b3989a4dc