Pith. sign in

Paper Citation Record · LEDGER

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

As of 19 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.10903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10903 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:18:41.054671Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T09:29:45.021656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:30:48.177781Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a816238d-e1ab-4c7f-adf7-5d6b99775fc0 · outbound

This paper cites Draft, sketch, and prove: Guiding formal theorem provers with informal proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Draft, sketch, and prove: Guiding formal theorem provers with informal proofs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.195737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.195737Z digest=sha256:7330f51c25a54daeefb977636c0e5be85c4b145299becb17f680b40170e14b47

Observation 137b8a50-ca7c-482a-951d-c2fb60725cc0 · outbound

This paper cites Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.397037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:32.275985Z digest=sha256:6e933aa2f5b9681387cc87ae60bd0cdd2091c29c2543d26553746f95ae77689d

Observation 91dfdb80-5796-4738-addf-3a8f02fb9094 · outbound

This paper cites Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.420330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.420330Z digest=sha256:4e60ebb95387672b58ef7a83862af2fc66892ba6158f0cd6f1564b8124aba1b7

Observation d272895d-4865-4805-93a2-6e13de144499 · outbound

This paper cites Improver: Agent-based auto- mated proof optimization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improver: Agent-based auto- mated proof optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.191854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:32.532344Z digest=sha256:5eeeeadb0c609509ea277d1e1dd2e7d5198b06daa24d61a949d2e8c16d2047ae

Observation 75d131e4-b583-4ee6-a264-8d4735dd8f7a · outbound

This paper cites Learning formal mathematics from intrinsic motivation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Learning formal mathematics from intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.950599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:32.672032Z digest=sha256:d1ebb6dbe45c2d090c8a046e541c45c249799c9ea14e3bfe3f236fa6a048bcbb

Observation 35677c87-0245-4d30-b063-71530db8b180 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.811609Z digest=sha256:68920eec84350a5c626a8cd52221744a87c1bcee5c30765bbacdcf5ee5e300ea

Observation ed735183-5cc2-4980-b111-503812ddb3c8 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formal Mathematical Reasoning: A New Frontier in AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.934134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.934134Z digest=sha256:84ec35c9b1b2a58b86a64b2050193b0b14a86ce4350eea5bbee919754cfc8ca4

Observation f4062c19-fdea-48f6-8940-82595d81160d · outbound

This paper cites Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.046527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.046527Z digest=sha256:0bc6010b5c8233f226d1fb568105a76a9691aece81227f6b6f103dc4a6c396a1

Observation 18ad6655-73c1-46cd-81ce-1d69c186e9d8 · outbound

This paper cites Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.164701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.164701Z digest=sha256:fa629d508eee3109de5edfaf69990480441a40fd770b0a619d4f76ebb25fc8c0

Observation dc537e8e-e477-42a3-9675-eb46c219e13b · outbound

This paper cites Dennis, and Andre Freitas.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Dennis, and Andre Freitas

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.647213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:33.300555Z digest=sha256:4d8bfaa8d9b7182befe09800b6679b53239056f48b82f549e36f37d9b957c444

Observation 102261bd-0257-4fae-a262-88dc703e0259 · outbound

This paper cites Autoformalization with large language models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalization with large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.409341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:33.622889Z digest=sha256:5bef138b360a4e64620301d619281c16e18b51bc68180b7014c0f934c555da56

Observation adf526b7-37f6-45d8-bf35-408ab94636e4 · outbound

This paper cites Consistent autoformalization for constructing mathematical libraries.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Consistent autoformalization for constructing mathematical libraries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.157181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:33.790467Z digest=sha256:c9f48536d92455211120f4ead3eca21255c5848661078d42bfa12de21a47cd4c

Observation 8e6ae0de-335e-4f1c-819a-132ccf2cc1fb · outbound

This paper cites Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.907155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:33.924912Z digest=sha256:acd9331ab985df96d741e2944a72c630fa1530f2010fe0a0efbd470ac6a7c0d8

Observation 001367c0-3118-4ae0-b04d-ac25c0afdc2b · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:48.600638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:34.098846Z digest=sha256:3bd764cd967022bd7a51984d4ff02fb4da9031efa0e1263c7f8846dfb5376c42

Observation 63f2c65f-bb2f-4d83-afa7-582da3a4167a · outbound

This paper cites The lean theorem prover (system description).

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning The lean theorem prover (system description)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.397535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:34.249390Z digest=sha256:30fcfe7b24961386c521e3a7bf2824759aceb12ce50ce8564d63fc30de79894b

Observation 1571152a-433d-4204-9f1e-cd191b04572b · outbound

This paper cites Gonzalez, and Ion Stoica.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Gonzalez, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.145267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:34.432097Z digest=sha256:129e9bbcbb4082b21c0a26889687daa71668ed0b71c764c8037bd09df7bc4725

Observation 74021821-a3e1-4a4a-bbd9-23a8a17e825f · outbound

This paper cites minif2f: a cross-system benchmark for formal olympiad-level mathematics.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning minif2f: a cross-system benchmark for formal olympiad-level mathematics

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:34.562778Z digest=sha256:88f6a794354014e3d31f8f2ae4c48c982de957c6b60872259eeb6972d93d0b72

Observation 47aac0b0-bdfe-42a6-9f1b-9f7864f60632 · outbound

This paper cites Ayers, Dragomir Radev, and Jeremy Avigad.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ayers, Dragomir Radev, and Jeremy Avigad

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:34.919040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:34.919040Z digest=sha256:9b585a057cb4040259fab21deebf734d9957b3aa427f8a95834518ceb6562d01

Observation 37032d3a-37ac-4c07-9d44-6cc6ad0ac01f · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.048175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.048175Z digest=sha256:4ea386959440aa784215ab9cf550546d1e52031bfec7417a2a8747c4d9169d25

Observation ae2b1546-3a2c-4ffd-9afa-c804c94c7b7d · outbound

This paper cites chrF: character n-gram F-score for automatic MT evaluation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning chrF: character n-gram F-score for automatic MT evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.147923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.147923Z digest=sha256:4d11e43c7d2a04b4cbd99c611eb0484fd299c99d730c426ae6f03c0c48618ace

Observation a9f4a789-209f-490d-9a2b-81fab12a788e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.273746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.273746Z digest=sha256:40468af434da6871be75484c701550956e7074cab9e37580b25143b11a043570

Observation 2dbd116f-a8d9-41e4-aba6-5a76d2fb2468 · outbound

This paper cites GPT-4 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.413820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.413820Z digest=sha256:2167da07977702533a23f3bcc384a536f609ba33a28cb08f66fa6e98ef0783dd

Observation b52d544c-1183-4060-9c2d-f7634d36decd · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.595478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.595478Z digest=sha256:b2edbf29a3b765614c10aa454fef37ed62dabb57a77fae3f914c68bc9170c10e

Observation 3bea1d13-df8c-486c-bdb5-ece8dbb8bfed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.713877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.713877Z digest=sha256:67111697305f3d35b92019dcb11b876dff5497e64f5c8952c6405eaded145e06

Observation 333f42d6-2093-4662-8ca7-c489eabde370 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.867509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.867509Z digest=sha256:60c6d125476a843bf422f8007d113f99829536628e3a2ace033b2ce92c859035

Observation e0644526-4e3a-4db7-b468-0c8f226d2275 · outbound

This paper cites Enhancing ethical explanations of large language models through iterative symbolic refinement.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Enhancing ethical explanations of large language models through iterative symbolic refinement

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.378180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:36.004769Z digest=sha256:b54afe22a4e7a6a5a18a99b8b77d9e5f9688f2104881c16945a28a612bbc7afc

Observation ea25568f-b3a0-4d18-b498-36ba5a2aeae9 · outbound

This paper cites Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.811020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:36.305644Z digest=sha256:cc630040535e818052d58b2b2695faa3e18d1fb8379e20c2cfc86960995c542a

Observation 21afc887-d59f-442c-82a1-ae126ab38535 · outbound

This paper cites Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.491908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.491908Z digest=sha256:2b6cf67fad5ee5b1a47674087166a9c8bd8072b5b4035a39601ca95b7f8c2bdd

Observation beb74f61-cf34-4f7f-b3da-6085559f02c7 · outbound

This paper cites URL https://aclanthology.org/2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning URL https://aclanthology.org/2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.098812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:36.188422Z digest=sha256:23b06d00d46e4917d132587908f38bea9f9c01bbcbcccdacf6f683f8d3051d56

Observation cd3f2379-5699-4520-b467-8e93aef866b3 · outbound

This paper cites Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.618070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:36.813396Z digest=sha256:0dcc466c4421ac22670aadf968ac4cd9c809b879ccd41e43257076b5ed029f99

Observation c3cdeb5b-e500-453a-b856-ef46c5be2a5b · outbound

This paper cites Process-driven autoformalization in lean 4, 2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Process-driven autoformalization in lean 4, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.352503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:36.948080Z digest=sha256:2ba0b4a38649e4be56adf11a2ce7060c6c298e1767034346f10be67c46faa283

Observation 1d0d0b32-4316-4192-aeff-b148aa2ea300 · outbound

This paper cites LeanDojo: Theorem Proving with Retrieval-Augmented Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LeanDojo: Theorem Proving with Retrieval-Augmented Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.682669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.682669Z digest=sha256:c7c03b12f17457c4db8ff8cfe7fe4a0ba04a3fc3176a2563e78290911e5083eb

Observation 6ff906c5-70ef-4cef-bbb5-3508baaec954 · outbound

This paper cites Jiang, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Wenda Li, and Mateja Jamnik

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.159363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:37.243294Z digest=sha256:24517de2357c12057ed50bf02f3b3b5ea4c30c85c73a0f44b0b2208f231820f0

Observation 4d3ae176-89da-43b3-b7e4-412ae99385df · outbound

This paper cites Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.933592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:37.400350Z digest=sha256:0451263cf3805167777501d4450f34414e638d0dbcd6b7680d889151dd62e6a4

Observation 0c7db4fb-86ba-4368-8d0f-462518c4710f · outbound

This paper cites Autoformalize mathematical statements by symbolic equivalence and semantic consistency.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalize mathematical statements by symbolic equivalence and semantic consistency

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.081378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.081378Z digest=sha256:403280c04f142b62f2340957a0e6d03594f6619ee9cf42ea494ddf6148cd6740

Observation 1eaac6e5-82b5-4339-91d3-8e3d656f47f4 · outbound

This paper cites FormalAlign: Automated Alignment Evaluation for Autoformalization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalAlign: Automated Alignment Evaluation for Autoformalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.768435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.768435Z digest=sha256:558fc2f6407aae63b2358170275228e813bfa3809393c0d9a9c9d3ce76ff58ff

Observation 6ad1e9c1-7d6d-4634-a860-c8a043088537 · outbound

This paper cites Branch-solve-merge improves large language model evaluation and generation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Branch-solve-merge improves large language model evaluation and generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.937243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.937243Z digest=sha256:879156a1b435d3666936e89bc724988ead0f22bb8263a0f1a747d21504504d5c

Observation f6fc7ac3-3d55-4304-b3b2-e556f1ec9ee3 · outbound

This paper cites Self-Taught Evaluators.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Self-Taught Evaluators

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.084976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.084976Z digest=sha256:7daee1a9460cf109254e6bc17f6245209e772452edc697779b9767aa23b93ae8

Observation 60ff8893-e2a2-41f5-ae70-28be40973b6a · outbound

This paper cites Improving autoformaliza- tion using type checking, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improving autoformaliza- tion using type checking, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.664687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.664687Z digest=sha256:2237521f56c6c62d67bc42f972ca7c5ba24c554de3370a82af84bd7493ffe060

Observation 1d1f828f-2576-4cfa-bf83-6ca7805c7c51 · outbound

This paper cites J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.375950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.375950Z digest=sha256:44b9cce9117aaf0d3af85bc19713aa57d7d670b808e1def2ee385c70776bc94a

Observation 15a2d426-3228-482a-9c22-0ce3cae0da50 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.572175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.572175Z digest=sha256:9d45653e8c32d3f7206f268112c2716dfe1a77ee56a53ed7079afe4839a10c87

Observation dd5a74ee-294e-4314-a397-44d9d4c9fc11 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning A Survey on LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.698535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.698535Z digest=sha256:91d73ff216dedac9238ad79d47568b0105ef3b19818c865b6eeee7b4d2bdecbd

Observation 24134fc3-59a4-4513-8ce4-0d407d1b94e2 · outbound

This paper cites Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.239873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.239873Z digest=sha256:487df72d822bacb250c4f6a39705b7e9e08f062f6de4644eb6c63ec30b899b77

Observation 433988b7-2bc8-4b3d-8835-dbc4e74d7e36 · outbound

This paper cites DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.132366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.132366Z digest=sha256:8c522d5b9e3345fe6599927326fb8dda65571709128f35eee782351a95f240b2

Observation e654143c-fe28-4cd0-84c9-265a5007d2e5 · outbound

This paper cites Assessing judging bias in large reasoning models: An empirical study,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing judging bias in large reasoning models: An empirical study,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.671838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:38.874413Z digest=sha256:9c2ee22dfbe3216a5b325f0b8f936e6d50f6f19dcfb30df02a13ad90275c035d

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · outbound

This paper cites Assessing Judging Bias in Large Reasoning Models: An Empirical Study.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:4d5b8a8b0504260d6fd4455258d204e0f41a4d42d56cb9a0a7b7b7ba91031d97

Observation 4c1bd57e-0c22-4e90-a74f-5c29b3d838a0 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.443347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.261252Z digest=sha256:237b8165a3f480e1597aeb764707ec402b1ecb690562402b5b045681ea1a005e

Observation 9a85e2f3-ae8e-401a-ba33-8a3d30b96e77 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.103398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.461393Z digest=sha256:73885c68f2db6edae6787a74aea78267918a563550cc230cb929c7f1a8bcfe12

Observation b7e0f182-6e81-4afe-bbf3-58e4cf7353f8 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.912106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.614636Z digest=sha256:373c7db67799ce621057ea2a505589565e3d5bd972b87af00b777fa8d5518efa

Observation 781f4874-c6f8-4d24-9488-e7af0ab01019 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.668304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.767627Z digest=sha256:f8c39a8b35819b8d46fa61f5301bba5eeff32411089a36edf6e5366e1c476c08

Observation 107593b5-e04c-4598-a95c-eeff9355911f · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.419887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.871380Z digest=sha256:a89e2883c757692ba3995e8276da58320ba7133fd5199546109ddea27127ecd0

Observation 17c5e063-8659-4a54-88b3-01b9a2313b6c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:39.985638Z digest=sha256:225cc0e8ce566b50fbe0305cdf28f89a6e04b1b6610ccfd352cdbd6207af3b6d

Observation 02ace0ff-3be6-4064-90f0-162cfda66110 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.942277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.098574Z digest=sha256:635f900e4ed97b214ea300a95577b43a9f9ae7409f2ac68b764b27888f0a3b7c

Observation 6d05781f-064a-4e99-a2e0-8036bf1498e3 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.688713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.263334Z digest=sha256:423ddc434d43a5090bc3c72b82f18d924a61e5d0efd3262278bc625d92ecadb2

Observation 261fb21b-acc0-4aff-a9c7-ddeb5b01092c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.407700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.380673Z digest=sha256:d119518b2f8b040d6c1bc1561bb272fe4c8e098b62d2c62cbe29329b895e0786

Observation 8118648f-d720-48bd-80e6-5e44868d4b3e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.093832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.502132Z digest=sha256:61ce37ba0346438a4968de6a5be662366ee883c0991f96559248f299d28765b3

Observation f8b51ee0-2a3b-467c-b8e0-6b74b0b14180 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:42.853290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.606262Z digest=sha256:2f06e0487b0291d5fe7c17c837aaff45dc6c1d63a693d8a6ae63c0123130691e

Observation e6df49b0-4b5a-4383-b759-d4e3c3a6c432 · outbound

This paper cites 15 Purpose Content Basic You are an expert in formal language {formal_language}.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning 15 Purpose Content Basic You are an expert in formal language {formal_language}

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.549214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.724866Z digest=sha256:78575e1dbdbf63d89b48d68c0431a2a9a4fa6c4d26eabc9b31a67e5accad5fb9

Observation 79f0eca7-498c-42f9-a217-827f750abd28 · outbound

This paper cites True" or.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning True" or

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.266962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:40.886817Z digest=sha256:498104654817fcf40b573bf2e9dd4f0f05ee2ad405c22ca41ee5a403a12edf96

Observation 5f14c5e3-8597-4858-8fe3-ed2c9e6a390f · outbound

This paper cites real ⇒ real.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning real ⇒ real

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:41.953448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:41.054671Z digest=sha256:ed5c74db889414aa61595d93b12e47269da5390d8c6d5498195cdbdbd9b82a57

Observation 49864cd8-4a36-4a36-a88a-5d12268d63a7 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:47.645468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T04:18:34.757295Z digest=sha256:f7ecc28e9b7262820f6573761a02d207114cdb1f823cee456afad760d787a868

Observation 315b83f6-3c65-4b4b-84f5-c33ba707b30e · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.172.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning doi: 10.18653/v1/2024.emnlp-main.172

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.473040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.473040Z digest=sha256:586fb3eec96b86e6202c050765c19ce29ec46898c10eaf1eee50bee817c9fef1

Observation d0d82ccf-bfb5-44be-bf8e-eb045006c155 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.536531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.536531Z digest=sha256:a33b9c865fca871336f223f1670dcec10cd35bb93a031eaa7248763fcc640c76

Pith citing papers

Observation dc4654cd-50d5-40c0-9878-6cbf9d206818 · inbound

Monotonic Reference-Free Refinement for Autoformalization cites this paper.

Monotonic Reference-Free Refinement for Autoformalization Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.179885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T09:29:45.021656Z digest=sha256:8d0152532293831273d6dcbf34b0884e22a6a49d84a4c76edb2a0abf70ee1879