Pith. sign in

Paper Citation Record · LEDGER

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.10903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10903 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:18:41.054671Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T09:29:45.021656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:30:48.177781Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a816238d-e1ab-4c7f-adf7-5d6b99775fc0 · outbound

This paper cites Draft, sketch, and prove: Guiding formal theorem provers with informal proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Draft, sketch, and prove: Guiding formal theorem provers with informal proofs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.195737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.195737Z digest=sha256:d414bfca40fdf2990f4ba29c6d2c588558b1b4d1b504a2687736a7a0ede5ced7

Observation 137b8a50-ca7c-482a-951d-c2fb60725cc0 · outbound

This paper cites Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.397037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:32.275985Z digest=sha256:532c12893abfa35c9e73e3c932783bce148164a149f9da16a6f8ea9f97af723f

Observation 91dfdb80-5796-4738-addf-3a8f02fb9094 · outbound

This paper cites Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.420330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.420330Z digest=sha256:80c7b0b97a006108778d8a529c913809cde9bc032bdeaca570cc06a194f8a402

Observation d272895d-4865-4805-93a2-6e13de144499 · outbound

This paper cites Improver: Agent-based auto- mated proof optimization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improver: Agent-based auto- mated proof optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.191854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:32.532344Z digest=sha256:761f3f138c0723d999fccb3a13bc52e8bf7463eb8d0506006c4a92fccd752176

Observation 75d131e4-b583-4ee6-a264-8d4735dd8f7a · outbound

This paper cites Learning formal mathematics from intrinsic motivation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Learning formal mathematics from intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.950599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:32.672032Z digest=sha256:23e2cba249fdfc59041e42375c6c3a83620f51ad0a83bd9a183fe0c4c1118a1b

Observation 35677c87-0245-4d30-b063-71530db8b180 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.811609Z digest=sha256:39a86556086a779703b4c9a1fe4d4f3207de67a74ef8057894ecc6abde34460a

Observation ed735183-5cc2-4980-b111-503812ddb3c8 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formal Mathematical Reasoning: A New Frontier in AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.934134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.934134Z digest=sha256:452cd6a78434602953c40c5cffb4e34b056fce96b9f358d1c73ad743515627ab

Observation f4062c19-fdea-48f6-8940-82595d81160d · outbound

This paper cites Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.046527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.046527Z digest=sha256:759963c04b8d053b8a5dd092b95884c688feacdc765dff6cc45f3fcdca795822

Observation 18ad6655-73c1-46cd-81ce-1d69c186e9d8 · outbound

This paper cites Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.164701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.164701Z digest=sha256:ae311b945a38e96cc7ff975d0b53c609e8501575ff88a5dfc7ada5d5dc202091

Observation dc537e8e-e477-42a3-9675-eb46c219e13b · outbound

This paper cites Dennis, and Andre Freitas.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Dennis, and Andre Freitas

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.647213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:33.300555Z digest=sha256:e00442e5404ed96d3a58d0b13198a231b10194cfa03b0ba4c6dd432f5d1918ff

Observation 102261bd-0257-4fae-a262-88dc703e0259 · outbound

This paper cites Autoformalization with large language models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalization with large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.409341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:33.622889Z digest=sha256:4c54e810ec04948cd6d70613e39475edef354e94822ef60379d3ac195a16e231

Observation adf526b7-37f6-45d8-bf35-408ab94636e4 · outbound

This paper cites Consistent autoformalization for constructing mathematical libraries.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Consistent autoformalization for constructing mathematical libraries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.157181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:33.790467Z digest=sha256:8d138a84be0a0fb961210e9664e06ec0aa2185aa0bee45d0be2c527d89f06fa4

Observation 8e6ae0de-335e-4f1c-819a-132ccf2cc1fb · outbound

This paper cites Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.907155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:33.924912Z digest=sha256:95c799077266338afe3ab3248955d778d5754a4d77247d30577250c89bc36733

Observation 001367c0-3118-4ae0-b04d-ac25c0afdc2b · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:48.600638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:34.098846Z digest=sha256:48032caea54484c3cbbee4bb32bba78d649e0375748121f912190a4cbc537766

Observation 63f2c65f-bb2f-4d83-afa7-582da3a4167a · outbound

This paper cites The lean theorem prover (system description).

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning The lean theorem prover (system description)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.397535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:34.249390Z digest=sha256:5ebc98d8910ec1ae109d397352638b71c8b39684a465ad0b22ac67ed53699488

Observation 1571152a-433d-4204-9f1e-cd191b04572b · outbound

This paper cites Gonzalez, and Ion Stoica.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Gonzalez, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.145267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:34.432097Z digest=sha256:ca866025a055c2ccb96122dd3ceb06b9d739398aa223dcb2295350f8ecf2c4c5

Observation 74021821-a3e1-4a4a-bbd9-23a8a17e825f · outbound

This paper cites minif2f: a cross-system benchmark for formal olympiad-level mathematics.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning minif2f: a cross-system benchmark for formal olympiad-level mathematics

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:34.562778Z digest=sha256:0f68744345bc1abeb7b0c5335bce99afcc7baeff7b4f0b18b7ea0e91d767788c

Observation 47aac0b0-bdfe-42a6-9f1b-9f7864f60632 · outbound

This paper cites Ayers, Dragomir Radev, and Jeremy Avigad.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ayers, Dragomir Radev, and Jeremy Avigad

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:34.919040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:34.919040Z digest=sha256:919858f5fa03b3a5dd4555a3e4c2640f095343163cb830095e38d7bd7429a6cd

Observation 37032d3a-37ac-4c07-9d44-6cc6ad0ac01f · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.048175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.048175Z digest=sha256:b1da87c69d84f9c3915eeb4d3934daa8b43197c1094d6e7529d023b3860bba13

Observation ae2b1546-3a2c-4ffd-9afa-c804c94c7b7d · outbound

This paper cites chrF: character n-gram F-score for automatic MT evaluation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning chrF: character n-gram F-score for automatic MT evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.147923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.147923Z digest=sha256:865c5242cee2a47ebe772320301c5d2af600d6ed974cff7b88607e8286b2d866

Observation a9f4a789-209f-490d-9a2b-81fab12a788e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.273746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.273746Z digest=sha256:f339fc3c6c66f8913f0e11d08deed00ce81bae74466c9547388de041812a39e9

Observation 2dbd116f-a8d9-41e4-aba6-5a76d2fb2468 · outbound

This paper cites GPT-4 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.413820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.413820Z digest=sha256:5d63178cf9546e612e5daff5baa69a5034e86ebd718a233f85a4e83c20dad71f

Observation b52d544c-1183-4060-9c2d-f7634d36decd · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.595478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.595478Z digest=sha256:96eaf3a511489315ecb3379a018dea5848720f711827787bc3bb37131aec172f

Observation 3bea1d13-df8c-486c-bdb5-ece8dbb8bfed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.713877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.713877Z digest=sha256:392ddce0819a41a5ff013cbfd1600ef8dc65aaf597f66517fa20aa602d14ac76

Observation 333f42d6-2093-4662-8ca7-c489eabde370 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.867509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.867509Z digest=sha256:4f346374ea1a44045ceed288cc7817f93c4ea398f70f35d99ff9f080a2dcd7b0

Observation e0644526-4e3a-4db7-b468-0c8f226d2275 · outbound

This paper cites Enhancing ethical explanations of large language models through iterative symbolic refinement.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Enhancing ethical explanations of large language models through iterative symbolic refinement

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.378180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:36.004769Z digest=sha256:ebea6c788f830ecb31ea9863ffb1b7c3e0f0896c0ddfe235f3ab57b1ac786451

Observation ea25568f-b3a0-4d18-b498-36ba5a2aeae9 · outbound

This paper cites Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.811020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:36.305644Z digest=sha256:e0f88823730a5cf31fac348292c55f0cb59a6bbabf400393560cdd9d6e6f1d23

Observation 21afc887-d59f-442c-82a1-ae126ab38535 · outbound

This paper cites Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.491908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.491908Z digest=sha256:d056bba8b5c8aabf5be9e9dd0c399cab2f7fa14ab0f2eef9d568167f60f6aa1f

Observation beb74f61-cf34-4f7f-b3da-6085559f02c7 · outbound

This paper cites URL https://aclanthology.org/2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning URL https://aclanthology.org/2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.098812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:36.188422Z digest=sha256:76547b0bdc94ef8af29048d1336ccf33634560c2869b94a1fa12608d748fb3f1

Observation cd3f2379-5699-4520-b467-8e93aef866b3 · outbound

This paper cites Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.618070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:36.813396Z digest=sha256:5feb13fe4d03658e0145e50102bae09f7b80a0dbcc6db72781ec5e4e9a7a7ef9

Observation c3cdeb5b-e500-453a-b856-ef46c5be2a5b · outbound

This paper cites Process-driven autoformalization in lean 4, 2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Process-driven autoformalization in lean 4, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.352503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:36.948080Z digest=sha256:913e969f4b8903d41640574ccbae209ea214c0ebf7a784d195bed3d8f57d1685

Observation 1d0d0b32-4316-4192-aeff-b148aa2ea300 · outbound

This paper cites LeanDojo: Theorem Proving with Retrieval-Augmented Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LeanDojo: Theorem Proving with Retrieval-Augmented Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.682669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.682669Z digest=sha256:4c0ffe6c508c96da483e370f51ff38dcd31d8f691b05832ecde0c1b987e8ad00

Observation 6ff906c5-70ef-4cef-bbb5-3508baaec954 · outbound

This paper cites Jiang, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Wenda Li, and Mateja Jamnik

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.159363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:37.243294Z digest=sha256:b1523c643a2ba3d2ee26080fec09ee1d49bab764e6b3127b369c32e0a22d01ee

Observation 4d3ae176-89da-43b3-b7e4-412ae99385df · outbound

This paper cites Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.933592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:37.400350Z digest=sha256:0696f82ee4fd12b6531abf2dc216bb5973e7aef9cf446db33e2a65f090055aa7

Observation 0c7db4fb-86ba-4368-8d0f-462518c4710f · outbound

This paper cites Autoformalize mathematical statements by symbolic equivalence and semantic consistency.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalize mathematical statements by symbolic equivalence and semantic consistency

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.081378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.081378Z digest=sha256:3773820d464401119fbd472a052cc8f57fd60f7f45db89d4e5c5a3e9b81c417d

Observation 1eaac6e5-82b5-4339-91d3-8e3d656f47f4 · outbound

This paper cites FormalAlign: Automated Alignment Evaluation for Autoformalization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalAlign: Automated Alignment Evaluation for Autoformalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.768435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.768435Z digest=sha256:ccaf22d48719f91b1901d5c197cd8bb9c37eac0bbc1d0b0793a941b579739f27

Observation 6ad1e9c1-7d6d-4634-a860-c8a043088537 · outbound

This paper cites Branch-solve-merge improves large language model evaluation and generation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Branch-solve-merge improves large language model evaluation and generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.937243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.937243Z digest=sha256:b569a7cdba1f65f8c42e69766a4c7b7768fab8732e7e328e65a702aa9aea3a94

Observation f6fc7ac3-3d55-4304-b3b2-e556f1ec9ee3 · outbound

This paper cites Self-Taught Evaluators.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Self-Taught Evaluators

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.084976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.084976Z digest=sha256:f65c36d6ff526f2b5c27a96c6f4346f524e497d8b8072ca1f818073f4eb8bd19

Observation 60ff8893-e2a2-41f5-ae70-28be40973b6a · outbound

This paper cites Improving autoformaliza- tion using type checking, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improving autoformaliza- tion using type checking, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.664687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.664687Z digest=sha256:4e05e6c72c0f661aa4b98eab18f8a8c443fee571b48738b406b668a919cf0919

Observation 1d1f828f-2576-4cfa-bf83-6ca7805c7c51 · outbound

This paper cites J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.375950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.375950Z digest=sha256:3e6c5d8a169d08254080be9ceac470e06d6c29d5f465ecda2bc869dd0af8918d

Observation 15a2d426-3228-482a-9c22-0ce3cae0da50 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.572175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.572175Z digest=sha256:775bcaf638a6e24e212db1c4357f6a1e753fb552982eab5e3c2efca06cc71434

Observation dd5a74ee-294e-4314-a397-44d9d4c9fc11 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning A Survey on LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.698535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.698535Z digest=sha256:f44f6a69745bf7fafc891b62f39ac1ca8fd692dfb1dd1ad39472b30227f6dc80

Observation 24134fc3-59a4-4513-8ce4-0d407d1b94e2 · outbound

This paper cites Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.239873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.239873Z digest=sha256:d857d87a32c766a59a8a9c8bd4f97e5e26af4739afcbc243379ec421d1e90cd3

Observation 433988b7-2bc8-4b3d-8835-dbc4e74d7e36 · outbound

This paper cites DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.132366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.132366Z digest=sha256:18d1e81b52cc0ac02c329fb3bf1eae629f418cc9b88e30e26feff00fb2c4704c

Observation e654143c-fe28-4cd0-84c9-265a5007d2e5 · outbound

This paper cites Assessing judging bias in large reasoning models: An empirical study,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing judging bias in large reasoning models: An empirical study,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.671838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:38.874413Z digest=sha256:eb3c6fcad2fb69061e0c4df5f604a52d26cde44652e2223ece99f078b5e730be

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · outbound

This paper cites Assessing Judging Bias in Large Reasoning Models: An Empirical Study.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:9637e186ee41bea3ac58bd23a3a63d3ee2e73f8e3a5df3faa22170e34c212d51

Observation 4c1bd57e-0c22-4e90-a74f-5c29b3d838a0 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.443347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.261252Z digest=sha256:dfd501b362a9f9f3917d43c8fbb9ef7823c9d948779ec703970250e1bc82b41a

Observation 9a85e2f3-ae8e-401a-ba33-8a3d30b96e77 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.103398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.461393Z digest=sha256:fa827d89ea10aa744cb2dd7e4f96da7c1e992dc0c23b0c9cf1867e72222c2777

Observation b7e0f182-6e81-4afe-bbf3-58e4cf7353f8 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.912106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.614636Z digest=sha256:e3b65efa94eb1703b4210f506bc1aa77a29aefd3670c6d9c8ee8c0cc4da4d9d0

Observation 781f4874-c6f8-4d24-9488-e7af0ab01019 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.668304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.767627Z digest=sha256:c2726d538a72626e4485b8b2fdb169e648734b0d29becab4fc4502ef56131f58

Observation 107593b5-e04c-4598-a95c-eeff9355911f · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.419887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.871380Z digest=sha256:ee4c76c5a0f35ea47da817766e271694c83c40239a9372a9255f65e84cbeaa9a

Observation 17c5e063-8659-4a54-88b3-01b9a2313b6c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:39.985638Z digest=sha256:3ac8c10d3e9e40ece2ec88903cc1b2b5e654cf5e2deadd8619614dc1ef216e9e

Observation 02ace0ff-3be6-4064-90f0-162cfda66110 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.942277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.098574Z digest=sha256:f67df7003a4265d90874998fe9f840898f13cb4c77d76956fdbc37431fd15905

Observation 6d05781f-064a-4e99-a2e0-8036bf1498e3 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.688713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.263334Z digest=sha256:b97bc0d200ea1169194a78f88afd52462e7fbdf931fa3e7ec5cb204e6467bf97

Observation 261fb21b-acc0-4aff-a9c7-ddeb5b01092c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.407700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.380673Z digest=sha256:1b394ada8cbb5aa48fb5f5fe1eca04c3f245203f5072e9ed39e7eaa995232ffa

Observation 8118648f-d720-48bd-80e6-5e44868d4b3e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.093832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.502132Z digest=sha256:4afdf26915797ab049d54fe1b468a090e3226968bcc9300b474840ee99f3842e

Observation f8b51ee0-2a3b-467c-b8e0-6b74b0b14180 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:42.853290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.606262Z digest=sha256:30d9ed00b60cf72ef06aab7f81932c0a9020eab99e69fd2df7efc8a92dcbf9a6

Observation e6df49b0-4b5a-4383-b759-d4e3c3a6c432 · outbound

This paper cites 15 Purpose Content Basic You are an expert in formal language {formal_language}.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning 15 Purpose Content Basic You are an expert in formal language {formal_language}

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.549214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.724866Z digest=sha256:80268b8c58e01453a6333d3b278f5ee5ea5492e90521d1044aaadac5b6c9b8b5

Observation 79f0eca7-498c-42f9-a217-827f750abd28 · outbound

This paper cites True" or.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning True" or

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.266962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:40.886817Z digest=sha256:f84903a1eee4ff1e18b52494d75a397ea58c8d7ef0ce90619fe8f0e96f1d20c8

Observation 5f14c5e3-8597-4858-8fe3-ed2c9e6a390f · outbound

This paper cites real ⇒ real.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning real ⇒ real

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:41.953448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:41.054671Z digest=sha256:d3c09e1ce65ffab4216658bd82a618c3da14488eaa9ad15d69c7bc2dc91daa2d

Observation 49864cd8-4a36-4a36-a88a-5d12268d63a7 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:47.645468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:18:34.757295Z digest=sha256:3f8efca9d3a49c41cc57bb51c40014ec2cf1eabc44f3d0435eec97f13c64c6c9

Observation 315b83f6-3c65-4b4b-84f5-c33ba707b30e · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.172.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning doi: 10.18653/v1/2024.emnlp-main.172

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.473040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.473040Z digest=sha256:4adc813e120eae3f4e6e31388d95be30214f6da6340d67abfce5bbf599412c12

Observation d0d82ccf-bfb5-44be-bf8e-eb045006c155 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.536531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.536531Z digest=sha256:5ddcedcdc8b47283217fe4692e17a78a15c2b799e5e194091fac2b9ee9e957b0

Pith citing papers

Observation dc4654cd-50d5-40c0-9878-6cbf9d206818 · inbound

Monotonic Reference-Free Refinement for Autoformalization cites this paper.

Monotonic Reference-Free Refinement for Autoformalization Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.179885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:29:45.021656Z digest=sha256:acd2eb4a4ecbf4cc77c899d258ac83213288e1113504293e9e51ef57944cb6d5