Pith. sign in

Paper Citation Record · LEDGER

Psychometric-Based Evaluation for Theorem Proving with Large Language Models

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2502.00855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00855 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.728402Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46adc04e-6384-4ba7-8012-b5d099c8852a · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models A Comprehensive Overview of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.570661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.570661Z digest=sha256:a26c1222a61aff7246ad2849282511abd3d0c126a98839bf5188bd90272faab9

Observation d09e761d-77e3-4767-b373-dd821fc42455 · outbound

This paper cites A Survey of Large Language Models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models A Survey of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.575829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.575829Z digest=sha256:2b582667a0e1a650db89ed7e4eb5e684d78eda79b25043545347aad3c591b153

Observation 196b91ae-25b8-4233-9d40-4719144d5489 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Formal Mathematical Reasoning: A New Frontier in AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.579223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.579223Z digest=sha256:98de38a8bbd33bc46c3b4485d831b0e0b32ed704b9be34c79fd53e2369714790

Observation a792a89c-122f-401b-ac54-8d560c2bc5bc · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.582966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.582966Z digest=sha256:dcf5bd8287263efa49ba39e05f0b77cf010f30cd6c44fc8d985f64466a5d59a7

Observation 216c546e-e74a-47a8-ad92-f6e0602a77da · outbound

This paper cites an unresolved cited work.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:36:10.338408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.586909Z digest=sha256:7c24ccacdda2b9107865f6c134479b87367dc10b5717ac5f44dae275a5898066

Observation 724b3397-60c7-492f-8a31-f7ab8e55b027 · outbound

This paper cites LEGO-Prover: Neural Theorem Proving with Growing Libraries.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models LEGO-Prover: Neural Theorem Proving with Growing Libraries

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.590116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.590116Z digest=sha256:f4392a26baff5f7f273bd319c3b67b753a7adc8f80886180e3141d20aa83c24a

Observation aa175261-3913-4a31-b455-920faa242b10 · outbound

This paper cites Evaluating Mathematical Reasoning Beyond Accuracy.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating Mathematical Reasoning Beyond Accuracy

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.593630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.593630Z digest=sha256:194fe8efbdaaf9de1cb492c77417e44335b40842592f8982cf9dbb113985704e

Observation a96324a1-090c-460b-9c6f-6ba5004baa86 · outbound

This paper cites The lean 4 theorem prover and programming language.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models The lean 4 theorem prover and programming language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.597364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.597364Z digest=sha256:8ceca8e9a5a3b789fb5a012d1de5ef151015b5208125b047a9d15ff73074f803

Observation a0b5dd55-cb6d-4b01-8b33-2a281145c171 · outbound

This paper cites Isabelle: A generic theorem prover.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Isabelle: A generic theorem prover

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.600543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.600543Z digest=sha256:1caafab697374d0220486fb84b944c5c72c08201b3fa3f850aa9ecd5fa014aa0

Observation 4719425c-dcbb-4594-b15b-5a0e8466dbe1 · outbound

This paper cites The coq proof assistant a tutorial.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models The coq proof assistant a tutorial

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.603515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.603515Z digest=sha256:17dfcd471a81fd20bb81484905f043f6d514bc6fa02d8ad54a04cda136338c7f

Observation 9c63c7c0-c915-40f5-959a-5cab0960a528 · outbound

This paper cites Leandojo: Theorem proving with retrieval-augmented language models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Leandojo: Theorem proving with retrieval-augmented language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.606594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.606594Z digest=sha256:ce9b5996e80595eec8a2027cdb2254636aaa708d612f2fe3dd9f06bee9ae66e0

Observation 9b0fda7f-b129-4043-a78a-994dc9aa6a37 · outbound

This paper cites Thor: Wielding hammers to integrate language models and automated theorem provers.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Thor: Wielding hammers to integrate language models and automated theorem provers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.609615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.609615Z digest=sha256:a2c794bfd296c2a06b05e2831cacd518c2285f88d07d642a47248daa285c134a

Observation 9dfd4d5c-5d14-4ec0-b538-ab945a6f4055 · outbound

This paper cites Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.612982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.612982Z digest=sha256:c49a10c9705f124a95556a57185b5deed792791a74e3d90a2a9e8f73b436afc7

Observation 33370f08-92ce-4669-a42f-ae96a93c6fee · outbound

This paper cites Baldur: Whole-proof generation and repair with large language models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Baldur: Whole-proof generation and repair with large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.616356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.616356Z digest=sha256:5a247aaa493742feec5267dfdf98ddb126781af2c27065e95ca09537f36ea0df

Observation e0b55614-0113-486b-b881-35b0610bcf86 · outbound

This paper cites TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.619286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.619286Z digest=sha256:431d2851ccc4ce8bb982b1cd3ca6618854d54efe7d893db92166f0f7a4c164c2

Observation 401c5ced-95d8-44dd-a157-a5ad13ebd946 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.622596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.622596Z digest=sha256:ee615468b2d4aed08cc1002cf9ca38db859f71c8aaedc0f1911f482d839ebdbe

Observation 343d4693-79a4-4615-94f8-8ab3181eddb1 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.199928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.625662Z digest=sha256:358c2e0a80c071aad3ed59d9c03c8be49cb6bf445cbb7b363739370c6e4afcbf

Observation 5a350cd1-9f03-44ba-9458-0bbdcf69b8ce · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.628897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.628897Z digest=sha256:7b0867a3fc14e95c0f07177a14e8717f20dbcf6194b47d2cbcf2bb4e0ddf3bf4

Observation 50f8e8cb-e477-46cc-b372-20167b0f48ea · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.632242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.632242Z digest=sha256:adfc15096ce29623b548962e6a5ac722b0b103754eadc27bcdfd41e638412803

Observation 8d72be25-55bf-4215-8434-89bde6f7f85f · outbound

This paper cites Cladder: Assessing causal reasoning in language models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Cladder: Assessing causal reasoning in language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.635508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.635508Z digest=sha256:f1a312ce517bc8faa260cf0fd48a36f9c683bdeb27995843c1d1becc6117f4db

Observation 8f438116-dd13-4aac-bbbf-21fd26c5fcbc · outbound

This paper cites Using llms to facilitate formal verification of rtl.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Using llms to facilitate formal verification of rtl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.184417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.638581Z digest=sha256:31513cf57513d9de20d711063fd93b76accbec14f351a30002678b2465330105

Observation e398a153-d0d9-43ab-8532-405143c50ea3 · outbound

This paper cites Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.641336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.641336Z digest=sha256:3a02ca5a486b5c285a8ac5124334b2e9b523d1a3014343e379eb93188155884a

Observation 4ad8c173-7105-47c4-b0e0-307126fad0a3 · outbound

This paper cites Evaluating llms’ mathematical reasoning in financial document question answering.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating llms’ mathematical reasoning in financial document question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.644561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.644561Z digest=sha256:2b9ed5a55618e4a6e6e7c3c9626a37963814e337529d01004626163bd5b05718

Observation 68778ea7-cdff-46c2-b429-b893d4500153 · outbound

This paper cites Position: AI Evaluation Should Learn from How We Test Humans.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Position: AI Evaluation Should Learn from How We Test Humans

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.647330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.647330Z digest=sha256:963b5185fc9298ee7498a78306dd664e5598510a0fcfc13ea2cd4123a5ef3116

Observation 2673de03-43d8-4b24-994f-5e7c47ea5911 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.650394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.650394Z digest=sha256:4d3173c346d7d493530d8b0412610aa564f21075369cd706b713131250b5e461

Observation 87a1d1c7-2a9d-4572-9231-8c09383047a7 · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.653389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.653389Z digest=sha256:843089f6956ec33173ac1b97af09ef17420bb49b589f5d5da8a9f98c93010928

Observation 78fc73ac-c379-4366-8b99-5625f369895e · outbound

This paper cites Psychometrics: an introduction.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Psychometrics: an introduction

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.168668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.656619Z digest=sha256:7ec5976031e465b834e1b6e336cd28f06dcb6072ba3d5a8f03d8f11f44710cca

Observation 47d6942f-f917-4de4-a709-6d24bcb45c04 · outbound

This paper cites Diagnostic measurement: Theory, methods, and applications.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Diagnostic measurement: Theory, methods, and applications

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.159834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.659369Z digest=sha256:d52ebcc64be25f0d6854efab3139f2a3a8862ed7bf51ef703f10488503180961

Observation 85cbad2c-13ad-41cd-84cb-c09283918815 · outbound

This paper cites A brief introduction to evidence-centered design.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models A brief introduction to evidence-centered design

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.150856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.662336Z digest=sha256:c1689cd5805591dc4a369c7a6c68666e432ae542d0ca2c19000ed920adf9022e

Observation fe1e5f72-840f-423a-ab35-605ab3140e3a · outbound

This paper cites The basics of item response theory.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models The basics of item response theory

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.665270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.665270Z digest=sha256:52ab45f64703672e2f3e8a78a0c86e20fe1b437381186d15d78209c3c67bfea0

Observation ad463b42-6011-4d1c-b99e-f144a7832e5c · outbound

This paper cites Item response theory for psychologists, 2004.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Item response theory for psychologists, 2004

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.136025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.668171Z digest=sha256:94e6c9c272484aa7bd1165fc0903c6cfdc2b2b8aa7eb7a97173d0f4f0ddfa0d3

Observation 1896d037-6f6a-41db-9854-1a75eea3dde7 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.670922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.670922Z digest=sha256:78c5450a8abe7b01f14a42decb0c59d52c8a9eae5d52f68482e126180e1463b0

Observation 700bd2d3-198a-429f-a4b0-079ebb0cdfff · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.674013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.674013Z digest=sha256:02e22c52dbe9a6160310aab510d5dcd03bf8051fa0cd42f32e0bf5d5fcb4fa90

Observation 71011636-bb2e-40ab-8946-2254b05067d1 · outbound

This paper cites an unresolved cited work.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:36:10.126584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.677295Z digest=sha256:49a9ee867151e92d61f4cdd99abd8b420f95b495042bca0618d4cbdf3c825a72

Observation c6c94b34-ba09-4f4a-b1c9-104b90a3ff31 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.680237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.680237Z digest=sha256:418a25846a7d64e33f1f317690d3e1dac468e506cceb8bbc38763801b0d46445

Observation 081c51c7-9493-4b03-b3ea-a873f45d064d · outbound

This paper cites Item response theory.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Item response theory

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.683350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.683350Z digest=sha256:54c7dc8b70f43b2309b01d0cd037d1fac463133258053d00e12a2c3952a38186

Observation 85011a5e-a23a-4f63-b776-710aa2781705 · outbound

This paper cites A tutorial on fisher information.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models A tutorial on fisher information

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.111480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.686436Z digest=sha256:a0a50afb1b537670c3f4714d81fefb5664801d0b472b9b9232bd33c2ea9c65a4

Observation a57fdb7d-a90b-4869-9f32-dcfbd4540ea1 · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.102147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.689427Z digest=sha256:d28f21750287d07abc7cd0ee6f16243c54eeee20ab924d0562e7df1b5eb429c9

Observation e48c24e9-070d-45a2-8416-70530b903c9d · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.692301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.692301Z digest=sha256:87c79a167406240278f6fc8f151a29bb933080fcbe563ce3e515b7a5f57f23fd

Observation 45a75fc9-f008-44f9-b42a-4806fb871dbc · outbound

This paper cites Theoremllama: Transforming general-purpose llms into lean4 experts, 2024.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Theoremllama: Transforming general-purpose llms into lean4 experts, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.087594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.695142Z digest=sha256:0a750ea73d1594089de79292f2a8725cf2a6a473cf19e7c8c95c28abef5a58af

Observation a526a880-7541-4b2e-8e37-bb6dc2793119 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Code Llama: Open Foundation Models for Code

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.697940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.697940Z digest=sha256:a13ed739da5526d99d3755a6e75a265b5483a616d40d6e8062fc9ce378fff239

Observation c3e68a2c-1051-4db1-8559-8c665027fa83 · outbound

This paper cites Qwen2 Technical Report.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Qwen2 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.701260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.701260Z digest=sha256:95f010b87ed5872c33c853985de71fd9f8726c1f439ce9a3fb0b7a9a1784125a

Observation 34fdb993-c71c-49ec-b219-148b407bb726 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Qwen2.5-Coder Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.704602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.704602Z digest=sha256:aaf6a55c416f6123e9d244d69f7196a6173b724eb9d7a27531f9b35a45cc4b00

Observation 9b1b55d0-395b-4a56-b449-b2552450d9e8 · outbound

This paper cites Lisa: Language models of isabelle proofs.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Lisa: Language models of isabelle proofs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.707633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.707633Z digest=sha256:2af9de298bd278eb02a5a61ccaa1553c3fa45ed7aba7030acdbc7ca27f449299

Observation 47a8011d-b6bd-46a0-a65c-0c78c83290c4 · outbound

This paper cites ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.710677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.710677Z digest=sha256:74295368453f7cdcbed4fa4dc72506fb01c0c569ec5d80e5a78e337b3ef0d712

Observation cc6b863d-1272-4197-a4af-98b0f0d62e0b · outbound

This paper cites an unresolved cited work.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:36:10.071927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.713621Z digest=sha256:a6d74c2075bab9d693349dd86d4291e5e69f25af9f7c900d5a5f387f8153bff1

Observation ccc95c8b-89dd-4762-b411-bbb3302c9cc5 · outbound

This paper cites A.2 Theorem Categories The theorems in miniF2F are classified using multiple criteria.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models A.2 Theorem Categories The theorems in miniF2F are classified using multiple criteria

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-09T17:36:09.826170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.716426Z digest=sha256:9733591c9cae63430634d53ceb52ef305ef4543c30368811bbe1a372db3d7cc2

Observation 1305f129-dd40-4bd0-b677-d1563d2588b8 · outbound

This paper cites an unresolved cited work.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:36:10.062820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.719406Z digest=sha256:0965c4203c053cfd11f994b67fec80845c77d218a92838cefe9c59dc4a4f6bb6

Observation 84e96882-466d-4eba-9803-81063e4eb9da · outbound

This paper cites This is because, in the miniF2F design, the test set is reserved for evaluation, while the validation set may have been used during model training [32].

Psychometric-Based Evaluation for Theorem Proving with Large Language Models This is because, in the miniF2F design, the test set is reserved for evaluation, while the validation set may have been used during model training [32]

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.054541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.722439Z digest=sha256:4c0883e1e5406e205426c2027473c0fa63c4ec7dd12b8ea887cb609d39179bb5

Observation fbc1cf20-f683-405d-934f-a7e8adb54c62 · outbound

This paper cites This is because competition problems tend to be more complex, but their higher complexity also leads to a lower discrimination.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models This is because competition problems tend to be more complex, but their higher complexity also leads to a lower discrimination

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:36:10.046114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.725413Z digest=sha256:45f1440407f004925bbc63b116febd0f9f853ce9413fc298c7ca2be6577696a4

Observation 73b796a4-de88-45db-8199-08e445c71525 · outbound

This paper cites an unresolved cited work.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:36:10.037359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T17:36:09.728402Z digest=sha256:583edbea1e3a9f07bf98946e5408f1eb224fad83b989a918d5d784d9bb72129e

Pith citing papers

No inbound Pith citation observations are available.