Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

As of 12 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 15 inbound Pith citation observations for arXiv:2411.16579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16579 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:02:57.382411Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:37:23.509928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:17:40.141555Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved52
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b85b934-7e4d-455c-bea5-dc4b1d428ff2 · outbound

This paper cites an unresolved cited work.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.053027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.053027Z digest=sha256:af179b167cc4c4e21464825d31a5cf99bfa0006ca0d4e445e0e6c4d948b8bb71

Observation 94f492fd-d4e8-4322-87d3-31e39ed469e0 · outbound

This paper cites GPT-4 Technical Report.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.058041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.058041Z digest=sha256:85fdd1796901cd3bfea5768b419a9cd2e486a3cb8423dee79b1a9f27bd0df725

Observation 5c65238e-4661-4a89-9401-86cad5c54540 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.062438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.062438Z digest=sha256:fcb20885043c0bf1bd3a184f4050a2177de2996b6aa0c5ef44c197442c0a54df

Observation d1426059-7b5e-48ca-a452-7dc41a796e9f · outbound

This paper cites Mistral 7B.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mistral 7B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.066763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.066763Z digest=sha256:767e7cb769f04d61e3c1ed6a417e52bdd6240fb3a38fefe637d984285a939e07

Observation 93f973da-c684-4eca-8a3f-341b780182af · outbound

This paper cites The Llama 3 Herd of Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.071218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.071218Z digest=sha256:83556bcf9858825d4961543faaa3cfe83d7952874bd957a6cd4b9ffa7ca8445a

Observation 9f03ad74-ca13-4858-ae7d-49e78ade05a0 · outbound

This paper cites Chi, Quoc V.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Chi, Quoc V

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.075651Z digest=sha256:a3ba10c1a44b4d01a2ff8562aeb7980ca2f0bdd3e2b5f9129993a7e717b04c8b

Observation 0a1a82af-11c1-43f4-bf9a-5d7fbce0bbcb · outbound

This paper cites Self-polish: Enhance reasoning in large language models via problem refinement.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-polish: Enhance reasoning in large language models via problem refinement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.079955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.079955Z digest=sha256:0c69a723332aa00cbb76479907cdfdf9133a881b2872ed69b7996b7178eb8514

Observation ed50e2f4-042d-4662-b195-78c4e6c8d05a · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reinforced Self-Training (ReST) for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.083939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.083939Z digest=sha256:79d598882e73176f0ec42b237ee332dd99e3b946f9376ef4826f1f3afec84d9e

Observation 130fbaa5-bfa7-4d57-ae40-7764142d84e2 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Tree of thoughts: Deliberate problem solving with large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.088052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.088052Z digest=sha256:5cee862dbcee8b6167684e9cdad41f9e73fa94135fe51379a7fead0e5f1c0730

Observation f76c35f2-4c68-4bb2-bab0-c34d8c36d669 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.091954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.091954Z digest=sha256:8772cffda13e3afdc0f8c019dd90dc1d07cd7fcea2267c984df0cfe78b1bf728

Observation 22af9057-d6a2-4667-86a3-4fa58103ef42 · outbound

This paper cites Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.095809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.095809Z digest=sha256:7f8fb8e413ba8ecdd1ffb0c05b583ac4df3e11647578ba2b804e7a6a984d0f2a

Observation a45ee4b7-2e0e-44f3-92d2-e5f618b89784 · outbound

This paper cites Narasimhan, and Yuan Cao.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Narasimhan, and Yuan Cao

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.099726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.099726Z digest=sha256:be5cd93d17f2f0bbcac39d4a31cb9590c5b329ed809ea04593fc8c9039692f09

Observation 9c8e6dd0-4841-4059-8203-229da2653d4a · outbound

This paper cites Learning to reason with llms, 9 2024.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Learning to reason with llms, 9 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.224784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.107371Z digest=sha256:497e595b90ee0c149bdd9ef2de095b3a363d70b302a786dc167575034a8a56f2

Observation c16b122c-e1ec-4ca1-9f9f-7967a0689750 · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reflexion: language agents with verbal reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.214036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.111256Z digest=sha256:2020a9b77f36c1c5fd27be5dcba00763c2574116ef61c6abaee39e31aa53536f

Observation 9bc88d17-6261-4d2f-a612-74d0399dda58 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.114944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.114944Z digest=sha256:aa926eb8d5914be43b6c2e00b632d2ea27993fbeb3c8e877eaa12d4be788a9bb

Observation 2b490c3a-3440-4798-b6a8-a5a1a2c82b60 · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.118753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.118753Z digest=sha256:3aca6e52a64a8da45e2d6f412cf81641d5d776071b6291081970566d9a51e45c

Observation bdcb9261-86a0-42a7-bb2a-09f121e47e72 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Language Models to Self-Correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.122914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.122914Z digest=sha256:103bbe5c984982685e49c30224c99bcafab45c5f9909358802dd3477db1ed1da

Observation 012ce922-4711-4dd4-9118-70730d848644 · outbound

This paper cites Generating sequences by learning to self-correct.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Generating sequences by learning to self-correct

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.202601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.127244Z digest=sha256:396492cb6c6be530f783be509ce706c8b1b5fa2a13758b5fa99cbd1d9b510663

Observation 2144d6c3-d0b3-490a-b89c-ae48bf6d13f3 · outbound

This paper cites Pride and prejudice: LLM amplifies self-bias in self-refinement.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Pride and prejudice: LLM amplifies self-bias in self-refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.191472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.131026Z digest=sha256:7c15cb3b0ea30568882872714de5fb899d79950251ceb76b3fd555beb38eaec3

Observation bb67ad25-58e2-4fcc-b837-62227c7a50e2 · outbound

This paper cites Selfee: Iterative self-revising llm empowered by self-feedback generation.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Selfee: Iterative self-revising llm empowered by self-feedback generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.180246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.134721Z digest=sha256:826965a584a89ba0463c623f1b2eec1c22ffbf0dadff6aad1bb6ed6c15bbf66d

Observation a6521300-0f7c-4653-b495-f14d13535223 · outbound

This paper cites Language models can solve computer tasks.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Language models can solve computer tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.168099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.138447Z digest=sha256:ec20213b852da07703cac9790cca694592c04396e60b450768190c129a06144d

Observation a69af469-747b-41aa-b80d-523013afd16e · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-critiquing models for assisting human evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.142310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.142310Z digest=sha256:d8f76cea83a31f5be667bd0ad4f741926cb41df356c385c039e75cc8df3b6900

Observation d33bfbb6-a0c1-4f9b-bef9-9b41eebfd00f · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-refine: Iterative refinement with self-feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.156112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.146154Z digest=sha256:94f81af7c2a45b303e83a2d7bb85b526f94c7a77ca8983b1981024e25dff5441

Observation 64ae9e86-e13a-4b03-94fb-fd3203b108a9 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large language models cannot self-correct reasoning yet

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.143170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.149736Z digest=sha256:f9d277bed3c9640a15adb2e92a895da7e0677d1d2d9291814255ab38687a3253

Observation 7fba3881-f542-4fa1-b9e4-33b0aeca8e00 · outbound

This paper cites RL4F: generating natural language feedback with reinforcement learning for repairing model outputs.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision RL4F: generating natural language feedback with reinforcement learning for repairing model outputs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.131161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.153659Z digest=sha256:04c1c316cad978e4e19966beeec6f7868161759503c5832b3bcceb269a4dff2e

Observation af4c5e23-97c0-49c0-a684-1f0500b8f7f9 · outbound

This paper cites N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.117592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.157147Z digest=sha256:32b975cf03c9be36242e6b542946ed2b7f9311fb8fb343d8ce914b176f96ce21

Observation e194bb0c-ecc1-498a-96c0-b752fb219996 · outbound

This paper cites Glore: When, where, and how to improve LLM reasoning via global and local refinements.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Glore: When, where, and how to improve LLM reasoning via global and local refinements

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.104913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.161128Z digest=sha256:5593daacf84001a0c61822cd1a5ace5ab26c7fbd61e1adcd8fb822873f96986e

Observation fd04a305-fb82-4246-8db9-5efdef3f83b8 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Constitutional AI: Harmlessness from AI Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.164895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.164895Z digest=sha256:96e0699735ae914ef83031c4000b2c483ac7ccb16a0cb3ff1d7d5e04089d19ad

Observation 14a6d641-dcb1-4ee2-95f4-e69ed53ebdf8 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring Progress on Scalable Oversight for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.168900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.168900Z digest=sha256:36c49bc76669f63c3897320a991a77ee2c74f94c49db84fe3d024b02bcbb98c1

Observation af6fb549-3f18-4913-8f48-b96a6aef19a9 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.092948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.172965Z digest=sha256:e68f4a6558d8cbaf5b813e079ab25b816229d29ff7932914f63dd1e089641a57

Observation ec8316a1-d2a4-4c13-a741-f4d2efe5a023 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.176993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.176993Z digest=sha256:22663eca91632915b323a6d3f1c12194babfef6f2ffdf07b1beacbb7f97ad512

Observation 0d27df5c-156d-4edb-998c-ccdb77c6dcd2 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.181044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.181044Z digest=sha256:24d728da229d9d83ead8601f0c9320699cf3eb6652be9d95b2ea66e342a4fe13

Observation 67fb4773-3482-437c-af28-a94b36d1c356 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.185103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.185103Z digest=sha256:fd92324286b10995b456f39c5052b061280942e931f0af00c29cbd3802dcb88f

Observation ea4f1dc3-f386-498a-9adc-e96f14b8b585 · outbound

This paper cites Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.189007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.189007Z digest=sha256:b6fabfb7235100ace421ea8d14ee964a474f8cde408cad755d6b4e894e1e5fa4

Observation 82f5d69d-a297-47b4-b333-11544a25e142 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Verifiers to Solve Math Word Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.193025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.193025Z digest=sha256:09080f7865ec295e32bdfdcec837c18e9595e8374ad3236d1293ec4b93770654

Observation 48c564a4-6a9f-4e63-bfad-2d72e49e58c7 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring mathematical problem solving with the MATH dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.197231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.197231Z digest=sha256:929d5bb8b9e71e999125d0f4cb7bed5ec3e1f935c926c0f146edfef21198b29c

Observation 7242bde7-3079-44ff-961a-59baeece291b · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.201073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.201073Z digest=sha256:fbb41ac6a712f3fdf7528ebd579250bdd3329ba653c8b9ab15e13db8b1ac5503

Observation ef045c5b-3a65-46bf-9aea-a40dc452f808 · outbound

This paper cites Evaluating mathematical reasoning of large language models: A focus on error identification and correction.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Evaluating mathematical reasoning of large language models: A focus on error identification and correction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.064008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.205010Z digest=sha256:355b3241a7dfa29d352b6f86e13a155a9ccdb5965441c178fd568ab91fbe86f0

Observation 88d756a0-ec1d-41f9-bc42-cadb0297302a · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.208906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.208906Z digest=sha256:dcb7a519f26169a2d738844448ffbe86fd288059a0936d9af31206a5fb615520

Observation 87d0a53f-000a-47de-8884-68101bd7bffe · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.213083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.213083Z digest=sha256:35117f43d7dbce6251e8937022e37a72be79756fa3dec80dc1c578253b9cd0ae

Observation d6ab7f3d-ef8a-4b33-9cfc-f7418fbd81d2 · outbound

This paper cites Le, Ed H.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Le, Ed H

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.048170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.217728Z digest=sha256:293b863732ee2e6e22cc500dc5d71d447cdb5bfda6ee9e17ee451d66bafdd66c

Observation bb421daf-65bf-452b-8d93-5fd5540e663f · outbound

This paper cites Large language models can self-improve.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large language models can self-improve

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.033956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.221412Z digest=sha256:9c99592ee5083ec521cfe738d701539caafd1e1ee3e13b38a09d40116c3e1837

Observation a140117e-8ccd-4d50-820a-96af5bac56c3 · outbound

This paper cites ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.225484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.225484Z digest=sha256:e7459be51e5f31e7bd358b5589db7c970dc52dc422aa3d693154ae9dd25adfa6

Observation 9bf0d1da-e5ec-4971-8d03-d054db731fe6 · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.230261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.230261Z digest=sha256:df4cb125e034a75492156efd4cf60a04d2166a1e12e5faa21215c138d83ae5bf

Observation 3e969d2a-a5d1-4d54-97bf-24e39de11f97 · outbound

This paper cites Re-rest: Reflection-reinforced self-training for language agents.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Re-rest: Reflection-reinforced self-training for language agents

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:58.021031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.234166Z digest=sha256:2c5ecb261879e53badea96da5d483f8d223049ed55ca05365a0f5cfd078a7d25

Observation dbc49a4f-f48a-4f22-9da2-f6a508890715 · outbound

This paper cites an unresolved cited work.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:02:58.009509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.238113Z digest=sha256:f20e83d0b8cf78d3a448b2ca10883d96f8b8007f17ba439671e8f5d1f32a29cc

Observation 1691a1df-8419-4c99-bfe4-080641d630aa · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.241810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.241810Z digest=sha256:e6cc524af0f1032743c7aa9234f6151640575f93c747d57c54aad134a00a0dd9

Observation 34f8f316-6282-4d03-bfdf-25d850dbdcff · outbound

This paper cites Progress or Regress? Self-Improvement Reversal in Post-training.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Progress or Regress? Self-Improvement Reversal in Post-training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.245849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.245849Z digest=sha256:fadfdd68820178b8650035290ead1209eb7858f684a6f0c630c417a1586572a3

Observation bf24d1b9-42ea-44e1-8556-595620fcb0b4 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Qwen2.5: A party of foundation models, September 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.252023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.252023Z digest=sha256:2f8f6a234970c15375621c4f31e181eba0829a20ead0b1c89b164784bdd74094

Observation 2b244a35-1f88-43a7-bcf1-afa4f169b5c6 · outbound

This paper cites Selfcheck: Using llms to zero-shot check their own step-by-step reasoning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Selfcheck: Using llms to zero-shot check their own step-by-step reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.256215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.256215Z digest=sha256:02500f92ec0aa93f5b5806385cfdea3d8333d4f78aa5aac87c710abbcf65f8e1

Observation 5fe4b374-f31b-4527-a686-3133d1c68f29 · outbound

This paper cites an unresolved cited work.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:02:57.982839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.262193Z digest=sha256:da13c973c50680296b84b230558626de238c28b860060b457043d8c208ac902b

Observation 1ae59c0b-ab6c-4aac-b878-98f5b50aadb7 · outbound

This paper cites Mondal, and Jyoti Prakash Sahoo.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mondal, and Jyoti Prakash Sahoo

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.970806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.266388Z digest=sha256:d0952a9be135f01cd4c72603c3fc5f72eb1f85550443c5d745c350ccb0254811

Observation 24fb9d47-5b30-42ca-aec4-858dd01ad312 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.958655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.274458Z digest=sha256:b86fbf573fe4ffefe0052c5d26516f9b916630a8476d7bf469d51b97bfc15a10

Observation 599203ec-1368-4c74-ac60-8e2e8612a6ba · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.278896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.278896Z digest=sha256:21b2faacc0c679748df25b848f976535389f88cd2ccb629bd97978399574a373

Observation ca763e23-2533-4af3-843c-50c52e14451c · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.284688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.284688Z digest=sha256:0fda9946aa9511ceba2f20bd2fbc87dab3b5282e758ab0434adc381573e16340

Observation 39f744c1-b888-4e41-8a83-2dc4df53fa70 · outbound

This paper cites Baraniuk.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Baraniuk

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.946968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.289486Z digest=sha256:fd35d4cef2344e72656fb0589e453fc6c7576bcf9d5b36d38f559b07942c08b6

Observation 07f524a7-62e5-4f5e-adc4-807c59ae17bc · outbound

This paper cites Let’s verify step by step.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Let’s verify step by step

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.293636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.293636Z digest=sha256:ce4558f94a930faa8aab297a1d7a7ac81fe60931f8ea697b104129064fff9b4d

Observation eb958596-548b-4030-a356-a856f6a6f555 · outbound

This paper cites Improving discriminative capability of reward models in RLHF using contrastive learning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Improving discriminative capability of reward models in RLHF using contrastive learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.927529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.299359Z digest=sha256:5b52eaa7249ec10dabd7570373a37aebd109def953187211a63f07f6fb7d248d

Observation 91ae101e-4aa1-4e22-b5dc-fa35983840bc · outbound

This paper cites Improving Code Generation by Training with Natural Language Feedback.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Improving Code Generation by Training with Natural Language Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.303363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.303363Z digest=sha256:1151108a50746308087699faf4c2079473f3174b782573aa320541271feec454

Observation 954f32ce-ea5b-4f23-9461-f87c5059f4b7 · outbound

This paper cites REFINER: reasoning feedback on intermediate representations.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision REFINER: reasoning feedback on intermediate representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.915523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.307731Z digest=sha256:daee1a726d6e14b1ac135b0f02f0b2027988a2baceadb5fa4af6a67c1bc4b057

Observation 92a539d3-a388-4bf8-9cce-fcb2a9ef830e · outbound

This paper cites Chain-of-verification reduces hallucination in large language models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Chain-of-verification reduces hallucination in large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.311887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.311887Z digest=sha256:69537c5e5b9a036a9ebf782250843ca9dc6195d5b6bca90c61d4ea6564ae481e

Observation e1948d6d-46a3-47b5-b3e5-08d72aebc32f · outbound

This paper cites Shepherd: A Critic for Language Model Generation.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Shepherd: A Critic for Language Model Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.315505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.315505Z digest=sha256:d52aabf2c3e125a478077f362d59732359a0ba34ab2b529c7ae30d8d8953b80e

Observation 795dfabc-4141-4ee9-9b55-b4a768b7ea63 · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.319981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.319981Z digest=sha256:f23e2da9891a0d679645c195d6a35fa4e343543d2f4b57a089031bacf6f062af

Observation 774731a3-8338-41ce-b126-8f6c637b6933 · outbound

This paper cites Solving challenging math word problems using GPT-4 code interpreter with code-based self-verification.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Solving challenging math word problems using GPT-4 code interpreter with code-based self-verification

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.324170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.324170Z digest=sha256:d31d11b2b9c995651ee2694692b800b6801780aa374fffb7f0b697c0d885bb12

Observation b434e77e-61b7-4d83-b395-88e19cc9580f · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.329005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.329005Z digest=sha256:327fbf939a6f75e0527fc08a3b362e9d65c36f1ef7dae9c2165502f14d57a6bc

Observation 422a3b00-9233-4aac-8e9b-fab3048c7751 · outbound

This paper cites Critique-out-Loud Reward Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Critique-out-Loud Reward Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.332836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.332836Z digest=sha256:fb83e3eec5f995bdcd005315f0dd1bf339c6cbf54db3344134e7cda0495e36e6

Observation c8ad3372-b218-494d-b01e-37cc0c3a104c · outbound

This paper cites Critic-cot: Boosting the reasoning abilities of large language model via chain-of-thoughts critic, 2024.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Critic-cot: Boosting the reasoning abilities of large language model via chain-of-thoughts critic, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.891035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.336781Z digest=sha256:6b7f55181560a6ac82450e2b4ee862dd4619c7349935d6f05b34a91497042f0c

Observation 5fac2cbc-9969-4738-8c1b-220248c459f6 · outbound

This paper cites Du, and Beibin Li.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Du, and Beibin Li

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.879844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.340817Z digest=sha256:7feb6315ff9269da1d9439c68ad7d48616b73fe6d88affd81e58871f3ac0e3bc

Observation f2361f55-f8e0-445e-9b32-f684091edaec · outbound

This paper cites Ovm, outcome-supervised value models for planning in mathematical reasoning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Ovm, outcome-supervised value models for planning in mathematical reasoning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.867564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.345800Z digest=sha256:d8fa2a7fb49d876094146c0cc4f87a648591e7d4d0ea16bc0bed4b5314315391

Observation d0182e75-c839-4538-8707-a8e4e9b60d2e · outbound

This paper cites Reasoning with language model is planning with world model.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reasoning with language model is planning with world model

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.856336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.349713Z digest=sha256:ce4aa9e8616aa2a7d8444193f7f58a776985dd40532fce7ae17f9b476596df37

Observation 20e0f25b-9982-411c-b1cc-8652aed6fd47 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision AlphaMath Almost Zero: Process Supervision without Process

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.353698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.353698Z digest=sha256:870c89da1f7e22c5dc40ea6699058cbd0313aaf99c1ce3d0605c4257497fa4a2

Observation 4a67b251-f41c-42ba-a47d-927695d4d923 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.358278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.358278Z digest=sha256:e0afb9318f2de83dbe903ae5f25d3f0b19e5b75ef5fc58042865aaac1bb90140

Observation 67e3159a-8a46-4147-8b9b-7bb9fed0ef89 · outbound

This paper cites Measuring massive multitask language understanding.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring massive multitask language understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.362740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.362740Z digest=sha256:8a138f4231e04c2ddbc341bcb6bf95e44794b579a18e0dd15b1dd5ddeb761865

Observation 57bb5f76-3bdb-46ef-bdca-8e8ef5a1beee · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.366306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.366306Z digest=sha256:cc047d59691edec243671bda8c9be0b7a525dd48c43d775147f01b1c8507d9f9

Observation ef4192da-68bc-4261-84f7-0de63c9cbb10 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Evaluating Large Language Models Trained on Code

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.370629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.370629Z digest=sha256:2fc963e259da135ecee75328a34fef1b1f86757f5d7ba6bcc531bd3c6f4dc8d7

Observation 93d8110f-bc3e-4cf2-b6df-4d47b64ebe83 · outbound

This paper cites Program Synthesis with Large Language Models.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Program Synthesis with Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.374858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.374858Z digest=sha256:8590c6529fc220ceb87037da1ffd9f296cecc50110b0b049df5b46e3d437bfe3

Observation c771502a-0222-42a2-b935-c432d9d54c64 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:02:57.837662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:02:57.378894Z digest=sha256:b6e8bb735ee873fa052c2d2d96a97faa21d49f3261eca942eafc114e9e05f873

Observation f26ec808-d6a6-4b73-a46f-b56e2d85372a · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.382411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.382411Z digest=sha256:4de5e8e93369799d4361f937400eb70212f661d09a01021f4dbcff81ae4a94b9

Observation 04e7144a-3a3c-4394-bdd5-e78bc2a7f891 · outbound

This paper cites an unresolved cited work.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work

Reference 2023

Resolution
parse uncertain
no resolver link, observed 2026-08-12T13:02:57.103229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.103229Z digest=sha256:19bf167435119d8491ca25ab8385619cd85b81a196f757149dfb2d276a06e0b0

Pith citing papers

Observation c5b677c2-a98c-4230-8e5b-7383d98fe7b3 · inbound

Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying cites this paper.

Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:23.509928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:23.509928Z digest=sha256:97c350a6890ec9d333d06e43b36d31e9dc8b9ffd6b81c646f013dcc6c651b20b

Observation c206319f-92cb-444b-9adf-3a15c3694d30 · inbound

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training cites this paper.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.691028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.691028Z digest=sha256:27b2aefc1f9baa6adb0495bf4d5df787af1455deb7b5062732a948f547c0a785

Observation d2ac63d0-277d-4136-855b-3533293a6cef · inbound

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search cites this paper.

Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:57:48.778636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:57:48.778636Z digest=sha256:bde78611346d83f93c89066991dfded91a79bdc2a0b2d93f92e0150510c36bb1

Observation 5b592302-f65d-4675-a45c-b0b19ad6b29d · inbound

TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling cites this paper.

TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:38.405708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:00:38.405708Z digest=sha256:6fbc9e123de5e1de7e32d6953a787f697cead77eea42e2f57b0f4659df44b7eb

Observation e23b3e40-b19e-44cb-9332-01a0262f1503 · inbound

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning cites this paper.

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:19.417118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:45:19.417118Z digest=sha256:b88908bd981fb7f2ec40bae55c4fcd457f91f886f6ceaf72507180c1003dc24c

Observation 1f10f4ed-9f2c-40c9-bd52-d9d96a11b99d · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 289

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:10.512934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:10.512934Z digest=sha256:46f7a4381f7ee762742ed28735b536ddc3ff28be96456b3876d220b7f079b1d2

Observation 1e88dd56-b72d-4bbb-a544-7ee3bc701ca1 · inbound

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving cites this paper.

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:59.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:48:59.112150Z digest=sha256:af82af5509b128004f91387aa104d44ea23687db11b7585a302683132dc62de8

Observation 9615e3cf-cb0d-45df-8835-0d74c2c1830c · inbound

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning cites this paper.

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:47.251015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:30:47.251015Z digest=sha256:fc5a5a0db58ebcc482818c687a3ee78ee480c3c1bc520301ebc248578e319a14

Observation 83c9f6db-1dba-49fa-94ed-9c3d312980ed · inbound

R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems cites this paper.

R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:28.118937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:28.118937Z digest=sha256:787b7ef1ba830706ba438d51f44f8ec9485862c0d99ae30544b176888f397375

Observation 7aacda37-b046-471e-be59-73e724a8bec2 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:14.944730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:e371c3746e5725215e45953650562c3796af1c4f31c0f6d3d0267ad7c3b0017b

Observation 3396766a-ad2a-4185-b315-f00feb8a3e76 · inbound

Learning from Natural Language Feedback for Personalized Question Answering cites this paper.

Learning from Natural Language Feedback for Personalized Question Answering Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:56:53.045875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T22:55:16.228037Z digest=sha256:ac310178bad1a3f35bb8bad9e2bb666ce192597f58957d6b11fb6fe8c71a047a

Observation 6eafc153-32dd-404d-9036-6b9ae1198fbe · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:12.780313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:12.780313Z digest=sha256:1b2f614ab22b3f7292c749963f391b8867a7acf95d8aca98dd3bc64a13c67961

Observation f5360fa2-2d32-4f45-aa09-c34df55eb6aa · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:03:04.330035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:255a36f0b28853bb87c63a062f2460ac5574ca6c191375beffce108a965279ca

Observation 6207f2ce-2e63-428e-8042-23e0caf6e0ae · inbound

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning cites this paper.

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:49:00.673885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T20:47:16.236629Z digest=sha256:9904b115c247c91e1756463574708660bd40770fd9a65a859bf8d4853cc51d30

Observation 39681fe5-9020-4b3e-8b23-27317878f03d · inbound

A History-Aware Visually Grounded Critic for Computer Use Agents cites this paper.

A History-Aware Visually Grounded Critic for Computer Use Agents Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:17:40.143173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T13:20:32.432002Z digest=sha256:9077591dc34207f8fe69aaa072db94d98c09558b7bf3c423128a00a36b3c7550