Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:07.631907Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 8 inbound Pith citation observations for arXiv:2505.00662.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:07.631907Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:33:12.754141Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:59:42.593762Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e42d56a4-ed69-4971-9198-49baa36b08d5 · outbound
DeepCritic: Deliberate Critique with Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97709941-4e7e-403b-8b0e-d25b482973ae · outbound
DeepCritic: Deliberate Critique with Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974dbcc6-c96d-4fde-8125-1a26d0070a00 · outbound
DeepCritic: Deliberate Critique with Large Language Models Weak-to-strong generalization: eliciting strong capabilities with weak supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 58091933-36c6-4c93-8183-71165a77637c · outbound
DeepCritic: Deliberate Critique with Large Language Models Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834a1e4a-85ab-415e-953f-d961c5e5d81d · outbound
DeepCritic: Deliberate Critique with Large Language Models Process Reinforcement through Implicit Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94cf3d7d-e678-41fb-9ac3-b0e1ed4250ab · outbound
DeepCritic: Deliberate Critique with Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 1 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f1964553-28bf-472a-b7f9-3e833c5e482a · outbound
DeepCritic: Deliberate Critique with Large Language Models LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb44d9c-21ac-4edb-8823-fc9806cb49c2 · outbound
DeepCritic: Deliberate Critique with Large Language Models Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ada40ff-c93d-4c04-8283-0d75bf02d69e · outbound
DeepCritic: Deliberate Critique with Large Language Models Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4fd0dfa-5f83-4b1e-9743-9523b5310af5 · outbound
DeepCritic: Deliberate Critique with Large Language Models A Survey on LLM-as-a-Judge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86df7ff2-6425-4b53-a235-f777e8d1aef3 · outbound
DeepCritic: Deliberate Critique with Large Language Models Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa0f929-3068-408e-b4d8-5e08b9962d8a · outbound
DeepCritic: Deliberate Critique with Large Language Models Measuring mathematical problem solving with the MATH dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed360aa-ba13-4d0d-867d-da9a2c1792e8 · outbound
DeepCritic: Deliberate Critique with Large Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed949d4b-fa91-4c34-b842-1edb570e7e07 · outbound
DeepCritic: Deliberate Critique with Large Language Models Qwen2.5-Coder Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa2dc00-54d4-4ff4-8b56-14596bd123ed · outbound
DeepCritic: Deliberate Critique with Large Language Models GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396c3529-11e2-46cf-a865-73c2399f4c04 · outbound
DeepCritic: Deliberate Critique with Large Language Models SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations , 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb363920-1f2e-4432-b755-767485c1fcc0 · outbound
DeepCritic: Deliberate Critique with Large Language Models CritiqueLLM: Towards an informative critique generation model for evaluation of large language model generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c81debc-3edb-4cd7-b1ab-dfbd6175ec03 · outbound
DeepCritic: Deliberate Critique with Large Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e30f58a-6ea0-49e9-a8db-0ffef06400cb · outbound
DeepCritic: Deliberate Critique with Large Language Models Criticeval: Evaluating large-scale language model as critic
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 999f43f7-7c61-4f52-b32c-4b86dd9ae24c · outbound
DeepCritic: Deliberate Critique with Large Language Models Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da94829-b43a-4a47-82f0-61d7e8de933b · outbound
DeepCritic: Deliberate Critique with Large Language Models Let’s verify step by step
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdd6e85-1949-4ec4-a055-f6f4d0187513 · outbound
DeepCritic: Deliberate Critique with Large Language Models Criticbench: Benchmarking llms for critique-correct reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f19ef65-55ac-4db3-821e-0d069bf55d2d · outbound
DeepCritic: Deliberate Critique with Large Language Models DeepSeek-V3 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdaaf89d-ae88-44ef-855b-7de8c552d2b9 · outbound
DeepCritic: Deliberate Critique with Large Language Models Critique Ability of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81d4b88-4319-4aed-88e1-39143338ec4e · outbound
DeepCritic: Deliberate Critique with Large Language Models Large language models surpass human experts in predicting neuroscience results
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20d355a4-f854-4e18-93d0-8d46c969a5a2 · outbound
DeepCritic: Deliberate Critique with Large Language Models Self-refine: Iterative refinement with self-feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16947f13-c97c-49be-be09-eb5680b7a1da · outbound
DeepCritic: Deliberate Critique with Large Language Models LLM Critics Help Catch LLM Bugs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4b7420-ac85-46f0-9add-751533ef5f31 · outbound
DeepCritic: Deliberate Critique with Large Language Models Introducing llama 3.1: Our most capable models to date
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51b3585a-3780-4ce6-a985-b0df45b861ee · outbound
DeepCritic: Deliberate Critique with Large Language Models Gpt-4 technical report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2263a578-60f4-4b7a-9483-504e085311bd · outbound
DeepCritic: Deliberate Critique with Large Language Models Learning to reason with llms, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 43493cf3-7a27-4d1a-b6c9-8aa9d60ebfa6 · outbound
DeepCritic: Deliberate Critique with Large Language Models Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a338c6-d337-4365-86f1-3d9fc6a08dfa · outbound
DeepCritic: Deliberate Critique with Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 68b51c03-b63e-4cc7-8d6a-2a362ba22fc1 · outbound
DeepCritic: Deliberate Critique with Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 414be28f-41b5-4703-9fe1-20279d941be3 · outbound
DeepCritic: Deliberate Critique with Large Language Models Code Llama: Open Foundation Models for Code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 862ad342-2e92-4109-a88c-87d879528ab0 · outbound
DeepCritic: Deliberate Critique with Large Language Models Self-critiquing models for assisting human evaluators
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4398c8-083e-449a-80fa-388a1e743476 · outbound
DeepCritic: Deliberate Critique with Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252de94a-084e-4b9e-a2c1-3bcbb4c1fb6c · outbound
DeepCritic: Deliberate Critique with Large Language Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08fd443-e38d-4a79-8748-886103911a6e · outbound
DeepCritic: Deliberate Critique with Large Language Models Self-Evolving Critique Abilities in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eeb2d92-23d1-471c-b2a0-9f6a5df4b901 · outbound
DeepCritic: Deliberate Critique with Large Language Models Solving math word problems with process- and outcome-based feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04e9fba-3e72-44ac-8210-4af93296e8de · outbound
DeepCritic: Deliberate Critique with Large Language Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e53525-f12a-4b5f-b6d7-a937858a5c34 · outbound
DeepCritic: Deliberate Critique with Large Language Models Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2ab938-56c2-4a43-937b-90c45b2b7bd2 · outbound
DeepCritic: Deliberate Critique with Large Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8970a22a-d1f2-417a-9064-b1a31b231a7f · outbound
DeepCritic: Deliberate Critique with Large Language Models Scaling inference computation: Compute-optimal inference for problem-solving with language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f05c4ddc-291c-42be-9d43-0b10c6a9e8a7 · outbound
DeepCritic: Deliberate Critique with Large Language Models Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3e3503-df2d-442f-b5ac-f3c754546b83 · outbound
DeepCritic: Deliberate Critique with Large Language Models Teaching language models to critique via reinforcement learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e5e0e3-c221-4b39-81ab-fc65524dd9dc · outbound
DeepCritic: Deliberate Critique with Large Language Models An implementation of generative prm, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 96d41639-67de-44bb-ae50-0a44c4202d45 · outbound
DeepCritic: Deliberate Critique with Large Language Models An implementation of generative prm
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0bb897-b685-47ce-a3da-7324de9580db · outbound
DeepCritic: Deliberate Critique with Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04121c85-a2ec-4fce-afe8-8696e9dc1d72 · outbound
DeepCritic: Deliberate Critique with Large Language Models Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651ff0c2-1952-4f53-9df8-672228e0031a · outbound
DeepCritic: Deliberate Critique with Large Language Models Towards thinking-optimal scaling of test-time compute for llm reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aff4f65-2a08-4fac-81ce-af5f7bdf339a · outbound
DeepCritic: Deliberate Critique with Large Language Models Metamath: Bootstrap your own mathematical questions for large language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47a06cc0-8abd-4003-a8b6-86a8ea8635c7 · outbound
DeepCritic: Deliberate Critique with Large Language Models Self-Rewarding Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57c6abf-80d9-40b1-9b90-c4595c6e3bee · outbound
DeepCritic: Deliberate Critique with Large Language Models MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db65d28f-b8d7-4bdd-9896-2e52600b7259 · outbound
DeepCritic: Deliberate Critique with Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ff0c52-adc4-4134-a911-bedbda4c35c1 · outbound
DeepCritic: Deliberate Critique with Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 67ebcffe-0f8f-462b-a129-f234fbadc9c7 · outbound
DeepCritic: Deliberate Critique with Large Language Models Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd3c7c6-1689-44fa-a671-286c4057c856 · outbound
DeepCritic: Deliberate Critique with Large Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78dafb33-6032-45bf-8069-b20d261107b8 · outbound
DeepCritic: Deliberate Critique with Large Language Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9b9d03c9-3229-416b-a452-e3ab2c57deb0 · outbound
DeepCritic: Deliberate Critique with Large Language Models Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5a96f44-7a9c-4d60-8628-311264c86cca · outbound
DeepCritic: Deliberate Critique with Large Language Models erroneous
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ad22a85-1b4c-49bd-8103-98a1074fff65 · inbound
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DeepCritic: Deliberate Critique with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d0a0dc-20a9-4b37-b7ad-11ab42174270 · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepCritic: Deliberate Critique with Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294f5c9b-2f2c-4f5c-84ac-77f27ec6074b · inbound
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback DeepCritic: Deliberate Critique with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adfb6db-1cd0-4732-b701-6c23138684ff · inbound
Learning from Language Feedback via Variational Policy Distillation DeepCritic: Deliberate Critique with Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 77ea421a-ac94-4d5a-967e-d6d653576187 · inbound
Trust Region On-Policy Distillation DeepCritic: Deliberate Critique with Large Language Models
Reference 166
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e9dfcf63-bd6a-4d79-881c-fefce6088ff4 · inbound
On the Position Bias of On-Policy Distillation DeepCritic: Deliberate Critique with Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b1822564-c25c-469d-a5c0-99b73532b58e · inbound
On the Position Bias of On-Policy Distillation DeepCritic: Deliberate Critique with Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4be92d90-751e-4fa9-88f4-907765b67579 · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning DeepCritic: Deliberate Critique with Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.