Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:02:57.382411Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 15 inbound Pith citation observations for arXiv:2411.16579.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:02:57.382411Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:37:23.509928Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T05:17:40.141555Z
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b85b934-7e4d-455c-bea5-dc4b1d428ff2 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f492fd-d4e8-4322-87d3-31e39ed469e0 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c65238e-4661-4a89-9401-86cad5c54540 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1426059-7b5e-48ca-a452-7dc41a796e9f · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mistral 7B
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f973da-c684-4eca-8a3f-341b780182af · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f03ad74-ca13-4858-ae7d-49e78ade05a0 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Chi, Quoc V
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a1a82af-11c1-43f4-bf9a-5d7fbce0bbcb · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-polish: Enhance reasoning in large language models via problem refinement
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed50e2f4-042d-4662-b195-78c4e6c8d05a · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reinforced Self-Training (ReST) for Language Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130fbaa5-bfa7-4d57-ae40-7764142d84e2 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Tree of thoughts: Deliberate problem solving with large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76c35f2-4c68-4bb2-bab0-c34d8c36d669 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22af9057-d6a2-4667-86a3-4fa58103ef42 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45ee4b7-2e0e-44f3-92d2-e5f618b89784 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Narasimhan, and Yuan Cao
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8e6dd0-4841-4059-8203-229da2653d4a · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Learning to reason with llms, 9 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c16b122c-e1ec-4ca1-9f9f-7967a0689750 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reflexion: language agents with verbal reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9bc88d17-6261-4d2f-a612-74d0399dda58 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b490c3a-3440-4798-b6a8-a5a1a2c82b60 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdcb9261-86a0-42a7-bb2a-09f121e47e72 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Language Models to Self-Correct via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012ce922-4711-4dd4-9118-70730d848644 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Generating sequences by learning to self-correct
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2144d6c3-d0b3-490a-b89c-ae48bf6d13f3 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Pride and prejudice: LLM amplifies self-bias in self-refinement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb67ad25-58e2-4fcc-b837-62227c7a50e2 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Selfee: Iterative self-revising llm empowered by self-feedback generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6521300-0f7c-4653-b495-f14d13535223 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Language models can solve computer tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a69af469-747b-41aa-b80d-523013afd16e · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-critiquing models for assisting human evaluators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33bfbb6-a0c1-4f9b-bef9-9b41eebfd00f · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Self-refine: Iterative refinement with self-feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64ae9e86-e13a-4b03-94fb-fd3203b108a9 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large language models cannot self-correct reasoning yet
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7fba3881-f542-4fa1-b9e4-33b0aeca8e00 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision RL4F: generating natural language feedback with reinforcement learning for repairing model outputs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af4c5e23-97c0-49c0-a684-1f0500b8f7f9 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision N., Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e194bb0c-ecc1-498a-96c0-b752fb219996 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Glore: When, where, and how to improve LLM reasoning via global and local refinements
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd04a305-fb82-4246-8db9-5efdef3f83b8 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Constitutional AI: Harmlessness from AI Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a6d641-dcb1-4ee2-95f4-e69ed53ebdf8 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring Progress on Scalable Oversight for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af6fb549-3f18-4913-8f48-b96a6aef19a9 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec8316a1-d2a4-4c13-a741-f4d2efe5a023 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d27df5c-156d-4edb-998c-ccdb77c6dcd2 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fb4773-3482-437c-af28-a94b36d1c356 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4f1dc3-f386-498a-9adc-e96f14b8b585 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f5d69d-a297-47b4-b333-11544a25e142 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Training Verifiers to Solve Math Word Problems
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c564a4-6a9f-4e63-bfad-2d72e49e58c7 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring mathematical problem solving with the MATH dataset
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7242bde7-3079-44ff-961a-59baeece291b · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef045c5b-3a65-46bf-9aea-a40dc452f808 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Evaluating mathematical reasoning of large language models: A focus on error identification and correction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88d756a0-ec1d-41f9-bc42-cadb0297302a · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d0a53f-000a-47de-8884-68101bd7bffe · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ab7f3d-ef8a-4b33-9cfc-f7418fbd81d2 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Le, Ed H
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb421daf-65bf-452b-8d93-5fd5540e663f · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Large language models can self-improve
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a140117e-8ccd-4d50-820a-96af5bac56c3 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf0d1da-e5ec-4971-8d03-d054db731fe6 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e969d2a-a5d1-4d54-97bf-24e39de11f97 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Re-rest: Reflection-reinforced self-training for language agents
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dbc49a4f-f48a-4f22-9da2-f6a508890715 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1691a1df-8419-4c99-bfe4-080641d630aa · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f8f316-6282-4d03-bfdf-25d850dbdcff · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Progress or Regress? Self-Improvement Reversal in Post-training
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf24d1b9-42ea-44e1-8556-595620fcb0b4 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Qwen2.5: A party of foundation models, September 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b244a35-1f88-43a7-bcf1-afa4f169b5c6 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe4b374-f31b-4527-a686-3133d1c68f29 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ae59c0b-ab6c-4aac-b878-98f5b50aadb7 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Mondal, and Jyoti Prakash Sahoo
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24fb9d47-5b30-42ca-aec4-858dd01ad312 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 599203ec-1368-4c74-ac60-8e2e8612a6ba · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca763e23-2533-4af3-843c-50c52e14451c · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Curse of Recursion: Training on Generated Data Makes Models Forget
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39f744c1-b888-4e41-8a83-2dc4df53fa70 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Baraniuk
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07f524a7-62e5-4f5e-adc4-807c59ae17bc · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Let’s verify step by step
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb958596-548b-4030-a356-a856f6a6f555 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Improving discriminative capability of reward models in RLHF using contrastive learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91ae101e-4aa1-4e22-b5dc-fa35983840bc · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Improving Code Generation by Training with Natural Language Feedback
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954f32ce-ea5b-4f23-9461-f87c5059f4b7 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision REFINER: reasoning feedback on intermediate representations
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 92a539d3-a388-4bf8-9cce-fcb2a9ef830e · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Chain-of-verification reduces hallucination in large language models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1948d6d-46a3-47b5-b3e5-08d72aebc32f · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Shepherd: A Critic for Language Model Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795dfabc-4141-4ee9-9b55-b4a768b7ea63 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774731a3-8338-41ce-b126-8f6c637b6933 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Solving challenging math word problems using GPT-4 code interpreter with code-based self-verification
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b434e77e-61b7-4d83-b395-88e19cc9580f · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 422a3b00-9233-4aac-8e9b-fab3048c7751 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Critique-out-Loud Reward Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ad3372-b218-494d-b01e-37cc0c3a104c · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Critic-cot: Boosting the reasoning abilities of large language model via chain-of-thoughts critic, 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5fac2cbc-9969-4738-8c1b-220248c459f6 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Du, and Beibin Li
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f2361f55-f8e0-445e-9b32-f684091edaec · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Ovm, outcome-supervised value models for planning in mathematical reasoning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0182e75-c839-4538-8707-a8e4e9b60d2e · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Reasoning with language model is planning with world model
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20e0f25b-9982-411c-b1cc-8652aed6fd47 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision AlphaMath Almost Zero: Process Supervision without Process
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a67b251-f41c-42ba-a47d-927695d4d923 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e3159a-8a46-4147-8b9b-7bb9fed0ef89 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring massive multitask language understanding
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57bb5f76-3bdb-46ef-bdca-8e8ef5a1beee · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4192da-68bc-4261-84f7-0de63c9cbb10 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Evaluating Large Language Models Trained on Code
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d8110f-bc3e-4cf2-b6df-4d47b64ebe83 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Program Synthesis with Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c771502a-0222-42a2-b935-c432d9d54c64 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f26ec808-d6a6-4b73-a46f-b56e2d85372a · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e7144a-3a3c-4394-bdd5-e78bc2a7f891 · outbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Unresolved cited work
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5b677c2-a98c-4230-8e5b-7383d98fe7b3 · inbound
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c206319f-92cb-444b-9adf-3a15c3694d30 · inbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ac63d0-277d-4136-855b-3533293a6cef · inbound
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b592302-f65d-4675-a45c-b0b19ad6b29d · inbound
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23b3e40-b19e-44cb-9332-01a0262f1503 · inbound
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f10f4ed-9f2c-40c9-bd52-d9d96a11b99d · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 289
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e88dd56-b72d-4bbb-a544-7ee3bc701ca1 · inbound
Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9615e3cf-cb0d-45df-8835-0d74c2c1830c · inbound
Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c9f6db-1dba-49fa-94ed-9c3d312980ed · inbound
R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aacda37-b046-471e-be59-73e724a8bec2 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3396766a-ad2a-4185-b315-f00feb8a3e76 · inbound
Learning from Natural Language Feedback for Personalized Question Answering Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6eafc153-32dd-404d-9036-6b9ae1198fbe · inbound
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5360fa2-2d32-4f45-aa09-c34df55eb6aa · inbound
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6207f2ce-2e63-428e-8042-23e0caf6e0ae · inbound
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 39681fe5-9020-4b3e-8b23-27317878f03d · inbound
A History-Aware Visually Grounded Critic for Computer Use Agents Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.