Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2503.07572.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:48.065006Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:59:44.665948Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 48f99603-5474-42f5-b4f9-e763540754af · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 67d9088e-a32b-4723-8867-4a9ef489ad46 · inbound
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d92f2bc-acec-4275-80f0-22abdc1e0a66 · inbound
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d80b30-3761-4e5e-b55e-c5e36827d8ce · inbound
ProgRM: Build Better GUI Agents with Progress Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eff9656-260f-4662-8891-409e8e3b4b38 · inbound
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca712478-cdd0-4219-a452-36072aac2daf · inbound
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523b62c2-f102-40f1-9ae0-22fbd4ff1678 · inbound
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffc267b-be9d-44f8-bf7a-2c865a31dcf0 · inbound
AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95308d0d-230e-4fb2-9a55-11b1e9187da4 · inbound
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ece831f-5459-469e-af83-01d6f0ccb08a · inbound
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2f3e066-ffe6-42d3-8866-d5cb62a0cfe3 · inbound
Sample Complexity and Representation Ability of Test-time Scaling Paradigms Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dde6b51-6090-42df-b014-a3bd8579b19a · inbound
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3b642e89-0914-4e15-873d-ee3d7fab022d · inbound
How Far Are We from Optimal Reasoning Efficiency? Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2ad100-cff4-41f3-bf36-373daa431c22 · inbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83245270-4548-4a11-8767-ddf00921963f · inbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4594d305-d6ef-4987-a7ea-12d9132564ab · inbound
Formalizing Learning from Language Feedback with Provable Guarantees Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e46486-f059-4c37-a975-434082a63560 · inbound
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278f6c69-2b5f-40b4-9fda-74409491e923 · inbound
Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41da26e3-4e5d-4fb3-8c07-362fc3b80f4f · inbound
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636d64aa-4260-403f-b138-5acc801f305a · inbound
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616a595b-8c34-490b-a32e-44a2b1fea9e9 · inbound
OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6491f5ab-6699-46ce-a21b-57d51d5e1cac · inbound
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b7ebd3-8f31-4636-ab99-73c31d9664ff · inbound
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0887f2a1-f622-4a95-928f-fac43d84f501 · inbound
ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd9fa6fe-b7a2-425b-a0b9-596688c4bcad · inbound
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4636812f-dd2f-435f-be1d-691f9d98dc4b · inbound
Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c72c55e-5d48-4e43-bf7b-f95054e9e5bc · inbound
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c1130d8-29e0-4132-a5fb-cceb8f533b3a · inbound
Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6698c8d1-37ab-4ce2-80f0-bd785d300b04 · inbound
ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ff69dcd-fcf7-4372-bb43-db31ef2cab14 · inbound
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1d05090-8e18-4e84-a6a4-0e01974bebef · inbound
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 445599e8-7336-47a9-84c1-fc24e4dd4631 · inbound
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 86f5545e-8def-4fb0-badb-00ca4cf91a46 · inbound
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 059e7f97-c690-44a9-9d51-8eb20603a599 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 63dd714d-a3d3-4cd1-967a-58432dc6f5b5 · inbound
Hint Tuning: Less Data Makes Better Reasoners Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e4808b51-4c5a-4ca8-9a52-22458dfbca55 · inbound
Hint Tuning: Less Data Makes Better Reasoners Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c12158d-10a5-4b28-bf02-1bdde217889a · inbound
Understanding and Mitigating Premature Confidence for Better LLM Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 17a0c6e7-55d2-4b2e-81e2-e9977bb01337 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d103120-ebce-4d84-9874-fd1eb5670917 · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3b0fc51b-ff58-405e-ad0f-42a374f2bc6b · inbound
Addressing Over-Refusal in LLMs with Competing Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da886d2f-5d1b-45f5-b92c-c7ec909f6a29 · inbound
Structured Thoughts For Improved Reasoning And Context Pruning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.