Pith. sign in

Paper Citation Record · LEDGER

Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2503.07572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07572 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:48.065006Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.665948Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 48f99603-5474-42f5-b4f9-e763540754af · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:29:56.717777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:e73a20a7339f754de42dc5cdb00465f2ceadc804e6cb62dbc2b232e66b871520

Observation 67d9088e-a32b-4723-8867-4a9ef489ad46 · inbound

When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning cites this paper.

When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:48.065006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:25:48.065006Z digest=sha256:94afe5c82760cb54a2d6523951d8a60be16aa3ce1827da98f7f9c8d770175ca4

Observation 1d92f2bc-acec-4275-80f0-22abdc1e0a66 · inbound

ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy cites this paper.

ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:52.237254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:17:52.237254Z digest=sha256:b91f86b9501e43354bfa5fe50580e954fb8fe7ad92da41ab7d6810b7df3fbd4c

Observation 18d80b30-3761-4e5e-b55e-c5e36827d8ce · inbound

ProgRM: Build Better GUI Agents with Progress Rewards cites this paper.

ProgRM: Build Better GUI Agents with Progress Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.502834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.502834Z digest=sha256:6dce20724086139ad3a70eb24a00724bd31f17c37e6a7b191756752ad1dcf92d

Observation 3eff9656-260f-4662-8891-409e8e3b4b38 · inbound

Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models cites this paper.

Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:58:36.490897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:58:36.490897Z digest=sha256:7598cb903f5423606cfafaf46fd4f0bb3bca797cfa698c6e3d0d96d8fa900511

Observation ca712478-cdd0-4219-a452-36072aac2daf · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:03.913311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:03.913311Z digest=sha256:60e7984532ad9b3a1b560c4dda5de8de9a16e1c40f2c96dbc6932fcc6769f035

Observation 523b62c2-f102-40f1-9ae0-22fbd4ff1678 · inbound

Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models cites this paper.

Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.577160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.577160Z digest=sha256:67f8c31360d8e0bf999b538b11093846751e61cf080c6ccb0571f15228a6bf2b

Observation 7ffc267b-be9d-44f8-bf7a-2c865a31dcf0 · inbound

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up cites this paper.

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.461138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.461138Z digest=sha256:fab52497a2248340ecd2d045acf6dc366918a9e1adefaa71d9703ab8085263dd

Observation 95308d0d-230e-4fb2-9a55-11b1e9187da4 · inbound

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time cites this paper.

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:43.664836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:43.664836Z digest=sha256:27fec4b50f4f6b6efafde667ef72cc8b7d5f1663153e3c2db38d5535db36f2b7

Observation 5ece831f-5459-469e-af83-01d6f0ccb08a · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:40.992654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:40.992654Z digest=sha256:f0dcb17e8b753340563715925b6726ce3981a1d6d33eb15f6cb023039b3b906e

Observation c2f3e066-ffe6-42d3-8866-d5cb62a0cfe3 · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:37.612022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:37.612022Z digest=sha256:1be424b3229156d01af25708873163434c44f122ee32a7aedafeb997ee380867

Observation 1dde6b51-6090-42df-b014-a3bd8579b19a · inbound

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity cites this paper.

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:10:31.523933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:10:31.440921Z digest=sha256:01ba903ffbe8084ce13a3f3f592ce821ae555b13845e5238d293ba1e1e6fa679

Observation 3b642e89-0914-4e15-873d-ee3d7fab022d · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.378369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.378369Z digest=sha256:1c49384948d6603083b2bae219470c69cd101dbe50f93116f783a45b6a093c0c

Observation 8d2ad100-cff4-41f3-bf36-373daa431c22 · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.229084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.229084Z digest=sha256:a9860c91347315d2bf8314e16a2449457ef7f82f7b61138c856c98fe358e3fe4

Observation 83245270-4548-4a11-8767-ddf00921963f · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.410502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.410502Z digest=sha256:1461cb035ded8033071228f1e30bc4f9d2b84cee6660cbb30bccd1809fe5f158

Observation 4594d305-d6ef-4987-a7ea-12d9132564ab · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:19.223327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:19.223327Z digest=sha256:3f8f25e52a7522d54d21a7b4d49fec933873e3c45dc77cc3bb59242fea8dc7c1

Observation 90e46486-f059-4c37-a975-434082a63560 · inbound

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty cites this paper.

Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:38:09.500490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:38:09.500490Z digest=sha256:029cd93e2609d49663cc2a616b9882c07e3e6f2140ec067d6fdbd8411e7e1d58

Observation 278f6c69-2b5f-40b4-9fda-74409491e923 · inbound

Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement cites this paper.

Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:49.571611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:49.571611Z digest=sha256:6cf99c62641c425a920f6b7417f93224b0df59864a795f3ea797d0fa647dda9c

Observation 41da26e3-4e5d-4fb3-8c07-362fc3b80f4f · inbound

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model cites this paper.

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:09.631754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:09.631754Z digest=sha256:f63ce7c55cf6d6381f6dcb8d64d990dc9e02a0b3fa41e47feaf4844bd04e4c2e

Observation 636d64aa-4260-403f-b138-5acc801f305a · inbound

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs cites this paper.

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:10.775809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:10.775809Z digest=sha256:7c0a0c48b8fa972fcf23dca0384f6b96e856ccde192440c7aee5a6d9e94b2819

Observation 616a595b-8c34-490b-a32e-44a2b1fea9e9 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:18.617898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:18.617898Z digest=sha256:019443ac992c2af7a398fad41c65fb5025baf327e850afebda4db3a0e38a027d

Observation 6491f5ab-6699-46ce-a21b-57d51d5e1cac · inbound

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning cites this paper.

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:37.201287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:14:37.201287Z digest=sha256:6448defb77efe0ff6e82c3568b8e620fa61be6f62f68f88cf138bedc22cdde00

Observation 65b7ebd3-8f31-4636-ab99-73c31d9664ff · inbound

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens cites this paper.

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T17:04:38.734731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:04:38.734731Z digest=sha256:90276a8a5264a8e122166a9c437762d098ba7479f790a266d5fadda0719b6695

Observation 0887f2a1-f622-4a95-928f-fac43d84f501 · inbound

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models cites this paper.

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:53.905315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:19:53.905315Z digest=sha256:abdf00f31c2a92c5ce9706533f0c27ec5452cc69bffbb48ce6ff7adb3d46f80b

Observation dd9fa6fe-b7a2-425b-a0b9-596688c4bcad · inbound

ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute cites this paper.

ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:05.726613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:05.726613Z digest=sha256:60b992598f0e04fc33bdbba593ea450fc6d8605791bf547c51f26e8dccf33817

Observation 4636812f-dd2f-435f-be1d-691f9d98dc4b · inbound

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness cites this paper.

Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:31:41.560251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:31:28.644151Z digest=sha256:c1ac223394559ca5e4badded9d32d57371b9a2b914c2a8827a9a901815c1a989

Observation 3c72c55e-5d48-4e43-bf7b-f95054e9e5bc · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.164859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:86750f1ef5fb2c8d13870a8207cf270d6ebe7658383839f1648b74170eb72092

Observation 0c1130d8-29e0-4132-a5fb-cceb8f533b3a · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:53f4bfe14dbe56c3da55e6c0c3be355969fffe813936bfe75c67fd730b6e56fc

Observation 6698c8d1-37ab-4ce2-80f0-bd785d300b04 · inbound

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning cites this paper.

ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:42.649683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:29:34.145855Z digest=sha256:8be6f80d14702bf80ca17638aac183654f85a98efe0d07059165bf11f7ac88d6

Observation 3ff69dcd-fcf7-4372-bb43-db31ef2cab14 · inbound

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning cites this paper.

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:58:13.702378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:30:40.368494Z digest=sha256:71c81bbe9dc1a602e6203a0e539f82d2717e667a973d626631c08282d44e4c3b

Observation c1d05090-8e18-4e84-a6a4-0e01974bebef · inbound

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models cites this paper.

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:05:09.170236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:56:35.796962Z digest=sha256:4dfe4e582bb0456762f9a5bc35b72382bae14c5a4110827e0c02fc6817107787

Observation 445599e8-7336-47a9-84c1-fc24e4dd4631 · inbound

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning cites this paper.

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.333051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:14:29.531510Z digest=sha256:c753d448bfd1785f1c179c00a3e4984eaecbc91565c752f73ea3f4ae40b62a13

Observation 86f5545e-8def-4fb0-badb-00ca4cf91a46 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:08.850034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:615e9a66279efdf1b0d766157860da28c45c9e7212c5d1f82a9dfecf14d44f1c

Observation 059e7f97-c690-44a9-9d51-8eb20603a599 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.204015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:fd9c5eccca1f4c57db6a1ffe9c423d84185d05b93870c614345189c3af1859e4

Observation 63dd714d-a3d3-4cd1-967a-58432dc6f5b5 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:26.324301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:54:13.146373Z digest=sha256:33b221042a2781a3609fb0bb5e226a8c55df62717c8dc212226cab0c99d2be4c

Observation e4808b51-4c5a-4ca8-9a52-22458dfbca55 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.076988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:34:43.785312Z digest=sha256:d88d4e899b4ef464d57677ed756ff20d8925402b7dcdb01acacce7fe70c7ca36

Observation 3c12158d-10a5-4b28-bf02-1bdde217889a · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:44.414204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:f21c7ec702de2be48c89ab5954b3248ecb75dd44071c8d390b469fe9c07dec4f

Observation 17a0c6e7-55d2-4b2e-81e2-e9977bb01337 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.725804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:50a19b9fa66aa8895e4f8ea41f363e4b6b1123b6ffd8da1cbeb779a739d4fb4c

Observation 0d103120-ebce-4d84-9874-fd1eb5670917 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.667858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:d2ae1b84f7b992610b623344a47f7484fa4b348fca7de93a87be972cb31bb3a5

Observation 3b0fc51b-ff58-405e-ad0f-42a374f2bc6b · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.165526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:23818da19c0f973b525e500c01d9af61323e6f83a3725cbcd41d0b42b0a5e931

Observation da886d2f-5d1b-45f5-b92c-c7ec909f6a29 · inbound

Structured Thoughts For Improved Reasoning And Context Pruning cites this paper.

Structured Thoughts For Improved Reasoning And Context Pruning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T12:08:05.502310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T12:08:05.502310Z digest=sha256:51727d272ef408d78ce3a88252de8307f6986df37fdb5c980d394fe5ea177c9f