Pith. sign in

Paper Citation Record · LEDGER

Scaling Test-Time Compute Without Verification or RL is Suboptimal

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2502.12118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12118 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:44.275296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.081466Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8bbf0e62-8788-46db-b714-22f2cd90189d · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.238636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:cde08505f593465380c2e414444200a924852233f75cd76c3be1f5017b068f1c

Observation 58108822-4319-4b0e-9fc6-2ea880285d84 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.540276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:4758ca2b68f9a9e7ed190560a2ef800ebae75ce4b6634a7ace9b1cc261aba911

Observation d61edaaa-bdcf-42ae-ab77-16a2fcdc166a · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.485438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:ec5d4c548674ddc617c1908fd912d0b92a56932335cb503b729dae3065b3c931

Observation 38e18a3b-642c-459f-8caa-ccc9876c8e26 · inbound

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection cites this paper.

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:44.275296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:44.275296Z digest=sha256:d7cb14f5db88421bf3148c4fad49e37e1a26233caf1de0d815a578e2f12109da

Observation 00262278-648d-401e-83b7-43a01aac0f27 · inbound

Faster and Better LLMs via Latency-Aware Test-Time Scaling cites this paper.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:22.796446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:22.796446Z digest=sha256:8c098e06ef2031e4b1d39b33a8f7e2c0fc8dc3b662bb5a8191faa81f93d46322

Observation 0ca12fc2-9a90-4756-9d71-03f3a2ca862b · inbound

Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models cites this paper.

Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:58:36.660912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:58:36.660912Z digest=sha256:83c9f7f570d9fc82b657399276040eb172fcc5b3d653a6fa1cade82ca82880da

Observation 4fd36060-3c10-44c7-b923-e4aacd158829 · inbound

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations cites this paper.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:47:17.990957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T12:43:51.019983Z digest=sha256:77c3ca41ee9e3731ecd770bf457104e3de2be074e25b2e979c0cb94687e90667

Observation 55ac6afc-efc0-4230-bf21-ea854c17561c · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.177987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.177987Z digest=sha256:192b97f4bc9be75a2f5b93419a17dcf33ad3695764bd4934fb6b97dd05e7b8f4

Observation c5cb6b72-cbe8-4b35-8f4d-b19e3acc3ea8 · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.015386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:41.015386Z digest=sha256:0f69685ea538332a175cc70a7f4be643a41d5d0fc784dc4d4ebc7bcf2b9f4360

Observation e0e4fd3c-792e-4adf-918e-215295894e38 · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:37.965006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:37.965006Z digest=sha256:678f2297cb67a33cbcd0abc64b859ae798529dd435c1a139ce04ccf69f7937e1

Observation f20f4c7f-9d7b-48aa-9f8d-bad1bab99af3 · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.219788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.219788Z digest=sha256:0c2481fda915c7e95b2e5226a150801ad671907d6b0989f55606947862324ea3

Observation d155d0d9-b405-46c5-9231-8f57d771e0eb · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.653994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.653994Z digest=sha256:452517ae4d9eb5b632bbf9a6b20a6e063b8b06f4d59aba6b4a570fcae300cb22

Observation c3185151-7874-4f06-bdd5-076ba5a0224f · inbound

Risk-Guided Diffusion: Toward Deploying Robot Foundation Models in Space, Where Failure Is Not An Option cites this paper.

Risk-Guided Diffusion: Toward Deploying Robot Foundation Models in Space, Where Failure Is Not An Option Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:10.024727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:10.024727Z digest=sha256:e7ace80f96419cb59adb76c145db103cc72aa5e6c9e9863aceecdf6869536d73

Observation db2f22e1-3716-4ea5-b409-8dafbf119d55 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:18.663876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:18.663876Z digest=sha256:77de4a00b4cbcb0fba67bf86a0584d0dec2f4d055ddafb2be4c17bd50ea1b55a

Observation cd3a6297-62fc-40ba-8849-979f98675793 · inbound

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online cites this paper.

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:50.698310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:52:50.698310Z digest=sha256:aaed26c0e152b5dc61399eb0579fe26b38bc3a8b722776d7c4306786a04b59bc

Observation c3121a39-00b2-46a5-a1ab-654f921f914f · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.862981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.862981Z digest=sha256:97898055cc91b3f9ec747316746adf1b38720e9970cb6333ae44839b7f8f5434

Observation 7a0040a0-e9bb-426a-964a-23a9867f6fc7 · inbound

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction cites this paper.

RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:59.185386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:32:59.185386Z digest=sha256:6ab72d142c0ed6440da2910b0740c9bee34623fe9af1ab47d5de71b7bc61810c

Observation 3cd3c06e-92ef-46b1-b172-6c3976ea4406 · inbound

Asking LLMs to Verify First is Almost Free Lunch cites this paper.

Asking LLMs to Verify First is Almost Free Lunch Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T21:03:18.958441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:03:18.958441Z digest=sha256:0f859e5473372e63aa88d595c9553b2900576ab8c944b3554ebbad044d480b24

Observation 93315f33-d916-44be-ae74-e752e30de6b1 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.910581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:dff4e402dd4e94412874573bc23e4462dba7f8edfdf00c48945bcd02f5817ef5

Observation 7eaa64f9-baa4-4f34-8c4d-6dd271c40469 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:05.777615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:32ec65b896fae7a292c916192334c412ad75f1c74d5be9401d357443bf7389c6

Observation deee3ea0-b457-41b0-a510-eddef6dc0fa7 · inbound

CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning cites this paper.

CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:42:38.364806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T15:39:56.255871Z digest=sha256:957a330d0ca4942fdf267cdced091ab38b163b0107b77f80ef969ae6f04fbc94

Observation 0c548352-46e2-484a-bb5d-e726d873864a · inbound

A Predictive Law for On-Policy Self-Distillation From World Feedback cites this paper.

A Predictive Law for On-Policy Self-Distillation From World Feedback Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:13:16.210460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T09:08:16.307499Z digest=sha256:b1b07e98d8faa24846ea8db9e1c07f7800aaf658901224ea54797ef0d29ab6df

Observation 959bee1d-0987-4be4-adfc-c3212bae2d98 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.924302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:c6dc726fb38d4ca3d7dd5f0849ac9ab09254d9dabdb3439f105b9e6ee0847cb4

Observation 49abea7f-0bf0-4190-a9ec-c0beb200826f · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.082787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:828c5e128097392dd82c13794e02da90ec78ee10d21c9d9577c39136b5a9cc27

Observation ec8e1e04-02ca-48cf-8566-aa37402a0fb1 · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:d07beadc1e3ede09fb01d60967f23075b1eaf5071bd18bdad18a13fd2dbe6d49

Observation 86aa1bbb-a3b7-404f-b30e-bee62b08caac · inbound

Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration cites this paper.

Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:46:56.344831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:46:56.344831Z digest=sha256:98424dba2d7fa16332969d60f3e0c4c81989dc2d27c5883188259cd353bf4379