Pith. sign in

Paper Citation Record · LEDGER

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.28457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28457 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.372507Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a825b8a0-ae52-4f03-be07-a13e398d6c2e · outbound

This paper cites Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.240335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.240335Z digest=sha256:77288d05c73277c2d0d2585f252f207f72f96fa58f778bd0c751aa491d20f484

Observation 1866da13-62ae-419f-8316-fa105773e787 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.244491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.244491Z digest=sha256:a05b204c61e54f91281fe0f9abd00633a82f85549a4a5e6fc646aeb27975f4d4

Observation e9f97217-0f6e-4d72-b4da-786b3a6a79a2 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.247542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.247542Z digest=sha256:cbc1c1e88d11a3c13294198408bb8b05eb436d238283ae4af8f1c4183e2af238

Observation 60bf378f-42f9-4f5b-8588-235043699475 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.250815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.250815Z digest=sha256:ebf5cc52a5bdf64addc552d61cbefd0b6f7e7b189fce62a5384a2168106e2c31

Observation ba0bab8b-2f5b-4b63-a1f9-71c082b6835c · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.253654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.253654Z digest=sha256:d939ab6a7c7c767b5a0ffcafcdca513c610aaf0c2315c4c228155ea199c348dd

Observation 2dc1fbba-3563-430b-8356-e8d315eb3bd6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.256338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.256338Z digest=sha256:f7e603ae8fb6df41435a159ff34d93eec0284e4b0ef76e58d757d328a0020812

Observation 7772c9f4-63dd-41c3-879e-f713757fdd92 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.259533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.259533Z digest=sha256:5e09f20ed9e8dadcf59d4f1871bb11352cd2e08fafee29cd268ccb220b9d63bf

Observation 7307b0b4-e442-4238-bc34-fa070b144c78 · outbound

This paper cites MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.262444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.262444Z digest=sha256:d35b7ac35084c2371b836fb76b88657b9af9e2b69330e227d1790bae2ad7f398

Observation c29a4b29-bf6f-4119-bda5-724646d05362 · outbound

This paper cites Weinberger.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Weinberger

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.265469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.265469Z digest=sha256:f208f2caeeacbd918b9a21eb31a1ec427bf68c5661de92131958b8b63e2c1ec7

Observation f6597c73-e57e-46a6-8bd4-6d90be0b67a2 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.267928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.267928Z digest=sha256:e27e84a2d344d0cac61515d7555390463cf476abeb0633580eb3c88277b839b1

Observation 547217c4-d886-4f21-bce1-a19450646a4b · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.270393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.270393Z digest=sha256:d3339d36782b15108a8e5e1140573b5608393034d558ca2411e79eda9dd52d92

Observation a89943fd-80d0-4c0b-a257-ea385b6cc01d · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.273153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.273153Z digest=sha256:7e0bc8b8367555e7185a6d9afd284edcf9383016927d73638f0ea2651388a50d

Observation 99d1fe7f-29fd-43ee-af3c-a4747163b302 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.275571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.275571Z digest=sha256:c4b8f6d4e0bbee9626f809d8924bb054f8d44b6bbb9ed989782532341f46182f

Observation d5dfbd5f-c099-4e32-b559-2fb2dc834d7a · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.278008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.278008Z digest=sha256:e273d1267377f4b5b8c8def8eb4da1d484879c553d15af7a771576e4c131f642

Observation 6dd1ccbf-491f-4809-9220-d1e08955db23 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.280511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.280511Z digest=sha256:b55f79616cdc67d4ab91437b843f947580712b68fcdeb3adeade83e13b5c2e56

Observation 91b87691-aeb7-4b45-aa10-3717a8b52558 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.282877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.282877Z digest=sha256:409627f407f29210519c6d36528d48de4b91d1160123e8ac8f1b61cf2ecb251e

Observation 99e235a5-2182-4e8c-9e4e-ab0c208aa240 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.288293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.288293Z digest=sha256:c4cbee698eef31701d53a1f2e5e5564f0457d246cbf8cf91853dfc7949c0c6ab

Observation 8690b733-71d7-42d6-8f5f-f3e1c0a19c52 · outbound

This paper cites Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.293988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.293988Z digest=sha256:216898b2e046185fca93ae39e4626d56f3f778d09b0041c513adf2531a38ce6d

Observation 931698e5-3d3b-4bf7-a4f7-17ee223f0808 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.296469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.296469Z digest=sha256:95acf72fb19cb49af604bcc095163b5da405394105469b23dabb76062078f4e7

Observation 36908b5e-647a-435f-935f-0a0d4693b699 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.302878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.302878Z digest=sha256:23a9ed4e4864943c895f89bc2664f0b92b3ee83230994244b6e99fecde474997

Observation 15dc2459-51eb-4145-8da7-549ff3b58b24 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.305452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.305452Z digest=sha256:f23bcea4b80a9bf531389450eb03294746e3d5e9573647ac313171b8f779d5ab

Observation 0642e10a-d52f-46ca-8c09-f5f437884a98 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.308067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.308067Z digest=sha256:5c5b5fe9dc54972b0bc692dcaf3d4f1fbbdb2ac30391fa30bf7e94c5ed604ce9

Observation e1361378-cf90-4eeb-a66b-945b5aae723e · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.311071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.311071Z digest=sha256:3f8b3fb4a798a777e7ce1ce70d076bccff3584580e9cedd41acec615da5662a4

Observation 25a29516-b769-4f04-b7b8-94af82f3eb99 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.314225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.314225Z digest=sha256:bf63520e38153ab3c9d1090735f30356a34f0e2d8a0b08794d22cfceb1af1446

Observation b5599348-6e7e-4a5c-9cc8-0105cdfacdb6 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.320439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.320439Z digest=sha256:013a254836b960f564d22cc855190c2f2b6a4a4950533fce52c474c40937dec8

Observation 669cfb3c-f3aa-4e79-8ee6-c4d087d527ad · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.323528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.323528Z digest=sha256:99b8262a5799c5d70346bb30c2a9ec8c6247cb7e96246eef8d855aac8fa8fa4b

Observation a14ef52f-9d3d-4df4-bf9c-759e86481854 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.326262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.326262Z digest=sha256:c3426dc947c06d750c59480c5475176c125a63d7d5dc2ee0bf10e669fe445304

Observation bdca9a63-aad3-4285-995c-87118c8e3c97 · outbound

This paper cites Process Supervision of Confidence Margin for Calibrated LLM Reasoning.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Process Supervision of Confidence Margin for Calibrated LLM Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.331963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.331963Z digest=sha256:07bb45ad652532b378ad168899ef52388941384368802414187b8d781aa828ab

Observation 365bfda1-39b9-4cda-9048-a94ec618026c · outbound

This paper cites Le, Ed H.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Le, Ed H

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.334779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.334779Z digest=sha256:309439de941d07cf295dd942467de6012a90d027c0957d111765d1d16ba613a3

Observation 5beffa7d-7656-4943-b9ae-fb73df30609f · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.337255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.337255Z digest=sha256:bf104561f304027bfd8fe1f279d5d122c19d44db4c336ddd5883becef16612b8

Observation d8cb41a4-2e09-4d63-aa68-afd154732010 · outbound

This paper cites Chi, Quoc V.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Chi, Quoc V

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.339944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.339944Z digest=sha256:66043dda1701dfc6d72a47050f3fe1326bb37aba26ddbbb83031e24259d83650

Observation f15e6020-67ec-470c-be4e-24586c11e5b9 · outbound

This paper cites Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.342336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.342336Z digest=sha256:365018f7be47d2106b450909943d977e7d46399b263c9684b272ba96d20ea8cc

Observation 0e09cf6b-017e-4025-93c8-2c079fc9a959 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.345083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.345083Z digest=sha256:94074edfa8972f4a6ad5e3e3e2d7f51220963a00a906a042e2e93fc82bb6793e

Observation 2c5f91e8-3d50-4bb6-9a3a-b23ec8d8f839 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.348222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.348222Z digest=sha256:ec3b38e9876cbe054708c2e1c114679f8ce351b2c6fea774c98dfd2dd7bff419

Observation aaed60f1-e0d4-4c17-9a79-efd6bcacb137 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.351015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.351015Z digest=sha256:986dfb95b3c2276e808ab246d97b5187a7064c7a3a74dc06a843c6997ae8861f

Observation b75c8d8c-e57a-4643-b252-50faf96d5b89 · outbound

This paper cites Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.353730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.353730Z digest=sha256:918b5155f879a8417757423055c01b38a8679f5d5593e8d55dfee832e7d4e189

Observation 3b431275-1549-4c2e-bd65-bde46ca7c758 · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.356525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.356525Z digest=sha256:e8414b20d5a8bede7ccb0ee2bfd7b284f3e9a01fab8f64c67b823208846e6d48

Observation c16beda9-ad19-464a-9e24-01e347fc532c · outbound

This paper cites Group Sequence Policy Optimization.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.359063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.359063Z digest=sha256:115b373f97e9804450e554c46ab85a153d329541069e627ece75d7edb26c1801

Observation e6bb76ba-9a96-425b-baf0-a5a718f980af · outbound

This paper cites an unresolved cited work.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.361965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.361965Z digest=sha256:35800bc40bd8bf56931a6822b4ceaa627f993bfc05902d7fadd659e9e0c7ddfd

Observation 8f988685-4a34-4678-969f-324a920868e9 · outbound

This paper cites The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-07-31T06:48:41.366260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.366260Z digest=sha256:05ca8ec59d7dfa99276ee2248176ff4ce9f2f7f4968e161e7e939a118ef406d4

Observation 20ca0fa7-283b-436c-8bdf-7354f5bc85a7 · outbound

This paper cites Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-07-31T06:48:41.372507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.372507Z digest=sha256:1e5608d9ccba4832bf8e1ae018d065ea9f427457c31f2fc1361186d7f1d36428

Observation b0fb69d2-8dad-4010-a0bf-7464807a44b1 · outbound

This paper cites Language Models (Mostly) Know What They Know.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Language Models (Mostly) Know What They Know

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.285434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.285434Z digest=sha256:1855552402ee4934726e296422bee302cb09e4bcb0af7cf4b19a2b394bd722c0

Observation fd618804-d1ea-4a5a-a40e-50c5451a5c1d · outbound

This paper cites Efficient PRM Training Data Synthesis via Formal Verification.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Efficient PRM Training Data Synthesis via Formal Verification

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.290989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.290989Z digest=sha256:e3b820c1f5988ea33d30ee3cf134ce94dbf8f0a88fe2098fee0228c9074237ac

Pith citing papers

No inbound Pith citation observations are available.