Pith. sign in

Paper Citation Record · LEDGER

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

As of 16 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 14 inbound Pith citation observations for arXiv:2506.17211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17211 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:03.097264Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.173572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.036901Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b615c96-ca91-46e9-918f-b020b315d4df · outbound

This paper cites Phi-4 Technical Report.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.866049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.866049Z digest=sha256:7d4b0fc6e98e06bd4214f70b755e406f0e8d25ab195c38f9c92109e3ead2cf1b

Observation bff033ec-fa38-4fdf-be7e-a22e310e3686 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.871865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.871865Z digest=sha256:565cbcf08d739e6dc3217460bf71379041e3475b0ec7f64fdf05ab43b35c4963

Observation 3ec446cf-31e8-43e2-9d83-7930ac67c78b · outbound

This paper cites Hindsight experience replay.Advances in neural information processing systems, 30, 2017.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Hindsight experience replay.Advances in neural information processing systems, 30, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.876897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.876897Z digest=sha256:a1e041a95ec51da6e7d7185ed7fb3005cd309f79800923d58e127ec8da1b0be8

Observation 40101499-a613-40ac-80cf-f556f124128c · outbound

This paper cites Thinking fast and slow with deep learning and tree search.Advances in neural information processing systems, 30, 2017.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Thinking fast and slow with deep learning and tree search.Advances in neural information processing systems, 30, 2017

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:04.072720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:02.881939Z digest=sha256:2473e200f58a538cac201c233776fc3651a3e6505c88ea5df6e98be633380878

Observation 53069f5d-a754-4445-a10c-53637e29997d · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.886610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.886610Z digest=sha256:3228082c0217b3654c747e09d12418a235b0b4ab7ffdf2c4f8f5f154e015ccb2

Observation 90e7283d-54d3-4a8e-a0c5-0a7f978140b2 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.891209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.891209Z digest=sha256:ee755f24a63a966cca19a1fd1be3c7902aa9c67ec013d3f68fbd36ece64172c4

Observation 46184443-4e7b-44fa-b5d6-6cb689cf7460 · outbound

This paper cites Cambridge university press, 2019.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Cambridge university press, 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.896490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.896490Z digest=sha256:c0da9df7e117a769647524eea8f951e5500612687aa967d4627c082b407ffe90

Observation b16e010d-9345-4da8-8280-419adc28b9b7 · outbound

This paper cites First return, then explore.Nature, 590(7847):580–586, 2021.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning First return, then explore.Nature, 590(7847):580–586, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.900845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.900845Z digest=sha256:c40fe00fd140e8c631cc0e1286edfac519436a4a912f2d01476ed1d89d359016

Observation bb6da3df-2c18-4ef9-b9d3-fb78a3111f55 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.905412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.905412Z digest=sha256:0c7b129331db7f48d9cc4d4e3328fecc3490f650626bf4b8e221c483d853ce67

Observation 95f2eeb4-fba4-4c28-b47d-64dbf3543f6e · outbound

This paper cites John Wiley & Sons, 1991.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning John Wiley & Sons, 1991

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.909702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.909702Z digest=sha256:2bf69aaf963986353934560a61ec78f9afc364f07f5d25bbe7c0c4a0101c88d9

Observation 5dffc416-524d-4e91-abcb-c34974ad2785 · outbound

This paper cites Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:04.016087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:02.913842Z digest=sha256:e1f6f4374cac7a83ea72adb80263fa7bb6d44d201c1fa8ebc47114ad6fcc4884

Observation c63d430d-3448-4853-9228-fb9d71aa8f35 · outbound

This paper cites Last updated 14 May 2025.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Last updated 14 May 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:04.000613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:02.918543Z digest=sha256:7142b8f9576392e605dc31f94dffc42177b9f85d88241e4ce56ce01eb7d65206

Observation 60bd52a3-f3e7-4bf1-b07b-3ae0b83e56ce · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.923811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.923811Z digest=sha256:a6c2da695b0673c4bedcbda5e71d9128080e7d49ff85f06cb54e3d8a60dcc214

Observation 887f31d0-d506-47fd-872b-fb9905231a5a · outbound

This paper cites Language model cascades: Token-level uncertainty and beyond.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Language model cascades: Token-level uncertainty and beyond

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:03.981982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:02.929119Z digest=sha256:20df1629ca50dd60c80153e44ad517d4fd254489808f48f823028c0c42c6b81d

Observation 8b12159c-15aa-4e5f-9e03-c6baea6a4a4a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.934244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.934244Z digest=sha256:a6df1a0a36a5cef5cc19dad5f2c19d64c9c764a9da62e1edd474a926a696b83b

Observation 64f3349c-3e61-4df2-9aab-ae5b96d69797 · outbound

This paper cites Training Compute-Optimal Large Language Models.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Training Compute-Optimal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.939368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.939368Z digest=sha256:74f856c0df2c0b7e6f42e09ee517cf25220c44d4daaeb027cae788cd56718723

Observation 593e18c5-091f-4cbc-bda4-82ccc16f7fdc · outbound

This paper cites OpenAI o1 System Card.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.944860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.944860Z digest=sha256:ad57d68bdef3cb37102915e2e6dcca0597aa839038cb4413c0f5f0b7ff31caa5

Observation 1ecc5a44-9a45-40c7-8ef7-f3a14d1e6ffd · outbound

This paper cites When to trust your model: Model-based policy optimization.Advances in neural information processing systems, 32, 2019.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning When to trust your model: Model-based policy optimization.Advances in neural information processing systems, 32, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.949516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.949516Z digest=sha256:e631be4b48b78bcebe087d7f8e8bd7850ffc5087f2a490619ad0fb7f31dbfcb7

Observation db10e8a0-6ee5-4b4d-9b18-f884620abb2d · outbound

This paper cites Gemini 2.5: Our most intelligent ai model.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gemini 2.5: Our most intelligent ai model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:03.955605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:02.953808Z digest=sha256:e9315acbc1779a203767f6b3442d173dd029dc9e827779a5c03386b83f14fe2e

Observation 561abd4e-f2d3-4c36-ba0e-16811971c35b · outbound

This paper cites Adam: A Method for Stochastic Optimization.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.959368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.959368Z digest=sha256:a716750e4c370b1bff553f29f9735e82201dd6ce7df15f25352035516def27bb

Observation b8ebe66f-f197-404e-81f2-68354b35d613 · outbound

This paper cites Numinamath.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Numinamath

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.963817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.963817Z digest=sha256:39f56c0f55b1822131ede69eb332bf590df8b30e75305f82a998948d310601b2

Observation 8a63cfc0-16ac-42ff-a5f1-6966c7208a85 · outbound

This paper cites Small models struggle to learn from strong reasoners.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Small models struggle to learn from strong reasoners

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.968436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.968436Z digest=sha256:5d906aea10bc626cd8d2bb3570f62ce72a7003e3c9c92b3c54e85ef09a0f69d8

Observation 31d7cf98-6696-4645-9e55-af9e5cc7b8fa · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.974542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.974542Z digest=sha256:c4f1a714a2d7e9cedfb60889b86125d81f0d8a09776eeb3ed2cc9a6306d4cbb0

Observation 750e1d5a-a226-47c7-9ccd-5fb331ba6c16 · outbound

This paper cites s1: Simple test-time scaling.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning s1: Simple test-time scaling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.979296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.979296Z digest=sha256:1710ead9bc0b8215a0e9533a3770c490ff0431b78ca71701c8ab9cb05b47cf6f

Observation 79ecc836-e87b-44cc-95b4-944d64316795 · outbound

This paper cites Policy invariance under reward transforma- tions: Theory and application to reward shaping.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Policy invariance under reward transforma- tions: Theory and application to reward shaping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.983665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.983665Z digest=sha256:56ead4b9b3c4e154976a6e9b1de44590210bd7f53cc92b5296cf63aaa69b7cf2

Observation e71d2799-3c7c-4413-bd60-515ceabb671f · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.987946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.987946Z digest=sha256:f737810b5bc5796cbf0dec4bd56c10e28d268af34a9de9f2a6fb5ab3b655d6f3

Observation 90f601ef-e550-4ff5-9a70-5cd29c1831fc · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.993576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.993576Z digest=sha256:d4f6415e526aeaca4114e6321c2cff3d19a871710358b3de1874b584a2ad9264

Observation 0b12fe50-36a1-4ed4-a833-24583098b27d · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Gpqa: A graduate-level google-proof q&a benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:02.998282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:02.998282Z digest=sha256:89deb5bd6e1c0705701c19e4b4c5e74a30cca6ee080aa81a97d89a728c342e8f

Observation bfce9b3b-e663-4831-87a3-f3aa9ae05b31 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning A reduction of imitation learning and structured prediction to no-regret online learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.002579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.002579Z digest=sha256:9bb813d41dc16631aa2a8304e407ba42d5fba9176637da5520d6ad30c894829d

Observation 5a39a7a7-8e3e-4c8f-a3e2-68d6bea889d8 · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.007339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.007339Z digest=sha256:a22e03086155071a9de762d92f91ae3507b1d4ba17e0bbff1cc15f8dc9f52a8d

Observation 9ac2be1a-098e-4b41-9f47-060728cf6955 · outbound

This paper cites Reasoning with Latent Thoughts: On the Power of Looped Transformers.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Reasoning with Latent Thoughts: On the Power of Looped Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.012015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.012015Z digest=sha256:c460ad4183bc3418f14d34946afdcb50b054908ca6ab6de06fd30cc74b49be71

Observation d093823e-5fac-45d9-906e-229a3be8aed3 · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Kickstarting Deep Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.017958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.017958Z digest=sha256:e6601c08f8265250a7539b8cc802ebdf2554def810d5fe623856b92eded9fafb

Observation 2e4a7f8f-8989-4f57-8d08-2f8206a9bb91 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.024037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.024037Z digest=sha256:0d26714c796a5a28b61bbdd22265a4b17b83ed989d4498dd16ddde3353b0c844

Observation f6081e96-a962-48bf-a4ab-1701783111e6 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.029744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.029744Z digest=sha256:1a92e47a2cfac89e2b49e6a1fb5a96a703680af23b2a909e6f71254315afba7e

Observation ebe15589-435c-4504-918b-7ce1d9456038 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.034557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.034557Z digest=sha256:3cee5283c6087b53bc79970627a944c49ea19677ab8c03f8f42af9acbcd1e098

Observation b70788f5-3f86-4eb3-9235-0ec5698f3ca6 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.039661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.039661Z digest=sha256:a7a8362c201911216ef6842982e055ce829fe400d5428e2a421994f9e326ea6c

Observation 10649a21-2386-426a-b0e4-a1094e740573 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.044142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.044142Z digest=sha256:787703a80b4f168765566cabd0a03286024faf97f75f04a7dbd4282d7007579f

Observation 5469da06-763c-47c3-8758-f199384891e6 · outbound

This paper cites Dump: Automated distribution- level curriculum learning for rl-based llm post-training.arXiv preprint arXiv:2504.09710, 2025.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Dump: Automated distribution- level curriculum learning for rl-based llm post-training.arXiv preprint arXiv:2504.09710, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.049112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.049112Z digest=sha256:43e5722148a66a7924ed6b903f76a73185162ba8f75866a995fb722fe94aa27e

Observation 0876de9d-3300-4349-b82f-43b06698d209 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.058965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.058965Z digest=sha256:b56ef0b0261bb3646f8b03dfe4be514b2d023dfab0241b970b78a62254d3c105

Observation 18c17d81-78aa-4fc5-9c1d-c31bce1bb703 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.063494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.063494Z digest=sha256:e65f58ccddd0c431c9c6711284b99731e4b7293a9c55e78424dc3a3c94dc3194

Observation f67e7e05-8d2a-4de4-942c-4630b9cd4e06 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.068079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.068079Z digest=sha256:a2e44ff27605a03612a2fcb065631249c08f74692da442c9ba10d5ebdd0d49cf

Observation 1f9b2c27-53a9-4594-988e-304daa042a40 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.072738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.072738Z digest=sha256:67976856ae2154c72730fafcb161c80f04e11faf3a9e3bf9ec8ac1c1c7ba1d3e

Observation 6820098a-2d10-4b21-ae17-00996f34dbae · outbound

This paper cites Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.077573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.077573Z digest=sha256:bf63f4308a55be644a8a5776ee714ec7f3f11f6b6d7b22ba040e29a2e71f8dc5

Observation f0536afb-af6a-499c-8e9f-d05b9d33ac9d · outbound

This paper cites ” or“\n.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning ” or“\n

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:16:03.880067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:03.082278Z digest=sha256:4fa1ff0550ef1bb418f300e9e8de9258c45f809add85981836da8738339b8fbd

Observation c520faaf-ed47-4b73-bd70-c212a9150b9e · outbound

This paper cites an unresolved cited work.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:16:03.861787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:03.087736Z digest=sha256:bb4fd67768aaf3910581f0bbc4827ee4af6fc17bf46f42fa2b1ecb32c943b203

Observation 163ff7ed-6a39-4c59-a15b-3d761b97bdcb · outbound

This paper cites an unresolved cited work.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:16:03.844225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:03.092614Z digest=sha256:3d791b7407c3bd3b745407afd87bac03092d6e2cff308ccd3d80a7f1f38bee8b

Observation 46d42dce-87aa-4f7d-9e79-b9470f81b9ae · outbound

This paper cites an unresolved cited work.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:16:03.826514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:16:03.097264Z digest=sha256:5c4288a1aa0f4878e19130e250b90936750270cf30562ae90d07bbbb01a62b4c

Pith citing papers

Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · inbound

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models cites this paper.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.173572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.173572Z digest=sha256:e05e73bf825217755c23ef3a673e477f8f300011ee13f40682d629baed8bdded

Observation ce6002ff-3b9b-4330-81cc-1e7f8f8433b4 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:29.735455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:29.735455Z digest=sha256:99efc38ed5f5563755f54f514cc9c0d0b3f4695a8a30293d8e99dbbfac33f98a

Observation 5f7d246d-fdc6-4496-bda4-305a41ec8b45 · inbound

CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning cites this paper.

CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:28:24.110253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T20:25:35.573805Z digest=sha256:d36ed06328616a21ee61e867f9d56b5628c00eb10585d4cc58f45bdca0762b53

Observation 1f90e798-f156-4265-bbf2-6bb762263edc · inbound

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment cites this paper.

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:27:52.342148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T12:23:56.318846Z digest=sha256:e0530e4d78a40404caa53355b2e98385bf6c0f5fe75a732abdd5c0ebe7c50695

Observation cfa911e6-f72e-4014-9b58-18e0a0cdc9d4 · inbound

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment cites this paper.

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T09:22:48.248328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:22:48.248328Z digest=sha256:b816930317c27a8a152d547e4b7b6527ad22f563c5fd851e2ab6d837d8af41e0

Observation ac947f1f-e450-4296-bbf3-914c59795de6 · inbound

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning cites this paper.

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:27:44.472349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T10:26:45.961575Z digest=sha256:7370dd866d33fbc40412300c3a2e9d6459464b95a06eef89efd6cac433e2111a

Observation 4868544b-4871-43fe-87d0-83e5fc4872bf · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 176

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:49.398090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:491b7698d64c7d1b38221b181674acc7755b5a01357b1067e278dbb1891a2b66

Observation c3895e82-69c5-49c9-b926-1d6b03232eeb · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.454864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:c8db1fcb03866f45b09efe70609fdb748316c5df336d1a9aa1bb562704200b97

Observation 75445301-2356-42f7-b965-df56af58a95b · inbound

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR cites this paper.

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.388644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T19:18:52.845177Z digest=sha256:5fe50ec8edfae5e0b475f35365411a472015e3b7e93a07abbe359298648585ec

Observation 6a6d3607-e3c6-4c8e-bcca-e0ce1ebe180d · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.194931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:9bdfeb335a92d3fb97f5dc45b8a2b7f22f4f10ce60c4650c5af22b008a601427

Observation 8a73fa41-b0de-4b66-9516-1b598333ed61 · inbound

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning cites this paper.

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:57.038563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T00:39:29.703711Z digest=sha256:c12e8311bc3fe6e45d28b0e91d41dc8f5495190af05290d262fcd6254c3e1fe9

Observation f74b1d37-d4a9-4f65-a79e-667f30daddb4 · inbound

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches cites this paper.

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:21.888557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:45:21.888557Z digest=sha256:33d59851bc85ca12a75e4336b4610843235916ef2c3e77a12adf0cc675262516

Observation 4499d63a-74c1-46f9-9299-604148447797 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:36.704591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:36.704591Z digest=sha256:821372adaa17bf22f7bd0eca99f512e485369567bb0e36b84c74433f10b28992

Observation cabe3706-fffd-4c42-8540-0b7c87b376a6 · inbound

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding cites this paper.

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T21:13:06.807478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:13:06.807478Z digest=sha256:f9693f0eb1e47dcbcf44c85ec3443cb9f6e8d12fd4d3d0fa38e2ff00996c7c5a