Pith. sign in

Paper Citation Record · LEDGER

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.27787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27787 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:28:57.123789Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.596043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.596043Z digest=sha256:6a153cdbfe09ede09ad550b36832a80233d93e2bfdd692da69676dce3d32cee7

Observation 1e50381b-8495-4fee-8f2e-e3fd4db4bd9d · outbound

This paper cites arXiv preprint arXiv:2506.07527 , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2506.07527 , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.668707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.668707Z digest=sha256:6de3d95f3598d696711caa01f45a02c07d284f22f91c1aa7e117d30d26877747

Observation 5db1f480-06ca-44b2-9dda-d0bf5756f576 · outbound

This paper cites Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.739458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.739458Z digest=sha256:5f19b140d67905e610f513995c2fd7faf9273edc976e410a31f482ac12573f88

Observation ed2134da-3e3b-48b4-a765-437228c091dd · outbound

This paper cites arXiv preprint arXiv:2603.23871 , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2603.23871 , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.804128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.804128Z digest=sha256:3f4894556164b25c253fee02cfd7a41dc0758a821001d6c9a9153f72da9f2270

Observation 3dfaeff7-2e76-4d4a-aa44-f287a72a3801 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.868005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.868005Z digest=sha256:c7d6f35052534197b7cd70ac0b976a9aa61bb33d0214fcbbfb83aac6f73a9de4

Observation 8d577706-0b84-40fa-a279-107c03bb98fe · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:55.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:55.933165Z digest=sha256:310089b9c3bd9d28548885c66e221466610e4eadbae85a918f9e101ea60067bc

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.048240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.048240Z digest=sha256:1f2cc18420235496a9599eb35c1b1d2ecc451a1bddd61a436b6e4e0bfac9d21b

Observation 9a609134-d9db-4a73-af55-5bd7091cdc5c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts LoRA: Low-Rank Adaptation of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.118443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.118443Z digest=sha256:d455eb20045ba2fdc29bcf477b1a906a49dea3a542ab73c874d883603f076021

Observation 86296fed-8a99-430f-88c2-a5cd2006e074 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.176653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.176653Z digest=sha256:ea6a38fda4e68d143ac47572ca12d4e130cd563fccf32689ecd86d16e7120de0

Observation 72fd49f8-e1e1-40ae-961e-20159b36095b · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.250910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.250910Z digest=sha256:556f1b82d42bb51f046b62b4d85cad9d8b962fc618985a41c75871398a3b922e

Observation aab03436-8a35-4566-ac31-d16a11b3dfbd · outbound

This paper cites Neural Information Processing Systems , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Neural Information Processing Systems , year =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.314615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.314615Z digest=sha256:acf734c1b63bea15451b659aa943dfd9bf41ed9dd4ba39f4d744c010738cc205

Observation b2ab55e9-f415-4fd1-b962-59834df789e2 · outbound

This paper cites International Conference on Learning Representations , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Learning Representations , year =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.378939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.378939Z digest=sha256:2bbe3c0ef636b26dcf4f531d8010dac6a8ceb72c2558ac3c862fd2b6ebbf5e6d

Observation 6ec18f94-a102-46f5-b48b-f740d5d70b10 · outbound

This paper cites International Conference on Learning Representations , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Learning Representations , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.405734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.405734Z digest=sha256:6c005301fccba4d14477f013509d323b0839a9945d591d88d215fde04abfd615

Observation 15a322e6-43e8-4c23-82bf-bb4f8775ecf0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Evaluating Large Language Models Trained on Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.470597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.470597Z digest=sha256:e8c1433343fdff33216bbf8dc83bd2662d1c887a95376928a07c52e2b6687392

Observation 4495b5d9-cd5e-4933-b98b-7ca4606f0f15 · outbound

This paper cites Understanding.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.627338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.627338Z digest=sha256:27e74fb65701f9734d9d0b7ff0432ba5b22dc79b2ea3901d15eb8ebd64a433d2

Observation 994c961a-99d7-4a9e-8378-21d7cd9ceb46 · outbound

This paper cites Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.693276Z digest=sha256:d7f90c9549b58533a91054dd8f1a91d5027674a34cbe5e6d5d650f786873e64e

Observation 4ec73280-de7c-4484-9790-8ab6114ce01c · outbound

This paper cites 2026 , note =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2026 , note =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.748724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.748724Z digest=sha256:451709ec23321397c082b6f80ed143d93fc547a8f612254f25b6ab82a927694e

Observation a28dfa1e-ec7f-4daf-9161-39c79fab2afe · outbound

This paper cites arXiv preprint arXiv:2602.03143 , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2602.03143 , year =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.816872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.816872Z digest=sha256:000fc1e11ad1a53f2ad4b085d7efb8af5cd848c75942f4856f8204ae92897950

Observation 44774df2-16d2-449b-ad90-d338b1fb4e2f · outbound

This paper cites arXiv preprint arXiv:2604.00698 , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2604.00698 , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.878791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.878791Z digest=sha256:db451bf7353ccecc8ca0418bf61ee8f28b3e900f5015d53d17ede19b97782a0f

Observation 36266788-8891-4384-989c-d055085520ef · outbound

This paper cites Nudging the Boundaries of.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Nudging the Boundaries of

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.929960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.929960Z digest=sha256:54fa395cb7233e9ca8f84d102a83976842f7c383706e13bcae676d1bf5a25326

Observation 96aa2ad7-4baf-4a65-8b3c-8389cd68567b · outbound

This paper cites 2025 , url =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2025 , url =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.003113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.003113Z digest=sha256:be87df102b23bb2d02d0d0ea49f5926f75b80bca7fdd137d4356e03532e88953

Observation e44efb7f-7866-4efa-bc4f-8b03875d0489 · outbound

This paper cites Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.015362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.015362Z digest=sha256:24540d5b496b8a8ddf7f13b456bb5490de574dea69a33d10fa014981b7b39006

Observation f8002cd4-72a9-4660-8c2a-48c472e5bda0 · outbound

This paper cites 2025 , note =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2025 , note =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.021388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.021388Z digest=sha256:0d1d8a97281187fc7309566884304495970abc584b3c39671fb20d8405b1b498

Observation febe30bb-ebf8-499c-800f-b25b943b293e · outbound

This paper cites 2026 , url =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2026 , url =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.026690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.026690Z digest=sha256:eb25170cf1177a8bf5f53e0a716ac55b99ad2e98a88d16942d60b66f80d3393d

Observation 57629963-d642-40e3-bd50-702dc651009b · outbound

This paper cites and Jeon, Myeongho and Vu, Kim and Lai, Viet and Yang, Eunho , booktitle =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts and Jeon, Myeongho and Vu, Kim and Lai, Viet and Yang, Eunho , booktitle =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.032975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.032975Z digest=sha256:b4e16d7c8aaef0d4a6459d9254c294644434d2e59cab9e4a23a17fb46316bd9f

Observation 2383483f-77f7-4825-baa7-7a6489d929fa · outbound

This paper cites Don't Waste Mistakes: Leveraging Negative.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Don't Waste Mistakes: Leveraging Negative

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.038231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.038231Z digest=sha256:b7fae18913f9bff1cfe44ecf88369934c4a0426553f8723e0742cb45ba3625dd

Observation 8371c59e-caf1-4dd7-8099-7f61e6cc8e03 · outbound

This paper cites 2026 , url =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2026 , url =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.043917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.043917Z digest=sha256:f85302eedb74feedd3bf1bd849bd7eccf568fcdab069347c00af2a278cd2f1c6

Observation 867c3a41-5dea-48b4-90af-b8473fb3ac32 · outbound

This paper cites Learning-Zone Energy: Online Data Selection for Efficient.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Learning-Zone Energy: Online Data Selection for Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.049520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.049520Z digest=sha256:c1fbb0cfb2db063057e7c0db39d820465e3763fd1ffb1f134fce0b68b455182f

Observation 4f2ddb01-de73-45dc-9642-0555e779ac66 · outbound

This paper cites Advances in Neural Information Processing Systems (.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Advances in Neural Information Processing Systems (

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.055572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.055572Z digest=sha256:06343e78eaabc0c438d18d3f259a7296705adb36611cc0c4a7f4bd35b367e834

Observation b4b379e5-4de0-40a8-8c4a-8349696534d3 · outbound

This paper cites , booktitle =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts , booktitle =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.060474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.060474Z digest=sha256:5f09fb091d364cf1dc60ede8bf3560975f0147052bf91f43be0d3d700fdd49da

Observation 1d9fe803-993c-41df-8b11-df54279391a9 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.065296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.065296Z digest=sha256:327329a2a8999965a946bd54c4a5e3228e8d778f3e7f9814dec4f47130530a7f

Observation 4d3af5fb-81d3-4b49-8cb9-37789a57780c · outbound

This paper cites Reinforced Self-Training (.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Reinforced Self-Training (

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.071439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.071439Z digest=sha256:d8195fb230ee409b3be7c407f1ec69e26ed189877a4ef7ff842146607658190a

Observation 08879f88-e5ab-4680-8ffd-5309f37c4f48 · outbound

This paper cites Transactions on Machine Learning Research (.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Transactions on Machine Learning Research (

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.076222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.076222Z digest=sha256:86922ec035c2dc54c5e387a130279e18629aabe06b5ce5c308c1eaa0859b2c87

Observation ce414c73-cf32-4c85-863a-2c5b02663067 · outbound

This paper cites Learn Hard Problems During.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Learn Hard Problems During

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.081170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.081170Z digest=sha256:f4ab537a624a6c8b81b288209e6d24080938511810626a549bc123102cf9257b

Observation e738eace-afee-47af-b3bc-b827e9956532 · outbound

This paper cites Robotics: Science and Systems (.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Robotics: Science and Systems (

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.086165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.086165Z digest=sha256:584d05f2f3716a620331606cedbcaeb599037a930c69a16ccc9028f5483573f5

Observation 08a5c976-f6cc-4178-aab0-586be606dc7a · outbound

This paper cites Advances in Neural Information Processing Systems (.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Advances in Neural Information Processing Systems (

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.091777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.091777Z digest=sha256:4ea56a3ec61fb39dfe55e93b2ac512e43532d3e9a01d98c4dd6ac6ce812d8cf2

Observation b7f6980c-6371-495c-8c23-170743bfbe5d · outbound

This paper cites 2024 , note =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts 2024 , note =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.098425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.098425Z digest=sha256:86e66cfcff84a907d5f5c37c0296fd109c51ec14af98458411a859aef0a7436d

Observation b5355ae5-d164-4197-8444-a5fd2b447f52 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Process Reinforcement through Implicit Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.103793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.103793Z digest=sha256:30065d1845339f3addb89493a776c6c4499574e78ed218020725e7bcf6b0ba89

Observation 1c9b497d-6b11-4188-b807-6a05c317ccd3 · outbound

This paper cites Efficient.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Efficient

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.109916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.109916Z digest=sha256:62bc7d7365c3e32e5f8796d5f4cde9c884add94506859d0c58b429cf33e2a426

Observation a0005115-698a-4cb8-bfa5-8e2ecc2e6567 · outbound

This paper cites International Conference on Machine Learning (ICML) , year =.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Machine Learning (ICML) , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:57.123789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:57.123789Z digest=sha256:79ad10ca241b757b82b720ef1d7d2ea282051d495054ed5669a1a8d6103eefbd

Pith citing papers

No inbound Pith citation observations are available.