Pith. sign in

Paper Citation Record · LEDGER

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

As of 17 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2602.06422.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.06422 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:58:50.730090Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T20:38:28.328436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:35:47.308416Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4823fd-3f9a-4bc0-8804-05f6c0a19a57 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Training Diffusion Models with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.478031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.478031Z digest=sha256:09bfba394efb576186b51301511f45c12182b25d5fbe3457077699ebf28c4ba8

Observation 97819a72-f58a-435f-a44d-2f374c6a8b0a · outbound

This paper cites an unresolved cited work.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.730090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.730090Z digest=sha256:bb48fc1644be0a48c5a42217081fc7a20b36a79395ab0bfda99846d8615d97f9

Observation 2e78b1ca-bb65-4826-9c13-564fee38b66d · outbound

This paper cites Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.835551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.835551Z digest=sha256:ca7f94f78169f158ccf897ea0262b8a929233f0ae066475e1dd347aefc8fcb52

Observation 90896383-cba7-480c-8ca7-4346902e1dca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.943935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.943935Z digest=sha256:8a917231f3c84e4ad46ae9429876edd937ea673f3cdf0c9d04c245aecd03e473

Observation f2b5f785-2e51-4051-a603-adbd2bd6ba1d · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.067234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.067234Z digest=sha256:5accf2233d0a4ea1e38b5e1f29134effaf4324f6e39d9c238bc6ff2116e3a7d8

Observation 5be4e67f-1567-4acd-9a94-175afaf311c8 · outbound

This paper cites an unresolved cited work.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.688606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.688606Z digest=sha256:9ef2d2624e2c4f84402e127e3e5cf140089ff295dfcfb43864942d689f4b60fb

Observation cde24e99-d1f6-4383-bea0-86b457ff8e3d · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.354792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.354792Z digest=sha256:a6d797196abf676ea0366d6167fea74c7804a12c96580b8cd55b7ef8536e6208

Observation d3bd812a-b04d-4bc6-ae04-74353ab868a8 · outbound

This paper cites Flow Matching for Generative Modeling.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Matching for Generative Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.455396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.455396Z digest=sha256:c701d4e977c36aa25f258327464bbbeaf9002abb026cadcdb9326faa1544d52e

Observation d961a994-2c86-4e5e-9bca-f7390645d529 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow-GRPO: Training Flow Matching Models via Online RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.630032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.630032Z digest=sha256:0679752498941b17d21712f2c1729aa30ccadd99f1e68fd1c881eaa4eb429f72

Observation 6e4b85ed-65c1-4eec-8794-4dff5e214a8d · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.747831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.747831Z digest=sha256:9528b2ee93d1b36dde789f536b0b2f8d57c4c9daeacbab4b1257d321d14a6bdb

Observation 639bd713-4120-44cb-adc5-b7d95bd6d50b · outbound

This paper cites Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint arXiv:2510.22319, 2025a.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Grpo-guard: Mitigating implicit over-optimization in flow matching via regulated clipping.arXiv preprint arXiv:2510.22319, 2025a

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.127534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.127534Z digest=sha256:78fabdc40e1a0d17e4bd7ca9494fb9f94b1d198c69203215881b44f778ed4f0e

Observation a50eaa08-f5fb-4fdf-bef9-62c3a9ba542a · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DanceGRPO: Unleashing GRPO on Visual Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.281427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.281427Z digest=sha256:066497be58b9b0f5dd61818ac36404170739f20c69f6af0e28038d0d65f70a39

Observation 33790b14-018f-4f69-ada6-8934bee8ee34 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.427811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.427811Z digest=sha256:21c90b484efd3ae9f9cb2b12f477fdad80a9483c27f6a640870b2492a5226589

Observation 9f97412d-913d-4f1d-a1f2-cb5ea11b10c7 · outbound

This paper cites Group Sequence Policy Optimization.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.564346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.564346Z digest=sha256:22b7e0212eb4c9316a172342c175d579cfc266fc6d105d087f2d4fe4463dfd66

Observation 5fe96f48-1dd4-4a36-9acf-bbd97a657803 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:50.055891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:50.055891Z digest=sha256:853a7247fddf95f5dbdb8a2b284666cbcce78ca021dbea16a4b91ee10ece1d32

Observation cdac7abf-3952-4199-803f-d6e647d6d7c0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Proximal Policy Optimization Algorithms

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.865910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.865910Z digest=sha256:7bc0daad7edcf25f96633de21e8bfbb4dcd505e39cb635a964df3262efd95d9c

Observation 9b7e307a-80b3-4b1a-b673-c4319ec040da · outbound

This paper cites OpenAI o1 System Card.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO OpenAI o1 System Card

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.120237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.120237Z digest=sha256:8a1ab7ea6f51b6e328c9161e860fd6bdef7cef40c88ce535e1474e4de8267612

Observation c77a963d-838a-47de-9571-b24da31e2d72 · outbound

This paper cites PaddleOCR 3.0 Technical Report.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO PaddleOCR 3.0 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.541300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.541300Z digest=sha256:1574fcb525f0658d9c71b14d4720ae5222af921fe2f990a04373a5d6e6e52dcb

Observation 777f14c1-41b3-445c-ac38-471d792bea4d · outbound

This paper cites Guiding a Diffusion Model with a Bad Version of Itself.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Guiding a Diffusion Model with a Bad Version of Itself

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:49.195604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:49.195604Z digest=sha256:c617ffcc20751545d926b200835a5d9819f0a03075f60edbc1f3fc62b0e52058

Observation 89b3d90f-74de-4370-8b2b-36f1f8542440 · outbound

This paper cites Densegrpo: From sparse to dense re- ward for flow matching model alignment.arXiv preprint arXiv:2601.20218,.

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO Densegrpo: From sparse to dense re- ward for flow matching model alignment.arXiv preprint arXiv:2601.20218,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:58:48.719811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:58:48.719811Z digest=sha256:6e4aa7f1e09eb0a9a6ac0ac939ec9e192dcd0ebc036e059f58932e1598b18a4c

Pith citing papers

Observation 72c1ca81-5a6e-46b9-912c-239b718f6730 · inbound

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO cites this paper.

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:12.085101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T20:38:28.328436Z digest=sha256:b743e477dc52e4f66a46125878dae39bc8fc5a7ecf0ab5696b247868dd89c017