Pith. sign in

Paper Citation Record · LEDGER

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05600 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:32.517155Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74a6c8e8-bd8b-4f2a-982b-e23e62c00a24 · outbound

This paper cites gummy bear zone.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction gummy bear zone

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.556213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:54:32.517155Z digest=sha256:78104be6df2a089795a2e4d91de6d95d191b101a4d3cc7cddd7e730325e3c079

Observation 78cdb6ac-7061-47d1-9ffd-515270f00c1b · outbound

This paper cites Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.442334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.442334Z digest=sha256:a8b2c9856637026f7c7e0b1048da071480c6c947fb47391770a9b4f1cc6dfc15

Observation 2fbb39f5-cce6-4cdc-a516-4e2205c87a32 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Aligning Text-to-Image Models using Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.463629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.463629Z digest=sha256:6e824fc37a13579cac7878f4e956dc6b83d45b713e12082238eec08072d3543f

Observation 450e6ffc-4d27-4f86-a1e8-d3b8f04101f0 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.466927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.466927Z digest=sha256:87a667b712f3cad76bd3cef9de11295e67ef66cc66a89b42a3d29880c2374d38

Observation 39601ea0-7e1a-4235-846f-40869afebf2c · outbound

This paper cites The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:33.179122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:54:32.470552Z digest=sha256:45c01bc159f0e8cf5ab09159ec3ddee3863b59940e171c56211fb9c01f9ec26f

Observation d6aae682-450c-4d2e-ba33-86b016b218f8 · outbound

This paper cites Flow Matching for Generative Modeling.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Matching for Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.473890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.473890Z digest=sha256:9f71459a891e2d4459dd81022abb90e4805b152c425190ca123c413f721aa1fd

Observation 4a6fcdf9-3da1-4cfb-be82-b81ba54369fa · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.477285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.477285Z digest=sha256:a49f33fd1bc3e6123e2f9dd7ecf6ff30a71a6279015c8639764de62cb73eb9a2

Observation a895c95e-2f21-4848-8a64-1fc62b7c9d19 · outbound

This paper cites De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.480506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.480506Z digest=sha256:ba42dc6f78b26f611f2ddf489f2ed775a916bde42a865f860e12a30bdccbe7a2

Observation 01be9f98-9d40-457e-a09f-925b05f97ef0 · outbound

This paper cites FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.483595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.483595Z digest=sha256:4b95d76233caf78b02f076d1294f76b10267dad6cf5b3a048b76ae18da226396

Observation df2b5a68-f02e-4db3-8dcd-f24c8368982e · outbound

This paper cites Denoising Diffusion Implicit Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Denoising Diffusion Implicit Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.486891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.486891Z digest=sha256:6f1acaecbbaaf00443366029149b7b734f4a172ef2b8aa472546c3c29ec3abf2

Observation ffec1e08-447d-4820-9eec-013ba5787803 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Wan: Open and Advanced Large-Scale Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.490246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.490246Z digest=sha256:ab2eccc620be572b77985c1e3430d22922697bed0575286686aeaf5bd437092f

Observation b1c50b50-82fb-400b-8672-fe18ad1f0ac3 · outbound

This paper cites Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.493575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.493575Z digest=sha256:0afeb2ac642cdada384f5a69e45c266a6816f584ada8460377fc2a2f1c1fc5f3

Observation b1244178-d0ff-4c5a-ba0b-e271f9ac3452 · outbound

This paper cites Qwen-Image Technical Report.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Qwen-Image Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.497114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.497114Z digest=sha256:e48afb8c0434c12a0411f5eb0ba7d39caea28cee597e94b5ce232801d506db87

Observation 4770d7fa-9cce-4846-bb3b-8e84ac3e9d15 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Human preference score: Better aligning text-to-image models with human preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.500748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.500748Z digest=sha256:53ee5c3227e7b44f7df1b416c2b2100994192a2e56584f630707ef218b61c2c7

Observation e0be788b-bf4b-4f9e-9ba9-d24a224a7464 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DanceGRPO: Unleashing GRPO on Visual Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.503814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.503814Z digest=sha256:daa44afbeb78c7b549ffba3d80ffe8a51e9e70b9dc2cef3174085f6f7a0299e0

Observation fbb1a55c-c91a-40e9-b85b-18c7e94a1d49 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.566447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T05:54:32.507119Z digest=sha256:68863ea9b756f5f2f11d01803f0373793c28d5004b49ee80c8992d961e2ce261

Observation 2ccab37e-fedb-4b16-9eb7-fd2b44e3b0df · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.510791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.510791Z digest=sha256:ed6bb432f8bc3e03af223ca3c1a7402be8242dd0cad6d817f1816d47ada5f305

Observation 01ea36d9-80b1-401b-84a9-299eca4b0bae · outbound

This paper cites Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.514207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.514207Z digest=sha256:7e66f996ac7ae792ded6b91f03e65435a2473f3a136139f34c17ed214aeab0f3

Observation 6d317bed-9a7a-45c7-b002-0db3326cc03c · outbound

This paper cites Training diffusion models with reinforcement learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Training diffusion models with reinforcement learning

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.430935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.430935Z digest=sha256:0377a5113ec8d0792780bb22c67a3eb137f203f478432402d6089c7ea0663928

Observation 566e7546-1251-4838-9300-bbc11ddd74d9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.459894Z digest=sha256:dbe371cd52ee423fd73d846e73948defc5239b15c7dfbffe4823555a7970043d

Observation 16dac972-a554-4a30-9948-9036c2d111f4 · outbound

This paper cites Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.445788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.445788Z digest=sha256:dc87646c01344cd72aebb80c29b0dcb4ec159efda625ee0265cb6f2acf18c7da

Observation 5c179457-288b-4b1c-a484-f93c36261d2e · outbound

This paper cites Directly fine-tuning diffusion models on differentiable rewards.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Directly fine-tuning diffusion models on differentiable rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.438659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.438659Z digest=sha256:1e7d1de78a79f0bd093d92700ef4ce90af5ac13e2047c252df5ab725e7b866c1

Observation 8f03720c-298e-4372-b1a4-20b858522778 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.453132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.453132Z digest=sha256:f46647df052a8c4382a11645089671789e5c5930db30d2cc8de00c0a88b735ed

Observation e3072153-392b-4d9c-9309-5c258da56e75 · outbound

This paper cites Diffusion Posterior Sampling for General Noisy Inverse Problems.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Diffusion Posterior Sampling for General Noisy Inverse Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.434813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.434813Z digest=sha256:02840dd628a78461822e6795a6b47f0fe9bb9cdb91065cf030f4a1c550361e52

Observation 25164c7e-f160-4a02-9bfe-9ea5ab4f88b4 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.449343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.449343Z digest=sha256:39b1e8bc4c60568ec7279f2f127c39e5faf4a03c5e12fc566920798ff278733d

Observation 357a86b0-f2f5-4465-b63e-091143c5efd5 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Clipscore: A reference-free evaluation metric for image captioning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.456777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.456777Z digest=sha256:bc687b8b299fb8954ca23a6c3661cb8377fe330ab2266c117562373edb31eefa

Pith citing papers

No inbound Pith citation observations are available.