Pith. sign in

Paper Citation Record · LEDGER

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05600 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:32.517155Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74a6c8e8-bd8b-4f2a-982b-e23e62c00a24 · outbound

This paper cites gummy bear zone.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction gummy bear zone

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.556213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.517155Z digest=sha256:dca2584bb3358503547b0cc4d6a6f7da6cb05fdc75a123cc47ae0fe8b35a989a

Observation 78cdb6ac-7061-47d1-9ffd-515270f00c1b · outbound

This paper cites Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Densegrpo: From sparse to dense reward for flow matching model alignment.arXiv preprint arXiv:2601.20218,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.442334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.442334Z digest=sha256:db3eaa7ca5fcd766a9ba92b94590db90a97dfceb6733cb4b09822e7f3993dbce

Observation 2fbb39f5-cce6-4cdc-a516-4e2205c87a32 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Aligning Text-to-Image Models using Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.463629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.463629Z digest=sha256:5d7d576f6fb29bed653b7bde7ef3f45b64ad063ba8f1ccd26f92d1325ade90b8

Observation 450e6ffc-4d27-4f86-a1e8-d3b8f04101f0 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.466927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.466927Z digest=sha256:ac8a61728587dd48b6e870c16b5483eae7c3e6e3371f67a85697468a867a310f

Observation 39601ea0-7e1a-4235-846f-40869afebf2c · outbound

This paper cites The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:33.179122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.470552Z digest=sha256:399badeef4a4984eea0a2daed3b424d5b5d183c3a137face0e3ae9b0678b5d2d

Observation d6aae682-450c-4d2e-ba33-86b016b218f8 · outbound

This paper cites Flow Matching for Generative Modeling.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Matching for Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.473890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.473890Z digest=sha256:ecce939b40d5cbd17bb9f82f7a4a334f9a0a241ea2d2c46c5e24756fff6a0827

Observation 4a6fcdf9-3da1-4cfb-be82-b81ba54369fa · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.477285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.477285Z digest=sha256:e6482a3c6a54d3a57d5977f611aa1eb96760dae0ae70a3fd0fea808efe41d7b7

Observation a895c95e-2f21-4848-8a64-1fc62b7c9d19 · outbound

This paper cites De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction De- feating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.480506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.480506Z digest=sha256:a934c7c469de60d92e3704d57b2aefcd46c71ef964876eac175703487a4e1ba2

Observation 01be9f98-9d40-457e-a09f-925b05f97ef0 · outbound

This paper cites FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.483595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.483595Z digest=sha256:d6b7559158b052b4ab7acf0e42bcda3f5ffa7ebb645c23cda7b402a3d8f1ac9e

Observation df2b5a68-f02e-4db3-8dcd-f24c8368982e · outbound

This paper cites Denoising Diffusion Implicit Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Denoising Diffusion Implicit Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.486891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.486891Z digest=sha256:d0daf713d95542d4ca04f519f6ff75043702fe08bf8798895d6cee2cce09cd40

Observation ffec1e08-447d-4820-9eec-013ba5787803 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Wan: Open and Advanced Large-Scale Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.490246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.490246Z digest=sha256:582c40e8c299a881624b8af27456851e7a13610f3e1a410b071cf38b3a1fd1ea

Observation b1c50b50-82fb-400b-8672-fe18ad1f0ac3 · outbound

This paper cites Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Coefficients-preserving sampling for reinforcement learning with flow matching.arXiv preprint arXiv:2509.05952,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.493575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.493575Z digest=sha256:5d93c34b032cc3adcec58e242fc05d07905985c380434e4215013efd74d82437

Observation b1244178-d0ff-4c5a-ba0b-e271f9ac3452 · outbound

This paper cites Qwen-Image Technical Report.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Qwen-Image Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.497114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.497114Z digest=sha256:c2effb50ed8f7bb3545a343719b6919acfa49be9a328e7c0a8ec0b76eff2cdc9

Observation 4770d7fa-9cce-4846-bb3b-8e84ac3e9d15 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Human preference score: Better aligning text-to-image models with human preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.500748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.500748Z digest=sha256:8e718e4f4de35760e96469ac7c65e0684bb7fd6d0c995f915f5aff2c8c272102

Observation e0be788b-bf4b-4f9e-9ba9-d24a224a7464 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DanceGRPO: Unleashing GRPO on Visual Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.503814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.503814Z digest=sha256:9c5a4d0562d760bdc5d53ad36387c0a26f4909c9fec9f762dfa85c9e08435711

Observation fbb1a55c-c91a-40e9-b85b-18c7e94a1d49 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:33.566447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:32.507119Z digest=sha256:85eff2a64ddbe927628d30d55f2f695171bb7645d9b19e7b09ad9090de743243

Observation 2ccab37e-fedb-4b16-9eb7-fd2b44e3b0df · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.510791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.510791Z digest=sha256:3ec7ff48e4d40e9fbbd3b0f7bf6141633099e36fb2a55bb7bfe618291e5f3660

Observation 01ea36d9-80b1-401b-84a9-299eca4b0bae · outbound

This paper cites Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Manifold-aware exploration for reinforcement learning in video generation.arXiv preprint arXiv:2603.21872,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.514207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.514207Z digest=sha256:90ff7ba98165a4e7b12e9a32d93f0a930289eb097e3d8c1e90a712d6876c5949

Observation 6d317bed-9a7a-45c7-b002-0db3326cc03c · outbound

This paper cites Training diffusion models with reinforcement learning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Training diffusion models with reinforcement learning

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.430935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.430935Z digest=sha256:67ae079d9d06772a289a287c7b75475aa24f6b0684188e8916974fae8c2bfd83

Observation 566e7546-1251-4838-9300-bbc11ddd74d9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.459894Z digest=sha256:81f79cdfdbd757dd7d23c63ef46089d8ade3e8bd82b28682e0adcb9d18061fbd

Observation 16dac972-a554-4a30-9948-9036c2d111f4 · outbound

This paper cites Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Treegrpo: Tree-advantage grpo for online rl post-training of diffusion models.arXiv preprint arXiv:2512.08153,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.445788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.445788Z digest=sha256:f34321162bd6637f837397b1e678a2a030cf4f2d760b3186120565d68b43149a

Observation 5c179457-288b-4b1c-a484-f93c36261d2e · outbound

This paper cites Directly fine-tuning diffusion models on differentiable rewards.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Directly fine-tuning diffusion models on differentiable rewards

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.438659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.438659Z digest=sha256:616aadefef793aa7bd938090efe5673b31791683addd0991e0906768ead5a6b5

Observation 8f03720c-298e-4372-b1a4-20b858522778 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.453132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.453132Z digest=sha256:2c71ea7f1a4453b922f6bd323d480952b7c0a5d182dcd1aa04f83fb617d3cf32

Observation e3072153-392b-4d9c-9309-5c258da56e75 · outbound

This paper cites Diffusion Posterior Sampling for General Noisy Inverse Problems.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Diffusion Posterior Sampling for General Noisy Inverse Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.434813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.434813Z digest=sha256:b3e76e24f9fdcfe2338da331211c638d3aa2717e260ec81f9a2f6ddb269d7fe3

Observation 25164c7e-f160-4a02-9bfe-9ea5ab4f88b4 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.449343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.449343Z digest=sha256:fabfd07dada32015f04e666088900f1a7c7782d81ce104486837fc5a0c0d2c5b

Observation 357a86b0-f2f5-4465-b63e-091143c5efd5 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction Clipscore: A reference-free evaluation metric for image captioning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:32.456777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:32.456777Z digest=sha256:dbf0e32c2c80a986ff5ea6c0dd28e0682b4751653667bc23d473d0fa4c1d1212

Pith citing papers

No inbound Pith citation observations are available.