Pith. sign in

Paper Citation Record · LEDGER

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2502.06061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06061 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:57:42.654116Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.257162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T19:03:51.552172Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved7
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation add414f3-fad3-488d-a656-799906ed1732 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.064886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.556146Z digest=sha256:548b60f3521c7af5c6e55af031a2dbd89ce10f7ceb84324a92b39d5facb515f0

Observation 4047a014-4084-48fa-bfc8-b790c9635c23 · outbound

This paper cites To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1).

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization To address this challenge, we introduce W2 regularization, which effectively prevents over-optimization and policy collapse (Lemma 1)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.048103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.562201Z digest=sha256:39f860251fda7b181c421367bdd994a852d3b93e816bf7601bdb702c6709b047

Observation c6adba18-3860-4df4-88a8-87d193498cc8 · outbound

This paper cites left of”, “on top of.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization left of”, “on top of

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.081852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.550430Z digest=sha256:ce1e455b3ab4c8841f08f323ebd2d05c93a7a3d0433c4d05904e99ee088b3ff3

Observation 8848c0cf-3223-4696-b1c9-b9e4d71afab0 · outbound

This paper cites a cat in the sky.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a cat in the sky

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.029134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.567605Z digest=sha256:639a57e777c156b2c5329b50783e90549a9cd2c111b2dc3f1e1875b0de5f68d2

Observation f7cdfa02-8a77-4434-8a4c-3dc8e23792bd · outbound

This paper cites a train on top of a surfboard,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization a train on top of a surfboard,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:43.006766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.573522Z digest=sha256:09adec6e6cf1e56874779d6f1fb522601cf33d5bd1bf74a14636b48df90b1cf1

Observation b9e49203-b038-4437-af6c-c5399e991a44 · outbound

This paper cites 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization 3) Bottom row: Alpha Clip reward showcases consis- tent performance even with different text-image alignment rewards, validating the reward-agnostic nature/property of our approach

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.984986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.578823Z digest=sha256:31e4427b842d4a89c58314c8bf7074fd9661924227722e18d7c16a4e1a9704eb

Observation 7e8aa6a2-b43a-4b18-8778-2e6f50948a7e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.968182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.584564Z digest=sha256:9d0f64529c89b22726855fcf949fd0945d7db6270be0eb4afe0da7e4b73fd305

Observation 9f8a3694-bc36-4e72-90d6-3ead5fe9bcbd · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.950919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.589732Z digest=sha256:52de061dcebc2391220d836dcc7b36cda32ca1ab64172067f60c17a6599c898d

Observation d28eda71-a0da-4321-847d-cbecdb6d56fc · outbound

This paper cites on top",.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization on top",

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.933892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.595160Z digest=sha256:4eca41d3be90328093aa6264a61b82df5614bc330bfcc2fea0cdede385ddedbc

Observation 4066ae4b-8442-4f8a-aa12-ffbb3ade7660 · outbound

This paper cites Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Given that w (x1) > 0 for all x1 ∈ Xand attains its maximum at x∗ 1, we define: ϵ (x1) = w (x1) w (x∗

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.879540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.612598Z digest=sha256:527044106d0d30c2d800502df9b8e185d9d2b0191f4ff37bca0d6036b63c9c0a

Observation a396bb3a-3861-4a00-84dc-28a8157f6d16 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 15

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T16:57:42.861835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.617699Z digest=sha256:30e99559e4b612310d33a2fed3867acec1ad6bed6135ee65fcefcceaf34159e5

Observation 440d2089-c959-4adc-bfa9-168ca71fe817 · outbound

This paper cites Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can rewrite qN θ (x1) using ϵ(x1) as: qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) ZN (73) Then for x1 ̸= x∗ 1, we can have: ϵ(x1)N → 0 as N → ∞since ϵ(x1) < 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.844221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.623385Z digest=sha256:dc5dcb9ea13963566a3afb458a9b02cbe0b6a32ed3a4493d7e9b8e145a26af10

Observation 12896a55-1a87-4433-b125-2cd1f4707661 · outbound

This paper cites And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization And we can have the normalization constant as follows: ZN = Z X w (x1)N q (x1) dx1 = [w (x∗ 1)]N Z X ϵ (x1)N q (x1) dx1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.826271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.628335Z digest=sha256:3392b215bb6a439e86140d962a58874252e2d543a8d8ecaed5ea53752a210a9d

Observation 630b8d02-6064-4967-aa5d-b271f77edd9c · outbound

This paper cites Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Then, we can have the limit behavior: For x1 ̸= x∗ 1 : qN θ (x1) = [w (x∗ 1)]N ϵ (x1)N q (x1) [w (x∗ 1)]N q (x∗ 1) = ϵ (x1)N q (x1) q (x∗

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.807617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.633170Z digest=sha256:acd92abac1edd3e873d77f64eca85339a44215bd75d37995513c3e18f44cff22

Observation bfea99bd-4bc6-456c-aebf-059103fa33c8 · outbound

This paper cites (75) For x1 = x∗ 1 : qN θ (x∗.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization (75) For x1 = x∗ 1 : qN θ (x∗

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.789504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.637958Z digest=sha256:ec1e7415319ae8056b555dccc960b1e85268b7e674801860059f9dceda22b0b3

Observation 68dbfe68-f857-4469-8984-b3681be2010d · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.772687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.642634Z digest=sha256:f4e8d1f3ec3be16f2e5955d1969d62f557304f6aa5caa7e541348c1a76a71750

Observation 96598b70-5c5e-4947-b81d-82bee4ab30e4 · outbound

This paper cites cat") = pclip (x,.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization cat") = pclip (x,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.756262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.647660Z digest=sha256:c3f2909b75ec3b5307cb5a6668f0a274d81ae62bb69ed1f815bc6f4be3114534

Observation e95af9ff-2d57-4371-b120-ad8d9cf1bf3e · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.738323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.654116Z digest=sha256:f37aeeceeaf97c3119a82d06b9a5bb4abc95cf40ea7b1a9903868aac0ba7c785

Observation e5a1e71e-039e-4ec2-a381-5cf4f4b03a79 · outbound

This paper cites In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization In this paper, we introduce two methods to handle the Overoptimization and ease the mode collapse risk in online RW-CFM algorithms

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:42.897400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.606891Z digest=sha256:da7a862c65ae0a534e20024c49ab65a378a7b2f4f7daf39077757e9968ae8125

Observation b87d7f19-3b44-4ef1-834d-3d443c4d9783 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Overcoming catastrophic forgetting in neural networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.537531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.537531Z digest=sha256:74741cf8186901cd9e59148bfeee7d450d941cfd450cd03d9fb8833ca88110e5

Observation c132e1f4-b409-48e7-a311-447bfe713566 · outbound

This paper cites an unresolved cited work.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:42.916713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T16:57:42.600862Z digest=sha256:539aa864179b4e842985c76262781b152ffe79342e17c6fc5fc3a0b61b540fe3

Observation 1435c790-34aa-44ff-b440-baa7ed828949 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization Understanding the performance gap between online and offline alignment algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:42.544295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:42.544295Z digest=sha256:c0a9d4b8fef6d341447db9a963be2b46fc399d53d8a709b0bfebcaf9727f6e45

Pith citing papers

Observation d02ead63-07d8-4762-a248-384541f0024a · inbound

Flow-Based Policy for Online Reinforcement Learning cites this paper.

Flow-Based Policy for Online Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.257162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.257162Z digest=sha256:92e6fa8524ad2f84c7373564848e613748bb4ad990f683a081ed4cc53a5d71f9

Observation 5b73787a-0011-4fb3-8fcf-a88a0e1b9001 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.214222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.214222Z digest=sha256:cdad309abc14cbe390ed9aa1cd6ed91fc82c933abb6d1fb8535b1d30498c746b

Observation b909be6c-2528-444c-96c2-60f7217c6854 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:08.654870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:08.654870Z digest=sha256:38656f62e1f5ace9d36a8ebdc73844e1603d1117d5697de12bf1746aeb5af2d5

Observation 5a44649f-859a-42bf-97b9-d208d985208c · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.553716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:00f3ef24aaf203b69b0d33ecaaf32cae788f7710af4c35f7442d309fe7c09dad