Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 5 inbound Pith citation observations for arXiv:2511.19356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19356 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:36:08.401563Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T04:29:45.017327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:23:15.759179Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a26b75e-ccdb-4541-94ad-957a160d7f01 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.488757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.488757Z digest=sha256:5f40ef57c15468c33bbe5ede655de4bc61df9706911cb0c5987280797d5b04ab

Observation 7fbcfb2d-1706-4683-aed3-2eb00a406eee · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.524886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.524886Z digest=sha256:8dfa0602c5418e890021800f452d604c6dd4b791ac8ca27a74e201bbcc3aa42c

Observation 4ffd453f-6682-4f08-b348-8ab98838e177 · outbound

This paper cites Reinforcement learning in continuous time and space.Neural computation, 12(1):219–245, 2000.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Reinforcement learning in continuous time and space.Neural computation, 12(1):219–245, 2000

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.580757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.580757Z digest=sha256:4934c074b53f1778700e717c5ea1af044ddd95ca68895893205e509d16dcd58d

Observation c372b29d-e85b-44d5-88db-1a97d8c7edca · outbound

This paper cites InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.640778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.640778Z digest=sha256:d853ecfbbb687fae0493c9b839319d0f2d6ece50b5129f71db8b472693a755d4

Observation b8c870be-03be-4cf7-884c-a0d416dd60bc · outbound

This paper cites You only look at one sequence: Rethinking transformer in vision through object detection.Advances in Neural Information Processing Systems, 34:26183–26197, 2021.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets You only look at one sequence: Rethinking transformer in vision through object detection.Advances in Neural Information Processing Systems, 34:26183–26197, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.709921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.709921Z digest=sha256:639274933cd9c2a887cbd0e4fe8fc1628eb1b86b86113e8bff386d339b564f12

Observation d0b4da14-7099-4ac6-a8d1-ec0faacc4d3a · outbound

This paper cites Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.771107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.771107Z digest=sha256:4b87b75d58f03afc60913b75091758a86869cf6a2a05a35c149d32d384bfd4d4

Observation 8f136264-316f-4b35-bea6-c297c3a28190 · outbound

This paper cites Hierar- chical process memory: memory as an integral component of information processing.Trends in cognitive sciences, 19 (6):304–313, 2015.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Hierar- chical process memory: memory as an integral component of information processing.Trends in cognitive sciences, 19 (6):304–313, 2015

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.878239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.878239Z digest=sha256:832c2f60a972f94f8a09ea89255635b59d5e1ddf93b3ab3413e490c5e740f1fc

Observation 0ed3e48a-db57-4fd3-b1a0-0585347694f0 · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.891548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.891548Z digest=sha256:49aa531a2bbe2e9fd4df1602eb68510f32372b1cb897ab1c6bffc267a25f3764

Observation f9dcbf46-a30c-4650-906c-97ad8217fc24 · outbound

This paper cites Non-negative matrix factorization with sparseness constraints.Journal of machine learning re- search, 5(Nov):1457–1469, 2004.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Non-negative matrix factorization with sparseness constraints.Journal of machine learning re- search, 5(Nov):1457–1469, 2004

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.936110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.936110Z digest=sha256:233a5bdc59a80dd324b6c4ac155f1faea49868e29c687d6f27bb0cef2dd07729

Observation 2b19eb94-d592-44b4-bf51-c4ef1596d038 · outbound

This paper cites A reinforcement learning-based automatic video editing method using pre-trained vision-language model.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets A reinforcement learning-based automatic video editing method using pre-trained vision-language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.985497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.985497Z digest=sha256:2d59fda5c7b6838ac621fdfbf6868dc8f9b03d7eb5c617f02f89716564676139

Observation bb29fbe3-b2a9-4ce3-beb3-4f54649712ed · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vbench: Comprehensive bench- mark suite for video generative models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.095455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.095455Z digest=sha256:1ff048e036ee761a1e49fff600369e2f0e4976f279b93bf1051553ce04af5a27

Observation 3fe8ecbf-bf06-499f-8aac-c389af856352 · outbound

This paper cites istock video dataset.https : / / www.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets istock video dataset.https : / / www

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.157189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.157189Z digest=sha256:d1e4c0a046378239f87acfeb11bd4ffaf9ebf6b5e0e0e297b33ccdaf9cf4697b

Observation dfe3484d-e872-494c-8a88-fb3fa698eb21 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.225043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.225043Z digest=sha256:ec304ce5712546ee682419144c2576db5c45c453425ba883fbc0c2be7789f633

Observation bf740fa4-ed8d-4eb3-b7d0-30c45294ce47 · outbound

This paper cites Early language acquisition: cracking the speech code.Nature reviews neuroscience, 5(11):831–843,.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Early language acquisition: cracking the speech code.Nature reviews neuroscience, 5(11):831–843,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.329285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.329285Z digest=sha256:27a0ddf2c4534aec60f85e3ad2b28f893f50409c7baae474d58a661bddb50958

Observation 0ad24454-862c-4482-8ccf-81fa676d4ba9 · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.406540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.406540Z digest=sha256:0d4a10544b431e650715b0f48627a202aec6613c820799db07baeecb1ea8bd5a

Observation 5db676d6-2ab4-4175-bdba-39f7c6d3850f · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.500523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.500523Z digest=sha256:f13bca79435ee63ba6e6b6f6d5ea2c3440cb2aeb45df1bec5e9f9eadfc33433b

Observation baf2ad4b-590f-401f-978b-c3089fe149a8 · outbound

This paper cites Spatial-then-temporal self-supervised learning for video correspondence.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Spatial-then-temporal self-supervised learning for video correspondence

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.566506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.566506Z digest=sha256:8c6e5b307f86703bd87c45f8f14c4fa740f90d8cb774e1da90a62692c2007657

Observation fcce184f-2472-4bae-9d17-de9bcbb1404a · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vila: On pre-training for vi- sual language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.634807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.634807Z digest=sha256:932967edea4db82a8a09e1007f5bfc4bee5c99d2b9a9afb932fb312daf07057f

Observation eec0061d-2857-4c57-aad3-9bb6de700600 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Flow-GRPO: Training Flow Matching Models via Online RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.785918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.785918Z digest=sha256:07db6da0d982d23eaafe07ee2b72e03bc19f0a039a5569ab9d5f1a0f442f55f4

Observation 154dd822-2a02-43a7-99a6-6f69ec2495f7 · outbound

This paper cites Improving Video Generation with Human Feedback.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Video Generation with Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.855065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.855065Z digest=sha256:f4ec0a85639e621659673f382ccee297adefb5c2e01e02b0b83dfc19eba00ab5

Observation 27cf8f20-695f-42bc-84dc-dc927d85c6f1 · outbound

This paper cites When the future becomes the past: Taming temporal correspondence for self-supervised video represen- tation learning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets When the future becomes the past: Taming temporal correspondence for self-supervised video represen- tation learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.920341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.920341Z digest=sha256:d87759e4d492a436158fa42e25809cecf85ead2300af02a09fac7e81722e837e

Observation 8f458605-b810-4ace-a867-b7761e23c5cc · outbound

This paper cites Enhance-A-Video: Better Generated Video for Free.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Enhance-A-Video: Better Generated Video for Free

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.992333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.992333Z digest=sha256:8abd06a672e8c8ed9e8eb6f293330eaf2dcadf04c7ca62fa5ae98cf656178536

Observation 42c0482e-3c03-40c8-b80a-e1abebd19b69 · outbound

This paper cites Liv: Language-image represen- tations and rewards for robotic control.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Liv: Language-image represen- tations and rewards for robotic control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.038736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.038736Z digest=sha256:99af88a179972238b67eaf56210dca8d042cbef18e6ccd16ade5320fd73ae094

Observation 8b1d7051-8449-4d88-8b13-ffdd7cd19a90 · outbound

This paper cites Video Diffusion Alignment via Reward Gradients.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Video Diffusion Alignment via Reward Gradients

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.105901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.105901Z digest=sha256:5839e29fa923feda4bef2ccc5f7455c80170303b337af016db624fecd02ed9d3

Observation b3afd3ad-b9b6-4d74-bbbe-462bae08cb7a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.166691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.166691Z digest=sha256:625e17a38f6f34977d3e74d75eaf54a3b87471ee984c836808a1a2800824860e

Observation e65b5bb0-a5c6-4c4b-bef3-69403e135409 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.266533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.266533Z digest=sha256:007fce2c649f72bb28fbc645967385e708269ac4f71ea15b0cea2ef377d9541d

Observation bdedb479-18c1-4fa4-b884-2e3eabbf9305 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.330783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.330783Z digest=sha256:dc40a0f0ad9c71088331303c3c8b161f33a5a48d3284bd3d77ecea70a0ef0bde

Observation 51ec4549-e869-4fa7-860b-df6749cf4024 · outbound

This paper cites Predictive reward signal of dopamine neu- rons.Journal of neurophysiology, 1998.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Predictive reward signal of dopamine neu- rons.Journal of neurophysiology, 1998

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.378888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.378888Z digest=sha256:c806f0ab3e2adfdd5178b4f64ff463203f61dcf407bb8a4b99589eda266cc85f

Observation 7478f871-005d-4492-8f96-b64fe333e03d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.472490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.472490Z digest=sha256:2be0b249f584e25c1f5cc43538ad0c5df81fc6547b396d8bebe9b711ede690c8

Observation 09bffd77-3d7a-4723-9087-5940c2d7594a · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.684556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.684556Z digest=sha256:00926ab94ce29003c2c23d407e54b112ef03fa27de8deb904ff08469cfb61223

Observation ff7f33f3-eb8b-42b9-8c58-d3ba12d758c5 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Wan: Open and Advanced Large-Scale Video Generative Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.757082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.757082Z digest=sha256:7f3654b933bf4f798e9a3c8b84df79c9b223c672802235879b79fd1691f0b8b9

Observation 1148c233-707d-42d5-8580-ef6b65130e5e · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.941340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.941340Z digest=sha256:f16d282214d9dc45aeb1e0134f8ca6a4ab513c64ef4883b7ce4f30780f29de1e

Observation da874b86-f657-4855-9bd6-5e66fe2fced6 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.000682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.000682Z digest=sha256:ddf5c0430abe1e3569e99d60c9c978685fd89b9d59457c2c22986f59cd58a767

Observation 27b698f3-cd13-49ca-985c-5b610c9a5b47 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets DanceGRPO: Unleashing GRPO on Visual Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.085061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.085061Z digest=sha256:66b81d55e1f1bd6feee6d3c87bcfb7600576818479e9da24c7d7d0494b24c0a4

Observation f5292d51-25b2-4bed-b186-ec7a93c6829b · outbound

This paper cites Self-rewarding large vision-language models for opti- mizing prompts in text-to-image generation.arXiv preprint arXiv:2505.16763, 2025.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Self-rewarding large vision-language models for opti- mizing prompts in text-to-image generation.arXiv preprint arXiv:2505.16763, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.240538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.240538Z digest=sha256:a3ab8ff0715648f47886b9aaa4a683a347299bf74015005836c4176dbe8df461

Observation 2800b94a-0457-477b-b420-fa151df9509f · outbound

This paper cites {text_prompt}.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets {text_prompt}

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-03T20:36:08.401563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.401563Z digest=sha256:dbf698feb73ce9b8d054dda60fc7447c47fa60df94fc82829609eee91fadb3cd

Pith citing papers

Observation f1482365-50cf-48cd-a368-7d3c056db61a · inbound

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation cites this paper.

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:10:26.580964Z digest=sha256:2517888c9a372e8a3f905f99fe6187ec9cd17ade7834a819701a65afdfbdf8eb

Observation db3f62fc-c43c-45e2-9240-efef6d48f3fe · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:9dc77a87c31c8644f602b5e4d96980078edd8582cb271c1f653bc23e22be56af

Observation 0f6f0326-c438-4795-a64c-3eb17e0c1d8d · inbound

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning cites this paper.

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:14:20.558527Z digest=sha256:0dabfe2510799e19013ed562404d8c743737c87b048a4df571d6a665a7249b04

Observation 7a91ca6b-51ee-4564-8968-651b80867ef4 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:45.017327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:45.017327Z digest=sha256:d9f907506fc8fb6bf82fe878e501fb624853bcd8c6571187eaae64e0f80a5f2a

Observation 2a221d75-1be6-4883-ac0e-b10d67f04d23 · inbound

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision cites this paper.

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:49:59.034747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:49:59.034747Z digest=sha256:7bef50b14f514c40c60c400ef64178b57ac2c628fb402450e9b0a6e1c9b69110