Pith. sign in

Paper Citation Record · LEDGER

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 5 inbound Pith citation observations for arXiv:2511.19356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19356 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:36:08.401563Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T04:29:45.017327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:23:15.759179Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a26b75e-ccdb-4541-94ad-957a160d7f01 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.488757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.488757Z digest=sha256:f9648ed152b0a22f8e2d1661bbc858441e79cd6b21ade666e4df8485157a6c99

Observation 7fbcfb2d-1706-4683-aed3-2eb00a406eee · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.524886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.524886Z digest=sha256:572b19ba6eac82c807190c8be5603d54b4eaa6b9e9b14642234ea19f0d74e9a2

Observation 4ffd453f-6682-4f08-b348-8ab98838e177 · outbound

This paper cites Reinforcement learning in continuous time and space.Neural computation, 12(1):219–245, 2000.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Reinforcement learning in continuous time and space.Neural computation, 12(1):219–245, 2000

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.580757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.580757Z digest=sha256:cebb1bd6633333b2b609570b7f24a9b5255721acb5113aa098cbe7ce15e294a9

Observation c372b29d-e85b-44d5-88db-1a97d8c7edca · outbound

This paper cites InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.640778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.640778Z digest=sha256:6915b2e5293325e1e896a516a290011334c02536de06829609e650fd1d4274de

Observation b8c870be-03be-4cf7-884c-a0d416dd60bc · outbound

This paper cites You only look at one sequence: Rethinking transformer in vision through object detection.Advances in Neural Information Processing Systems, 34:26183–26197, 2021.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets You only look at one sequence: Rethinking transformer in vision through object detection.Advances in Neural Information Processing Systems, 34:26183–26197, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.709921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.709921Z digest=sha256:6e686d8678fb2737e0c220e321b865d0bc586e80359d6bacfb5a304dc363ca11

Observation d0b4da14-7099-4ac6-a8d1-ec0faacc4d3a · outbound

This paper cites Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.771107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.771107Z digest=sha256:b17cec6bd0220e99e53e32fbb19799b0f09606665ed56a86e6b67378edc67baa

Observation 8f136264-316f-4b35-bea6-c297c3a28190 · outbound

This paper cites Hierar- chical process memory: memory as an integral component of information processing.Trends in cognitive sciences, 19 (6):304–313, 2015.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Hierar- chical process memory: memory as an integral component of information processing.Trends in cognitive sciences, 19 (6):304–313, 2015

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.878239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.878239Z digest=sha256:9cf200c5c79eace572541ecce4eedd9724caab6d8a40f49719460dbc99d59926

Observation 0ed3e48a-db57-4fd3-b1a0-0585347694f0 · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.891548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.891548Z digest=sha256:73cf5a75cc6d15096dc717d6913e4967f04b325856421d2e8ed2bab3a8a8fca1

Observation f9dcbf46-a30c-4650-906c-97ad8217fc24 · outbound

This paper cites Non-negative matrix factorization with sparseness constraints.Journal of machine learning re- search, 5(Nov):1457–1469, 2004.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Non-negative matrix factorization with sparseness constraints.Journal of machine learning re- search, 5(Nov):1457–1469, 2004

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.936110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.936110Z digest=sha256:6ef4cf468c159d8703f64e6408cc412825bc8dc7eb385d4373cba187d7fe2e5b

Observation 2b19eb94-d592-44b4-bf51-c4ef1596d038 · outbound

This paper cites A reinforcement learning-based automatic video editing method using pre-trained vision-language model.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets A reinforcement learning-based automatic video editing method using pre-trained vision-language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.985497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.985497Z digest=sha256:bad341e1cdf2401e40a40521e03ac2e4da44b0db2c60f882dcd901a9520d1f02

Observation bb29fbe3-b2a9-4ce3-beb3-4f54649712ed · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vbench: Comprehensive bench- mark suite for video generative models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.095455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.095455Z digest=sha256:6e5eb8bdc6b5ee61f6da46321fb6bc2258da2b0f2176c2a841fb08c3055e0fd7

Observation 3fe8ecbf-bf06-499f-8aac-c389af856352 · outbound

This paper cites istock video dataset.https : / / www.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets istock video dataset.https : / / www

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.157189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.157189Z digest=sha256:d4f800538ce4c4bd3a9b14a157ca0df2d6deb6b1d0e7b1b89bf6bb37c9bd1888

Observation dfe3484d-e872-494c-8a88-fb3fa698eb21 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.225043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.225043Z digest=sha256:8ae7a39951f68dfdafc1ee1d47c7b24552a031fdb6aa56c13d32db640ff45fd5

Observation bf740fa4-ed8d-4eb3-b7d0-30c45294ce47 · outbound

This paper cites Early language acquisition: cracking the speech code.Nature reviews neuroscience, 5(11):831–843,.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Early language acquisition: cracking the speech code.Nature reviews neuroscience, 5(11):831–843,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.329285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.329285Z digest=sha256:62cdc02ebed214404d163fd951a754d6663bd9c6d2b6d0c9b1936f8bc9df6dde

Observation 0ad24454-862c-4482-8ccf-81fa676d4ba9 · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.406540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.406540Z digest=sha256:0a6a171ec66b92aa33da78bb145090a1682034b8fba4d315c6452352a3cea7c6

Observation 5db676d6-2ab4-4175-bdba-39f7c6d3850f · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.500523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.500523Z digest=sha256:524551c1ba1db7f7cd086392565e96726b6a7633b096bba03c65b8a7b5034245

Observation baf2ad4b-590f-401f-978b-c3089fe149a8 · outbound

This paper cites Spatial-then-temporal self-supervised learning for video correspondence.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Spatial-then-temporal self-supervised learning for video correspondence

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.566506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.566506Z digest=sha256:351394b6852c4dba454f6e20010b167b00337e143c4070bf5ac0e3c53043b155

Observation fcce184f-2472-4bae-9d17-de9bcbb1404a · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Vila: On pre-training for vi- sual language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.634807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.634807Z digest=sha256:551d01361c233394698292186bbd9838ac71d789a7480dc2e693ff65efac653d

Observation eec0061d-2857-4c57-aad3-9bb6de700600 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Flow-GRPO: Training Flow Matching Models via Online RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.785918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.785918Z digest=sha256:dbd3e5e6e9deb5f39ead6cce3ffb17473458d8f9d8562fab4cc6e8a15e03b446

Observation 154dd822-2a02-43a7-99a6-6f69ec2495f7 · outbound

This paper cites Improving Video Generation with Human Feedback.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Improving Video Generation with Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.855065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.855065Z digest=sha256:d228f9fe54d0df11eba9762e9eafdb0d35f372e0c9676d43c3c7717fef19b323

Observation 27cf8f20-695f-42bc-84dc-dc927d85c6f1 · outbound

This paper cites When the future becomes the past: Taming temporal correspondence for self-supervised video represen- tation learning.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets When the future becomes the past: Taming temporal correspondence for self-supervised video represen- tation learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.920341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.920341Z digest=sha256:f92fc72f5f4a8bf31937c02fcd7e2ad0e4bc7af4f307882c4036067bb208da03

Observation 8f458605-b810-4ace-a867-b7761e23c5cc · outbound

This paper cites Enhance-A-Video: Better Generated Video for Free.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Enhance-A-Video: Better Generated Video for Free

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.992333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.992333Z digest=sha256:54bc0f70f5e6df0a9091068a1158c476c068c03e380b2a856e79814bd800bcc0

Observation 42c0482e-3c03-40c8-b80a-e1abebd19b69 · outbound

This paper cites Liv: Language-image represen- tations and rewards for robotic control.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Liv: Language-image represen- tations and rewards for robotic control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.038736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.038736Z digest=sha256:837bcf3f1e128da88ddbb26e07c3c8ff1a7af0f28387ed75761d14d2dae9cfe6

Observation 8b1d7051-8449-4d88-8b13-ffdd7cd19a90 · outbound

This paper cites Video Diffusion Alignment via Reward Gradients.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Video Diffusion Alignment via Reward Gradients

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.105901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.105901Z digest=sha256:69f2ba9b648403402784fb6838fe60846043f978a1e4543243c40c0d026ea9cf

Observation b3afd3ad-b9b6-4d74-bbbe-462bae08cb7a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.166691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.166691Z digest=sha256:a12b50d44d8bf471b4920a534f1d698122ded0973387744d4051ecee1423171e

Observation e65b5bb0-a5c6-4c4b-bef3-69403e135409 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.266533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.266533Z digest=sha256:582de9dafe0a9a22beb5f4df2137742580cadbe03cc96d8bc62481794483ff58

Observation bdedb479-18c1-4fa4-b884-2e3eabbf9305 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.330783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.330783Z digest=sha256:843685731620da06a9f61ac30020d4824c5a9db9a49dbc5745565e53d77f3e6e

Observation 51ec4549-e869-4fa7-860b-df6749cf4024 · outbound

This paper cites Predictive reward signal of dopamine neu- rons.Journal of neurophysiology, 1998.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Predictive reward signal of dopamine neu- rons.Journal of neurophysiology, 1998

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.378888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.378888Z digest=sha256:206864003b020f024eb59ada1044ecf37ebb5ebb488454635514d55dbaf5e236

Observation 7478f871-005d-4492-8f96-b64fe333e03d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.472490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.472490Z digest=sha256:332e695a0fb54e4bd5cb3893d231ffb4fa6b554eebf036ff79fad3ebf26969ec

Observation 09bffd77-3d7a-4723-9087-5940c2d7594a · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.684556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.684556Z digest=sha256:bb4f2706efd348e464a18accf4be90a25abdf3c3651e4506671a83491139da00

Observation ff7f33f3-eb8b-42b9-8c58-d3ba12d758c5 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Wan: Open and Advanced Large-Scale Video Generative Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.757082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.757082Z digest=sha256:7dc5d0e7bf96629f7db6f954171dcd8f81b7d41643c7eac87ad699a92c1e4316

Observation 1148c233-707d-42d5-8580-ef6b65130e5e · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:07.941340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:07.941340Z digest=sha256:b4b40d3425d054e696bf3e5ec8ef4692970c5f57e449cb7fb5474337c159f98b

Observation da874b86-f657-4855-9bd6-5e66fe2fced6 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.000682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.000682Z digest=sha256:5d2fa530f1a7fcd701205e046667caafccfcee00e5cc5d7883dc0f96c14a9374

Observation 27b698f3-cd13-49ca-985c-5b610c9a5b47 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets DanceGRPO: Unleashing GRPO on Visual Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.085061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.085061Z digest=sha256:5c6d25ac6fc5550991ae652f54f3b0e4577b7165b49976dc51e327b723315b4f

Observation f5292d51-25b2-4bed-b186-ec7a93c6829b · outbound

This paper cites Self-rewarding large vision-language models for opti- mizing prompts in text-to-image generation.arXiv preprint arXiv:2505.16763, 2025.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets Self-rewarding large vision-language models for opti- mizing prompts in text-to-image generation.arXiv preprint arXiv:2505.16763, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:08.240538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.240538Z digest=sha256:2a1b1aa25796799c7d47d797848119ddfee618216aa824212e4ce8c452b05e40

Observation 2800b94a-0457-477b-b420-fa151df9509f · outbound

This paper cites {text_prompt}.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets {text_prompt}

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-03T20:36:08.401563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:08.401563Z digest=sha256:857287317d8856ed04759fd2282bbfb608d46f88a1850f561564cf5eb36168a0

Pith citing papers

Observation f1482365-50cf-48cd-a368-7d3c056db61a · inbound

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation cites this paper.

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:10:26.580964Z digest=sha256:507d7a53b5c0159a40378773a02f6c19b54c920cc50fb89186f4ddfda17dc9d8

Observation db3f62fc-c43c-45e2-9240-efef6d48f3fe · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:da886c072c8e227ad123165a9d098fae32d947bcf457cceb857666f82630f687

Observation 0f6f0326-c438-4795-a64c-3eb17e0c1d8d · inbound

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning cites this paper.

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-17T02:21:30.636526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:14:20.558527Z digest=sha256:186a86e7e9d7e8e3230c371452bdc725154a1e843a0d4405c37ba97701f4eac0

Observation 7a91ca6b-51ee-4564-8968-651b80867ef4 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:45.017327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:45.017327Z digest=sha256:2807ee704c045287acae5145bb7517485df07a6891949dc07d11d5ce6b3b52c9

Observation 2a221d75-1be6-4883-ac0e-b10d67f04d23 · inbound

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision cites this paper.

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:49:59.034747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:49:59.034747Z digest=sha256:500907a2297e27b60e885963d4f7eb92545888efa90c3897bd77cf047365d9ad