Pith. sign in

Paper Citation Record · LEDGER

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 13 inbound Pith citation observations for arXiv:2505.17017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17017 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:56.714854Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:52.913763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.820556Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3e0d43a-e781-4bb8-9f69-4309bee371a1 · outbound

This paper cites https://www.anthropic.com/claude/sonnet/, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://www.anthropic.com/claude/sonnet/, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.913770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:50.317351Z digest=sha256:2ea9f73a5a46a685f1d7b3e5ecdb5cc3b55d8d1849abe6dbbff8b5f6ef1f8b3e

Observation 359f5ec5-250e-4b04-a289-1452bb414858 · outbound

This paper cites https://deepmind.google/technologies/gemini/pro/, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO https://deepmind.google/technologies/gemini/pro/, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.726291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:50.378463Z digest=sha256:4c72d994809d1638c993246181420b5522a945d6710ae0f0397a36edb3f9701f

Observation 7f2ec122-7832-494f-8238-9c8fa8bef65c · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.431952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.431952Z digest=sha256:9dd797cf11aeeee62cdb73b0dd275fb0e49f897d98e4512366e52b3536319ffc

Observation 0f70fb72-5c3c-45e7-8989-7a6151e3e5b2 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.485153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.485153Z digest=sha256:63e7f9c6526f4a8b6abc01614b9b940d33fd7384a4daed1a2c6f796f2d04844f

Observation 865e8eba-a415-411f-8e38-0a8706c7f078 · outbound

This paper cites Program Synthesis with Large Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.558588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.558588Z digest=sha256:2789a64e6704a8f7195cc138f705d4950089e4d819b1c736fea763667064a7f2

Observation 9b3f9923-ebb4-4dff-955c-ad8be9d64559 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.610635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.610635Z digest=sha256:f9c09032bfa47cf4e7115476acd6a48aaf2c677092fd72dc0318145908e0d085

Observation 6b5d2924-ea14-4772-b4dc-cd1c2d13bccb · outbound

This paper cites MaskGIT: Masked generative image transformer.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MaskGIT: Masked generative image transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.581232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:50.698683Z digest=sha256:46213262f41a334f324be5eec1f1ac7dbb9f04cddf897ab738aaac6d0f7e4d4e

Observation 73efef0d-ae15-4a8c-ae56-c0b095bb9ebe · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.751073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.751073Z digest=sha256:0ab40479d792763f5c242b50ae25198d65879a27d61d04232dc02277733cedf8

Observation bf05aa7c-a1ea-4549-8892-0e238d00bbd1 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.825180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.825180Z digest=sha256:daafe96299cc0045265365b5b7e2232116ae2647f3f3b573bd08c77e9f076e8a

Observation 9fbdc1ef-c2a2-489d-88c9-3148afa46cd8 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:50.898025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:50.898025Z digest=sha256:91e46567087b4bdf74cadba7a5ddd58d1cb2f4e75c15f068259d073b2bb63f90

Observation ed3149c3-aceb-4027-bb4f-1ecaa2225239 · outbound

This paper cites DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.409929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:50.980864Z digest=sha256:b1cf1147099d73f0631c8663f6f5cd3413121159c0c6478b1fc13972144260d0

Observation b03dfef3-9e3c-4dc0-85f9-135553a7ca59 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.096255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.096255Z digest=sha256:aac53d4040e877bf1c391257d4c1df1e13c3ef2b350fd67f87373f3a151942c3

Observation 507abb59-e777-4a29-a576-b466f32e1ca6 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.104750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.104750Z digest=sha256:ed440c10499001f6a2e841a359e0a7e036314ae021058dc322159afc407e9ff2

Observation 9cbe2bcc-01ab-4e06-bdce-3af667445360 · outbound

This paper cites Video-R1: Reinforcing video reasoning in mllms, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Video-R1: Reinforcing video reasoning in mllms, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.209672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:51.147433Z digest=sha256:df6d5aae1839786d3572419d0b83b47a6915042f603da9a02ebce92b71ce9085

Observation c97fdc58-343b-4ef6-adff-f48df04e373a · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.197669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.197669Z digest=sha256:512ee2ba5fd0d7e08b53b11e1246f6fbed79a7a1cbdf674654b760044137d3b9

Observation d225a9ee-30c4-4bfb-b251-25ff1b66c382 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.296056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.296056Z digest=sha256:7ad22fb47dc90d7792b418000df45b079479410c957faee596593dca1069b06b

Observation 217e9aec-6b6f-42ec-b807-9c66c9954767 · outbound

This paper cites SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.415780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.415780Z digest=sha256:f1109ffc10c42adb75157b08f1239c73164261fe06bab25c1fb40997b03c1976

Observation 052310a6-09eb-4896-a7dd-72062cfa63af · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.520817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.520817Z digest=sha256:cfe2bd580ef8416b64a2c76eccffb3552839436b8644f01a920f1aeed4456654

Observation 2d65c360-5a19-443f-9765-18bcdeb960b2 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Measuring mathematical problem solving with the math dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.647139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.647139Z digest=sha256:169f857fa01407660f2c176fa0603160aa1c4dc336deea989e2540950ddf17d9

Observation 245ed011-309c-4e79-b538-9efa1570f074 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:56:00.020467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:51.759498Z digest=sha256:5d256115fcb4e1ba160b3997f8c4268d260cdbb8d94e3536db557f076613395a

Observation 972274fa-eef1-4abd-a1b6-0693320ef687 · outbound

This paper cites T2I-CompBench: A com- prehensive benchmark for open-world compositional text-to-image generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-CompBench: A com- prehensive benchmark for open-world compositional text-to-image generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.888512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:51.859745Z digest=sha256:4905c19fea719d8f12371ac6c4e759a595a5b99dea2a8652c31cb384612ea667

Observation d607f266-e310-41ae-8c12-c247a1e5c781 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.953592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.953592Z digest=sha256:efd62f23370805b0741bee10542e46f7a61647eb1aeb105e9320e279d081499d

Observation 6cae1717-2a14-4e02-8196-c49abafb20a8 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.031634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.031634Z digest=sha256:dbc926b065b9a764460e0c8a65a6f550afe07018de7b15df1336a674ad3cf23b

Observation 70ee7bc9-52fd-45dd-9250-797e0d1ace69 · outbound

This paper cites Large language models are zero-shot reasoners.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Large language models are zero-shot reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.179275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.179275Z digest=sha256:5089de0a6287c734cfbe3b5410ec4576a5760b325ae861b8e129ec81c59a6fa5

Observation 2be8be33-2dab-4e68-976c-8379b846a7ea · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.344176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.344176Z digest=sha256:ac6508509ff7999fe3e0615fa10ca3efac117ddffbd48c9ad3a4b88ff4709b51

Observation 69914c10-c486-40b0-a882-4fff1238d08c · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.466438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.466438Z digest=sha256:eb0c0ff196d8c74ee795af9bbc400ee9623a19deb3bcfeaa7b242c81115c2557

Observation 76765892-f7f3-4f0a-b521-9039ed32dcac · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.582353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.582353Z digest=sha256:2035ca7fa5975fe050b7ae34f26d22f29dd961a5f03f167741147a10a9a84f03

Observation 5c0dda25-23c1-44e5-9915-fb6ca1751178 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.695914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.695914Z digest=sha256:ab02ce2168249c7393b386523a28eefc4e854e265e7f883b0135501a803a11d2

Observation d6931dd5-8f76-4216-ba1f-aaaa4d0fd839 · outbound

This paper cites VideoChat-R1: Enhancing spatio-temporal perception via reinforce- ment fine-tuning, 2025.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing spatio-temporal perception via reinforce- ment fine-tuning, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.681942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:52.830543Z digest=sha256:84030a4fbafae987a19d5603a64fcd2d4e7ba4f95b7ca197459b28f1429695ec

Observation 768dd5fd-8570-4e7a-ad2c-dd089deafb31 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Cppo: Accelerating the training of group relative policy optimization-based reasoning models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.002220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.002220Z digest=sha256:9cb98893107c6c9155dba1d48af5e2d9d4d4ef05a014c8ea54828510df10fb9f

Observation afd2a74a-c2c3-4465-ba2e-b549029c4eea · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.166352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.166352Z digest=sha256:6395ce226559ec8972657b6bf7a5ba7aba0a165af56cac4972e12efc8277ba8f

Observation cd557c76-2f94-4dcb-a8aa-b6415f016d87 · outbound

This paper cites American invitational mathematics examination - aime.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO American invitational mathematics examination - aime

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.258103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.258103Z digest=sha256:7ad4980866c2f74fbc7073c770325126d7b8b34bccc98d29e3b0604ceb645325

Observation 6eddfe43-730f-4c0c-888d-ef2066bd308c · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:59.554579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:53.441712Z digest=sha256:daa8f725037dd1812f173d3180aeb75d4cb9c0612a70b5890ba9b23c52ca2d9f

Observation f8ffe76c-c587-4c75-be09-570240960d69 · outbound

This paper cites Hello gpt-4o.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Hello gpt-4o

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.582237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.582237Z digest=sha256:3000e052f8a712cee7223542783e733804cd2980c7d02b16036d82e6dc86d0a3

Observation c1e073c8-ab00-4c75-822b-d03d10e2ced7 · outbound

This paper cites OpenAI o1 system card, 2024.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO OpenAI o1 system card, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.407866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:53.699183Z digest=sha256:2b99963d6e7ccaf1812e73cccbb19692695cd9f25d2005bbab6eab8cfa8f2478

Observation 7805cc68-d011-4869-a0ab-e0839c948934 · outbound

This paper cites an unresolved cited work.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.844686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:53.844686Z digest=sha256:937e8b0bbadc04942fa812a8442a5c183f9e0be459771cb8b46d4739d0fe12d6

Observation f40e5c07-c350-4479-95d0-67056c5c6e2d · outbound

This paper cites Iterative reasoning preference optimization.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Iterative reasoning preference optimization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:59.219934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:53.982080Z digest=sha256:788cb8a3bdc8a707e30f6a89249507f0379ba6ecb669151e400ba54fcdd224db

Observation 493ee883-eac6-4594-b315-d8512b81770a · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.154340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.154340Z digest=sha256:0c6e70c93ffc68efc83823d85811f0e4f4c45dc7cba1c99e5ed635889787687b

Observation 996690e2-b131-4ce5-874d-2a7bc002187f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Direct preference optimization: Your language model is secretly a reward model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.311678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.311678Z digest=sha256:7a75d837dbc955d0e3abd088a64207c421d17291cc7d8a93bce5c3ce8fe35b1e

Observation 8ea8d8dc-3b12-4d22-97fa-e1300cf41597 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High- resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.426573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.426573Z digest=sha256:8e4f7d958834d7b2fed6aa0d14ce374a67e0e75a8e8abae8fa499f2147b75760

Observation 6246b52d-5954-485e-bd3d-2ebc214a4369 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.498153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.498153Z digest=sha256:77f1318d1986704ee7a87cfe9356c02f322db3304deedd5e71e6069283033dea

Observation dddb4f89-c7e4-43f6-92ee-f660d2653159 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.567116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.567116Z digest=sha256:6f7a89cb043258c0bc1b92b6ce7e3e53bdb6c64f6607cff299deae59fd891ad9

Observation 307c3867-8f6f-4974-92fc-8b3b1bc88cd8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.665299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.665299Z digest=sha256:0714f9a4938ec5662501db12a19a202a60439495c0d842d253de5f8f1a3491d4

Observation 68aa617b-2484-4b7d-8d6d-5518e7f48e6e · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.754970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.754970Z digest=sha256:3c971841c32ae3b09f0a02d271fb1f5cdd24ed07e35bef3b5d4a2a0dfc6c23ae

Observation 67ffbb2b-5402-417b-be92-d947d17bcc75 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.845570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.845570Z digest=sha256:a91a40c467e97e1d16997c08aab7f0ff47eaf569ef991d851ebb45d0cc128d52

Observation 12979f00-01d9-48cc-b3d3-742632bdd7ac · outbound

This paper cites LaMDA: Language models for dialog applications, 2022.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LaMDA: Language models for dialog applications, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.997167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:54.911937Z digest=sha256:35b89017acae7bfaa93a04b7fdc3dafd9b6c6678f16fa49fad048a749840a3fa

Observation 74e00793-1e6b-4c4e-bf61-f5b472092833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.983985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:54.983985Z digest=sha256:85ff66e8436602d92642a54f12c78ae304e27563049269b68127e52f4a904428

Observation 0609fef1-cc26-4865-a8ea-2ac1bf490271 · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.061900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.061900Z digest=sha256:6dccf9b3d4f36eb37c2ec96d655a86816601ddffa8f523dab443ed75f8452da9

Observation 1e187a1a-9e83-4952-b891-dad26567e6e2 · outbound

This paper cites Reasoning in conversation: Solving subjective tasks through dialogue simulation for large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Reasoning in conversation: Solving subjective tasks through dialogue simulation for large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.794682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:55.131047Z digest=sha256:5da04730546e2a2cd9e5abd48dbb56dc5a7f34e576cee66c89efed6173f66beb

Observation 5b9cc19e-30b8-4bc5-b80a-19cec2698a79 · outbound

This paper cites Le, Ed H.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Le, Ed H

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.573936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:55.222227Z digest=sha256:b50aea5ab5e171efcc80b649f77fcb83bb4dfbeb1348c71c800c6707ccac010d

Observation a7820c84-0c1f-4296-9edb-9640123c910d · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Unified Reward Model for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.299644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.299644Z digest=sha256:e0899b99c0f9ec908e04e1792cef97e3cf5717d84a7d1ed2421de28bb77895f7

Observation 4c0a8dfa-552e-48f1-b390-7b3b6fc9aa2f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Chain-of-thought prompting elicits reasoning in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.388496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.388496Z digest=sha256:e0e17e078c5d3071c7845f044bd230cb9439fc5372606a40a680dd8ba14fc43d

Observation 0c95efc2-6a09-4c19-a25d-aecb404deb1b · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.475472Z digest=sha256:f56924d40386aa169e5ac8244b30d74bd3bbdcf8fa371a2c90a02180c7df081c

Observation 9bbd3866-6666-4c54-bc94-565bacdcece0 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.543870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.543870Z digest=sha256:db3de5a3485bd3db47f93a083be2653bc0dda0f378766916e85f664ee5b6378f

Observation c0e88441-b730-427f-bee8-b976b37589ea · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.602449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.602449Z digest=sha256:419c3b28fdcfe34dea6c0470ede8ab1adc9375d093239a0120c6a90bd1af9407

Observation 5f2273de-bae4-489f-9dbd-b4131c80edd4 · outbound

This paper cites ImageReward: Learning and evaluating human preferences for text-to-image generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ImageReward: Learning and evaluating human preferences for text-to-image generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.664044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.664044Z digest=sha256:2e53078acbdfdb63d3a9dc9272c39e6a9db631e8e674c49f309a67379df76eca

Observation 2411a786-693a-4729-b382-0ad45550c05d · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.743601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.743601Z digest=sha256:852325e4857da575ebc23082a2ace48a51f65561913838294e440db65821a742

Observation 111ed8cb-763d-4fa0-8eda-bb83ab36e378 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DanceGRPO: Unleashing GRPO on Visual Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.854961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.854961Z digest=sha256:ded263cd59e9f32a30a07c103fe1d72d7ef6184770a6e6a1363ab56340a8bf5f

Observation 4ef71de6-d9f6-4576-8362-a8326dc81bc1 · outbound

This paper cites Qwen2 Technical Report.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Qwen2 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:55.929798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:55.929798Z digest=sha256:73cf46e12b099df582c5402f20a0b427d14765ec21f8e1501decd502a3949988

Observation d5c1b975-fb60-4432-b4ea-9fedb2b99cf2 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Vector-quantized Image Modeling with Improved VQGAN

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.042260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.042260Z digest=sha256:83860c689053a207e7daab9b2dd54d224903a65dbfccee6652ad6a7fbf1e448d

Observation 249c4d65-1e62-49b1-8920-e501a572bab4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.192183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.192183Z digest=sha256:a6b0af05e9528e948383d5f1f008d106d883ed4f810938bbad60537ddd2a5a86

Observation 88d1fbe9-1623-40cf-bcb3-e30f7da695b7 · outbound

This paper cites ReST-MCTS*: Llm self-training via process reward guided tree search.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO ReST-MCTS*: Llm self-training via process reward guided tree search

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.393634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:56.304989Z digest=sha256:2addb21a2bd281258655c835bc06535e838ac0188f1511491b35e5b6b3785d68

Observation 32fe3a93-ddcd-4ff6-9d99-efd0f5a22756 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Adding conditional control to text-to-image diffusion models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.411952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.411952Z digest=sha256:ba17ba0e379a506011cfbc146182e7b6fd2c61482fafc1a5efacd6d6e90601ff

Observation e1ee7f91-4b22-4b0f-9b66-5ce90750e7d0 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:58.222658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:56.497410Z digest=sha256:203118bd99caf8124dd90efd17a31b56691a3e9267e107c5de8fc35ab4d92590

Observation fd2712e2-69dc-46ff-b7c0-86a9b2a7b790 · outbound

This paper cites MathVerse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MathVerse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:57.995183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:56.612251Z digest=sha256:d5b1e87286d755a6af1e602cafc213184205a964e85fec9c71a644a69e328d18

Observation b529bdf1-551a-4362-aa4e-a75d6d4924ab · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.665148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:56.665148Z digest=sha256:ad4f922f11d1753e0d604746b46e9d7552f5d03a2806caeee961e518fcab9603

Observation feb97369-90bc-4b4f-8c97-12b7ce006be6 · outbound

This paper cites SafetyBench: Evaluating the safety of large language models.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO SafetyBench: Evaluating the safety of large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:57.728336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:55:56.714854Z digest=sha256:91ff4bb5eeb08dd5f225adbb2003f2dbf578f683c958ce48c16a324dcc5b68c3

Pith citing papers

Observation 215d5d21-18fc-4a75-a40d-6afad8aef6a6 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:52.913763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:52.913763Z digest=sha256:f0511b0212dedf65de662bd247412efcf61022420d2d07124cf048ec57958f25

Observation e5ba5c38-9473-4601-832d-f72cf168f5e3 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.937303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:1314bd44cfe4140c2c5ca75ec2be7fe03f15812c3c9779440dad9d12e09bb570

Observation f6720920-3ec0-48c8-ae9e-3d68f885258f · inbound

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE cites this paper.

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:27:50.062590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T13:27:50.031781Z digest=sha256:73f513662c845ee075ddc0d9651c110b2e9b934b05c28ef0235fb111b83fe698

Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.870519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.870519Z digest=sha256:82e675e38d31b0e107c6b84ebf5878c169f14f57783dbc232c6d11cb1c670156

Observation 56128ab7-bace-4d99-8095-329180615e9c · inbound

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning cites this paper.

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:31:50.190725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T20:31:28.371368Z digest=sha256:c6ed78c15fb8d8a507d11a20f915a21d075922c78c3dd6db4efa7b203ca93f51

Observation fa82dde3-e6ba-4cd9-9f27-3430c0e0db5f · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.397685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:95b16b03bacceae1354a65a41993920310bfd5b5cdc1c6c67820fbdd14266d64

Observation f372d45e-5e7b-40e5-94aa-87ed35887192 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.125091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.125091Z digest=sha256:85b87255d117073a68f598096d8580ef37ca3e4f80891dd762864c21fef03342

Observation d8fe2b4f-c45c-47d2-9105-f294673a21f1 · inbound

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation cites this paper.

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.149568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T12:50:13.764159Z digest=sha256:e8c39ad216de0d5269e0d5d6fd485d133eb0ea9fb95950229f92455b656696a2

Observation 9fd36647-1579-44e2-9743-98f1d1e369a0 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:18:17.195366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:b9a650244fa1cc2ca347a5f58a3efdb1784c9f1ab715bae3373b77e13f072538

Observation ab32f4e7-7889-451a-912b-01d5df6cd009 · inbound

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation cites this paper.

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:45:19.190774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T08:42:28.685159Z digest=sha256:b392c3733f5feacb80a689d50ca88f31e794854fb84e00b4ede9f57f7ba9c433

Observation d3492cb3-6d89-4e83-9cb4-6f17fc5ff95d · inbound

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation cites this paper.

HumorGen: Cognitive Synergy for Humor Generation in Large Language Models via Persona-Based Distillation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T22:22:41.696113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T22:22:41.696113Z digest=sha256:4a4bf28a026a8d01e6b21e01f27348a0e8b0f21ab519a9a53f2c5633f4de27ce

Observation 5476f687-e9dc-49c4-81ae-c0f378828bc3 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.824204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:443c476ecf290e40d4ce512f321f1dce8187c5ba2f983f57a22429c3d05d0702

Observation 7b1e81b3-7ddd-4e79-a312-382928ac04eb · inbound

Optimizing Visual Generative Models via Distribution-wise Rewards cites this paper.

Optimizing Visual Generative Models via Distribution-wise Rewards Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.692094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T16:39:12.711424Z digest=sha256:78a1e160b705775ebb9942d78a6e180ffe85369388ce3c6a7736813159aa3191