Pith. sign in

Paper Citation Record · LEDGER

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

As of 11 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 35 inbound Pith citation observations for arXiv:2501.13926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13926 v2

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:32:46.670055Z

measured 126 of 126 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:16:09.308555Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.175257Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c96e86b-f010-49e0-be11-62ba37de9d69 · outbound

This paper cites Let's Verify Step by Step.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Let's Verify Step by Step

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.099916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.099916Z digest=sha256:d00794eed81b1b29caed7489450ea295b4b7f438b7861563694ed21e479d29e2

Observation c2af53d6-7bb8-4a94-adc6-189188921310 · outbound

This paper cites Advances in Neural Information Processing Systems 35, 24824– 24837 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 35, 24824– 24837 (2022)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.107637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.107637Z digest=sha256:92b5efbb9dc5f50698a0ef00dae8e45bc7a7e27b436ff04e67bc178462559c7e

Observation 5b1264ac-5d5b-41e1-a0d0-73fdd1b5af73 · outbound

This paper cites an unresolved cited work.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.114479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.114479Z digest=sha256:1f33c9d90dfc9409420eb0f59cff73c7d6ca72459d159426dba4cd7b048dbd37

Observation 58823435-7a69-44d3-bdd5-ca5b88c71603 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.119385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.119385Z digest=sha256:f5f1b4b3009602ffffe3508144122c0c4d1f1e64c784f6559a2eafda5f2cc68c

Observation d715edc7-9054-449b-9b39-ba828e275aa0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.125647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.125647Z digest=sha256:05ed425de0f1225c042f3ddd4b5ce5d172f6f7234efe562755b69bbfc166a482

Observation fa23aeed-03d3-4601-8372-55be656dba4d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.131039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.131039Z digest=sha256:e1e93fe41167eff29468d1df303ad9ef61758d3912c3f4a58e3f4c944053c608

Observation 10ba5a3e-7c1b-4b64-b42c-67e984a7e63e · outbound

This paper cites : Language models are few-shot learners.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step : Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.136418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.136418Z digest=sha256:016b7cf496d6235bc062b0fe29e9afe441a68b1958aa0498d85b731681884ec6

Observation 4425fdaf-1360-4787-87e7-9bfec824f986 · outbound

This paper cites https://openai.com/research/ gpt-4v-system-card.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/research/ gpt-4v-system-card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.143412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.143412Z digest=sha256:216eb732ac20224e53a3e9ff457655efee72d406a4f51949865002058cedaa5e

Observation 326caaab-53ed-4ada-8b17-7716fc141ba0 · outbound

This paper cites In: The Twelfth International Con- ference on Learning Representations (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: The Twelfth International Con- ference on Learning Representations (2024)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.148880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.148880Z digest=sha256:cc9e468c6e11f096fe5e1c0e0733aeb6682a1b924fdf486cc8ecb73616e1bf97

Observation e4fa8aaa-6ff3-48a3-8e67-86dcbf5eb2e7 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.154470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.154470Z digest=sha256:1b0eccc1a3102540f562276bf1fd18ff6ae1c96fa6765df5ad6a3b132a762dc7

Observation 21664e14-e84f-41bc-9656-ddaaebd19ea2 · outbound

This paper cites https://github.com/open-compass/ opencompass (2023).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://github.com/open-compass/ opencompass (2023)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.159611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.159611Z digest=sha256:aafc679fb6069661ea5e81f726d167c813de8b03ee0b1466df1c4658c3c58929

Observation fe984375-daa3-48bd-8d2e-fb1f890cecf3 · outbound

This paper cites https://chat.openai.com (2023).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://chat.openai.com (2023)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.164810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.164810Z digest=sha256:d2b85b4074028be900e7394a17d2298226c256ad8a53d648abf5bbb529e024a1

Observation c6ae4786-356b-4f89-90e3-2a3cf7bc8fbe · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.170631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.170631Z digest=sha256:058fdcf8c1646bd9136469ffb9356a6faf8175530ef0dc5bf05ed587eea2d548

Observation dc0e5935-bd31-446b-9a11-3ed715c46f6f · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MMBench: Is Your Multi-modal Model an All-around Player?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.176609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.176609Z digest=sha256:0225f8cb59740d83f9cae14138fe7a2d39923cab8ef8952bf27afad10ec6dc97

Observation 47fdc94e-a1af-4865-b1cb-74b6ae89b56d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.182807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.182807Z digest=sha256:1f88f5399e0cbbde7efc5bd19bd88259d58ec134b7bbcfde6f4646b9f9f6f612

Observation 0041cbb0-3cf1-4b1f-9456-8f02d07b37e6 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ImageBind-LLM: Multi-modality Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.188877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.188877Z digest=sha256:efdd61cd6a44e94d84c1eaaa4ca765b1b2065917a384e82b2927e46fa2d0db01

Observation 6d5d7674-a9f3-4190-b90c-9bcdb95d3c78 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.195022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.195022Z digest=sha256:15e4fbeb6e78f58e057ba8eedc8f84268f72f839624092e035ffdd23147209f0

Observation d0d93621-ff99-4e46-ac45-d9913ce53f8c · outbound

This paper cites SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.200811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.200811Z digest=sha256:b2c7f55cc64ecb6cba729b1835657070bf69a440a57ae39cc640712c23f08d18

Observation 6a7eecf8-a246-4704-8e63-be92f9a99abf · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.207154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.207154Z digest=sha256:ee024b33e505ed2504b7e3aef4a524dfcfe62ffacd7face8aea30373a23696db

Observation d5ad7b38-04d4-4705-98d0-c893915a6955 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Multimodal Chain-of-Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.212962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.212962Z digest=sha256:2151b622c6204f1f6ce723066162993d4d5346d1a113c49c7ab4a2da341e5614

Observation 12f079a4-6739-44bd-878d-a4f3354393f9 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.219119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.219119Z digest=sha256:92f6e646478d2b3b3517e53be26c6dd6a91cea619dd1853fed2f502775e1cf6c

Observation fec5fb1f-0d80-48d3-9803-c6b332e0cb0c · outbound

This paper cites [Online].

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step [Online]

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.226490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.226490Z digest=sha256:8de380183f40a8c89a92391865637ff786fd549718166afc16b3c5fa4598d9f8

Observation 989c7960-31cd-46c0-9392-da436cf3c7b3 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.231628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.231628Z digest=sha256:420932c4518fd725ee449b1811fb3e03ec34097a762c78812509044fe8112741

Observation cbc0938e-0cfc-4806-bd53-3cbe8e4af1cf · outbound

This paper cites https://sciverse-cuhk.github.io (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://sciverse-cuhk.github.io (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.236682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.236682Z digest=sha256:bcaf90cf7b1e9b313c2168c658dae470adc167c5a990502341d4ebff4620fdeb

Observation 770f26cd-4e7b-420a-9302-2d62e8913be0 · outbound

This paper cites International Journal on Digital Libraries 23(3), 289–301 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step International Journal on Digital Libraries 23(3), 289–301 (2022)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.249799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.249799Z digest=sha256:ded72932a440eafc28f166afc52b7ad99ff02d0eb5ad22d9ab5859b3ca9168ac

Observation 3c790baa-3cfd-493c-894b-011cc57e521c · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.255826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.255826Z digest=sha256:51a3020f4858b33727bafe8d0889673c8b4b1a2ecd6bcaceda7261ca62f00400

Observation 38e3d0af-e775-40eb-85b1-2b3015565f22 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.261672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.261672Z digest=sha256:7f3397c88e21be805eebac8169836534e9dd22e5d8a9604e5a60636e07a3ab8c

Observation 83ab593a-2b36-4a53-99d5-cbc51f927e03 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.266755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.266755Z digest=sha256:b8ec7a565804bd518dd8341fd2657f2c31f9698bdf5f41e7878e1c0e0d976b39

Observation 1a2e9baf-e9e4-497e-83af-dcf948d9ed67 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.272390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.272390Z digest=sha256:b74e606ab0bcc470ef5a3c48b14e2fdd219b67a7b40998489e2f9cd67b6eedb1

Observation 84649c06-02eb-42c6-be3a-b353774fbe7c · outbound

This paper cites BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.277619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.277619Z digest=sha256:0bc96796099cecb3e7a188dff08488df6926f82947679334708f256467c3ce7d

Observation f4aab891-4530-402d-8c78-58c3679b272b · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Emu3: Next-Token Prediction is All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.283363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.283363Z digest=sha256:ff92528db0b90872780252bf35fbf947912fa06981fb59dc5db85037fd0ddae3

Observation fba80be6-eedd-461f-93ee-9997a6459062 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.289517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.289517Z digest=sha256:2e6be03b8dfa87e3cc2531f0a409b578b1c5a9c8a068ccd84558179192e744e0

Observation 107a212a-197f-4078-8c6e-a791668d6558 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.295785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.295785Z digest=sha256:e74034b3bfadc8768b0d0ce836f8644105e2ae7d60f676c762cd46674c5142e2

Observation 3d4b80b3-9e4a-4396-89de-ef5e58d877a8 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.169018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.300218Z digest=sha256:c819f0f1a4a5f4ec7d5a923b38f5d833e836b615ffe8630a3b97c34329b37382

Observation 5e8bef2b-c452-401f-8e51-814fde43fc4e · outbound

This paper cites Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.304329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.304329Z digest=sha256:e88aaf40458172fb1d7ff9ff4844edabe020ce5b681f70b2330c7cbdecd1cceb

Observation b3f98590-752b-423a-adb8-a0d733587b04 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Teaching Large Language Models to Reason with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.309279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.309279Z digest=sha256:ab93723ed7292cdd2aeabccb659dfd18967e5fce8f029d540118173f03070ddb

Observation be8edc1b-2d84-466a-aeca-f54a9f5890cb · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.315309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.315309Z digest=sha256:e29fc43f9bd0593651b131149e9a30c18adbe2f25ac9a56949b33295c27366f5

Observation 4f1427ae-6884-4a09-bc76-0d79c12e2649 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 36 (2024)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.149499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.319898Z digest=sha256:2a762b78bcc03108dade99ef1b17e795eece938fb688e09e2018a22fec439c33

Observation 18e7c0ae-9925-4d1b-a088-169b569ebbf9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.324481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.324481Z digest=sha256:786dbbbf6d4d2add0576bf8ac8bd9e7bc6fd9c488130c40fcd54b32885942f45

Observation 09ee2ccc-5b99-4313-9c42-3350cd6d931f · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.337555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.337555Z digest=sha256:57265ed7e5e914dc133ada737859df5466480193e547041782a9835fe78b32f9

Observation 73b2a888-cbde-4897-8090-6905b21c5e3a · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.345905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.345905Z digest=sha256:be74e3a31aa2fb4f22cf7aad6571df207572e74ef05994287814b29db4517a46

Observation bed49623-bcc8-4193-8992-de02fa411719 · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.351972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.351972Z digest=sha256:2a3734bfcdf04986e571cf9d52dc631a6d6c296e6316f265c66681f539322ac5

Observation 2ad8f5d2-9f3d-4e39-bd3b-bef64c88cc7e · outbound

This paper cites Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.358123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.358123Z digest=sha256:266547a3637bfca82561343462b82079b7e1028d1426aba097871595c657f3ad

Observation bc5eaa5d-9552-46bc-b1f1-ad47bd0a39a4 · outbound

This paper cites In: International Conference on Machine Learning, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: International Conference on Machine Learning, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.134016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.363676Z digest=sha256:0439ce380dfb98c79beff1590cb2723c8e24f7a8eaedea4685a7cb6ff5a4157e

Observation 64964ef5-ccb0-4393-9ac8-e7a758145b0f · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step AFlow: Automating Agentic Workflow Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.368693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.368693Z digest=sha256:5f5960b48812cf1fa26ed196b0cd06ffde44b2021c1a5c86040f4854484157fe

Observation 7ce4e38b-3bb0-459d-baa7-0fb85eb0556c · outbound

This paper cites ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.377436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.377436Z digest=sha256:5eb368eff8d37508a72ee5349df8cd15b1e7df7d4da8f75d009d0045885dd588

Observation 6f0fe4ad-7de2-45a4-bc9b-7cb90ad313d2 · outbound

This paper cites arXiv preprint arXiv:2409.01392 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step arXiv preprint arXiv:2409.01392 (2024)

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.383805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.383805Z digest=sha256:d7535446584d6310961231b68be90cc7cadd8bf00ba4a3d8a3a6de71aeb2ce27

Observation 29c6c530-c439-44ee-8e90-2d1218ebe5a9 · outbound

This paper cites Advances in neural infor- mation processing systems 35, 22199–22213 (2022).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural infor- mation processing systems 35, 22199–22213 (2022)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.119198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.389252Z digest=sha256:1665f5e24c951a3dafd747649873c202dc8b06d3b7e306778c0c62f5511e863d

Observation 245fd680-d181-4e80-8d3c-104f62301a70 · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Chain-of-Thought Reasoning Without Prompting

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.395722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.395722Z digest=sha256:830d785277ca9b53ee9e0ec39855191b13e1022cdaf1f58e7cea87874fc0fd63

Observation 9550f5ca-5715-461a-b489-d4be99ff68a2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Training Verifiers to Solve Math Word Problems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.401062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.401062Z digest=sha256:bc6f863ca645943e8714e92992414613046e8a531a0924ef7edfe458b32ff419

Observation abe25246-05e8-49ed-8ab2-8a362b467a93 · outbound

This paper cites https://openai.com/index/ learning-to-reason-with-llms/.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/index/ learning-to-reason-with-llms/

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.103117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.406595Z digest=sha256:bcf240c74ceb456b8af2c81b02559224d4e7752b2a075ab5901fd9b327ef0fbf

Observation 70124feb-1e7e-4d0b-b5b2-1f4228a692de · outbound

This paper cites Advances in neural information processing systems 30 (2017).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 30 (2017)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.086408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.413534Z digest=sha256:a7eaccb3ca9c95de1a0992f5506035bf6bc99988c735ef0a1cabe45764afc5a9

Observation 141cfdcd-5520-4890-8ba0-50dcfe5e650e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Constitutional AI: Harmlessness from AI Feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.419501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.419501Z digest=sha256:0f78c1dbe44f3bca5781fa10d7c01d437b30563e554534a53b4d8a2c08a02828

Observation aa0251be-08c1-431c-9df8-e8cee779af96 · outbound

This paper cites Robotics Research: Volume 1, 161–176 (2018).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Robotics Research: Volume 1, 161–176 (2018)

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.068969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.431765Z digest=sha256:f9144d735778e5f1478b23290e5b263197aaeec74dc3a0d8b2e764a74c8b74b4

Observation dd549d22-6eaf-4d51-b6c9-afd1c0b4758b · outbound

This paper cites Dueling RL: Reinforcement Learning with Trajectory Preferences.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.438395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.438395Z digest=sha256:f6981899e0efc52f77bf63f1527bd8d485173a1093390fe4cfdae2c6652d07bb

Observation 0bc97bd7-3bd9-4cc1-a4cc-ce7d25c1542d · outbound

This paper cites Advances in neural information processing systems 26 (2013).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 26 (2013)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.052863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.444302Z digest=sha256:424d2703777cd9112e0478111d896ada733996a5f62805559a90ba12a451b929

Observation fe872f6b-218b-4fbf-b8bb-544150c59223 · outbound

This paper cites Machine learning 97, 327–351 (2014).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Machine learning 97, 327–351 (2014)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.037129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.451017Z digest=sha256:c14592a956cf81cc7f021d06a939ee8b26518d021f3959a311f877840a1add2a

Observation 487da223-b5cd-4c8b-b54b-6a13d2a10d39 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.457214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.457214Z digest=sha256:a04afb011e6874627fecbd54ba34de781874f22c7bca97895d52a6992a5723e3

Observation a030b60d-d2cf-4e86-8157-8129fb437613 · outbound

This paper cites the method of paired comparisons.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step the method of paired comparisons

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:48.012302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.462981Z digest=sha256:30c001070ffcada39922b495b2808dd7f96ed0c829cd934cd01837929e42e08c

Observation f27b1868-52a3-4671-8ce8-c4b67b1c9b62 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Proximal Policy Optimization Algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.468470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.468470Z digest=sha256:5af45ea9460071d3b801f2afca39a90e1a5347a97a9c3786df5cfaf950d5136a

Observation 06a6ae0c-116b-420e-b4ca-318d7bb2800a · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in Neural Information Processing Systems 36 (2024)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.993699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.474113Z digest=sha256:d3eaab0bd2fe812ae2d8800b772e0769efaa907f994831a729c1c93867a79e5e

Observation d2d891ec-a4d2-468d-b4d0-980f3d974430 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.479044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.479044Z digest=sha256:5bacc846d2362c51a46645b00bbc17e131316ea431263883848443ea46dbaac7

Observation 1e07bc7b-13be-42f2-a374-0c0c02c58443 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.484359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.484359Z digest=sha256:7fc0c07ea4050ff396282ede097a9b0b35450c8b8fc47b55ed98943aad098820

Observation ee2a90fd-5316-4a53-beeb-f059953d7a95 · outbound

This paper cites Aligning CodeLLMs with Direct Preference Optimization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Aligning CodeLLMs with Direct Preference Optimization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.490134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.490134Z digest=sha256:47147a401e2b45438d0ae30512eca8eb1eb86a84ba0af780f3b4131d0b7b0b90

Observation 4bf62b0e-ded0-4ade-8c5b-5219cb248b25 · outbound

This paper cites Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.495881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.495881Z digest=sha256:9621c288a0d5de52ce5dc9eb0b0b2af947ec0c9aa8d73943d981714858fc542d

Observation 6d6e59f7-a3ac-4e3e-8d71-4650e0e8be58 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Improve Vision Language Model Chain-of-thought Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.501314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.501314Z digest=sha256:20a9a0e186f2866d2d44c8679648c77c868a198e749857c7fd9451834168ecc4

Observation 00e4ef6e-9293-4a0b-aa0c-5a4531ba478a · outbound

This paper cites https://github.com/meta-llama/llama3/ blob/main/MODEL CARD.md.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://github.com/meta-llama/llama3/ blob/main/MODEL CARD.md

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.975412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.506642Z digest=sha256:50e029622fcaee3570c3aeffcfd4b18c2ee7a72cae2fbb907dd854e175d11cff

Observation c85d5cab-dafa-44af-8aee-e51bfa7aad92 · outbound

This paper cites https://openai.com/ index/hello-gpt-4o/ (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step https://openai.com/ index/hello-gpt-4o/ (2024)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.958500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.511772Z digest=sha256:30fc7e3f2fc6c4a9e84c41ec23f3097b672e3ae21328e17bec0946074ef7ec31

Observation 9a87b054-f017-4058-a30f-7e36168e808c · outbound

This paper cites GPT-4 Technical Report.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step GPT-4 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.517639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.517639Z digest=sha256:603fb00e6037086206f7990e75589abe8be0ec1e7ab6e3a1e0d5b707a7d1cec2

Observation be0612c3-4cfe-41d9-bd7f-92ff9e86d395 · outbound

This paper cites In: International Conference on Machine Learn- ing, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: International Conference on Machine Learn- ing, pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.942121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.522937Z digest=sha256:417c84ca977cc071d6fa96209fc967af297fdb9d753859841cf5cc0b3ad16ff7

Observation 072dc709-42dc-48c9-9b4b-eafc430dfd08 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.527690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.527690Z digest=sha256:6ee6064bf6bef28c61265980da9c404af605ff8d46a4878c8e58220a675a1269

Observation 50d7ee21-67eb-47af-a8c7-c6240e8089d9 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.532426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.532426Z digest=sha256:f50d166b5f615bd8054d8eef338cb7f21ca1c0c406f039f6f95429ff2d55088e

Observation dfea0120-e7eb-4e71-bf51-1c91556f40bd · outbound

This paper cites In: The Twelfth International Con- ference on Learning Representations (2024).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: The Twelfth International Con- ference on Learning Representations (2024)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.925393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.538711Z digest=sha256:4aed090b820f0188f4fab91501b6e44f76d3caea8d27d47a8d76d9f28c54abdd

Observation 8a630b36-2620-440c-97a1-2abbec01dcf2 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.544049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.544049Z digest=sha256:4bd6f04ddeec6ec91da1ed69e0fc386d71e0b2ae00fb14692a0a0b5465135264

Observation ce03faab-d2cf-4f22-9a0b-445db9186bb3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.550031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.550031Z digest=sha256:c3819ae87514de9bf03f63bb4707179d60a5dcbb1891739f23a7dced012987fc

Observation e3af367b-9dda-49f8-88e6-d3f8d126f804 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.554937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.554937Z digest=sha256:bd63cc7f1458d40d584171a7086f0ba4941473a70ce929be0dd076ab2506c766

Observation 1f6435dc-bb0c-4923-9fe9-450e76224116 · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.561266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.561266Z digest=sha256:4b8d8bbb9d45b79228d0af935e88580332e15749df03d7678975fcc171dc6ac6

Observation 8ca7e623-3402-4938-8e5d-de2252e40115 · outbound

This paper cites Advances in neural information processing systems 30 (2017).

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Advances in neural information processing systems 30 (2017)

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.906998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.567076Z digest=sha256:df0d615f36e55237c9d8edadd7d26e1b5d555591856c4e2e113cff87f38e343e

Observation cd032d3d-c657-49ad-b66f-707e0762545e · outbound

This paper cites an unresolved cited work.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:32:47.890749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.572873Z digest=sha256:a2b6fe51a31127ed324585fb90959a5fb3e69a4d58bfb769410b371aa0d99682

Observation f935dd1b-0b4e-436a-9b2b-46624a795abb · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.580857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.580857Z digest=sha256:1ab1f93765a9919a8c8b84c0fad69d53b75bcbc35a0fb4d589aa50b988f03a77

Observation affa5a69-a4cb-415c-89b4-cca260047913 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:32:47.863859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:32:46.589733Z digest=sha256:e09ef1b4c974468e83073c85ed5f74906b4969762a63314aa286b84e42b5a6ec

Observation 4668aa73-4b97-4cd5-9c65-150ab26f9adb · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Autoregressive Image Generation without Vector Quantization

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.600985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.600985Z digest=sha256:c3a9bc950816092a8b85417506490d4ae826c219c8d8a710cf8b0b951d0ac947

Observation c164f945-7a93-4d79-9fd0-a2960160d058 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.617205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.617205Z digest=sha256:695205869abca0dc02c3fa4dd64b5b7d10a791a7fbd007d98d3e3336a5de4188

Observation 23f23d38-ebe5-4dc7-b7b4-7dd8d7d9b13e · outbound

This paper cites Iterative Reasoning Preference Optimization.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Iterative Reasoning Preference Optimization

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.622548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.622548Z digest=sha256:a69c3f5815c24b7713798f2f34829044c6f2dd57053642a6923eba348e253557

Observation ff62fa16-7ee7-4a7f-95ab-c2160dd40716 · outbound

This paper cites Qwen2 Technical Report.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Qwen2 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.629650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.629650Z digest=sha256:1e43ac8b3084e9d371597a10f761b52dfb30db75fc5e48d377b1dbb399b25bbc

Observation 4e40192f-9fe9-4985-93f9-d080ab23c06e · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step High-Resolution Image Synthesis with Latent Diffusion Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.637932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.637932Z digest=sha256:8f2f45ffc11e9a1750783f082c8db2bf2e928d2efcc6e15f98b0055c7f4f2669

Observation 5a690ba6-c7be-4081-b376-ac467472254e · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.646084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.646084Z digest=sha256:a0c55f6a2939fabc81604e1ade182ea99c7e4975ab9cfa751f70c7b78ed0b3c3

Observation 5c44f91b-121b-4bc5-a07f-7a3947dbb51d · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.654313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.654313Z digest=sha256:68770dc0e3ddc11b7e82f67deefab46eadada07353a3ac3568db2fda5cf4b10a

Observation f934ddae-9fec-4718-9936-829a27118b9a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.660605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.660605Z digest=sha256:5bbdbf87331854865450e0649aed78df78c358e9e4f4a9040bf360b68cbdf06d

Observation 1356bd90-dd7b-4f44-ac42-bb7ae7c803a0 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Training Language Models to Self-Correct via Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.665734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.665734Z digest=sha256:9d8325a9078cab5d18b891474205a5776fe8baa20c70d4684062d0423421d732

Observation 442640bb-bc79-4d78-8cda-1086fa7fb338 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Large Language Models Cannot Self-Correct Reasoning Yet

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.670055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.670055Z digest=sha256:1a2e3d0b60e58c4807a6c9cf5c331bfd3cd9cdff650c01584ddb4ba6bb8f78ad

Pith citing papers

Observation e42c13fd-978a-4186-a935-9deb61109a38 · inbound

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot cites this paper.

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:16:09.308555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:16:09.308555Z digest=sha256:18b6c1bc3316e247f60f2a71fbce2cb743082cedeed65ab06d4ac050830d4615

Observation 48f17002-9e99-4caf-acf4-ae4711e6c14c · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.448367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:c747c666f70e52b15410556796d962459a2eb6853ef5fe085227cab1d1fd37fd

Observation fba5499c-c290-4ca8-a6e6-37db20ab6514 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.651274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:b7e3b241e42f97c004403e382c5467ae88d98039d2e1efc2220bebcbf190f293

Observation c6f9f267-9f74-4ea7-9ee7-83052f813b8b · inbound

DanceGRPO: Unleashing GRPO on Visual Generation cites this paper.

DanceGRPO: Unleashing GRPO on Visual Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:28:26.031397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T22:28:24.929046Z digest=sha256:a4754ebc09813aede060ae6824586223e9384f13cb16ebcc9ce139e2062ffb1a

Observation 1fc057b5-f3db-4e60-b51b-77bf28e69907 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.137002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.137002Z digest=sha256:6156d5ac9c1a74e944772fe4d37f5cc69d078f5090cd3107acb00e9b45542249

Observation 052310a6-09eb-4896-a7dd-72062cfa63af · inbound

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO cites this paper.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:51.520817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:51.520817Z digest=sha256:a6a06b63e5a43c58951099da092b057193e6fa3d6ce65d18f56dfb2737abfc4f

Observation d2ea6fd7-5eff-4b53-b889-422edfe5faf5 · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:03.001071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:03.001071Z digest=sha256:19e65a3450502641c99fdfe275da95cc9a15f5ac54cb5a7573f2bd4f2c8a13d1

Observation 051c8f17-2752-4dd2-86b8-fe78be8377d2 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:13.310672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:13.310672Z digest=sha256:b791b6b9e5cd16cc12fc6d0bea203e2ded9acd8cce4602a33a697b488a41f8ed

Observation 5c976fdc-86e1-4356-bcef-5e81b52d8b8e · inbound

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation cites this paper.

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:36.335219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:36.335219Z digest=sha256:035c5790cc249fa1e3611e65fc64abb37cde206140a5ce1efe4e158f6766e3e7

Observation 3e8053ea-4c2d-445e-aebf-e7b2b1ee8a09 · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.895103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.895103Z digest=sha256:dc9ae53c79243816f3e179a024bba3f1fd2b5db451a290deaae143750eaa5ef1

Observation c30ed327-5349-46e7-b27a-bb7d9de92dad · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.044539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.044539Z digest=sha256:2b025f322917c911c77bab29f73b0db265a84f491487754d05601a631f1def3e

Observation 8a0adae8-f90e-48c0-a83b-bca2917ef604 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:45.740034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:45.740034Z digest=sha256:8fec23005958ab267ab7c0802a6f847a9e3ffed5106a1ad7a64714ccf3bb710e

Observation b3f0c90b-2bef-45b1-972f-8a68e39d3793 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:32.622198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:32.622198Z digest=sha256:a5f1aa82890faeeffa8f5ce4c0acae03b1e9db0ed0c44abf51bf34779f6c2113

Observation 081cfe8c-5871-4cda-a8be-8a5bfe6693bf · inbound

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL cites this paper.

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:21.683953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:21.683953Z digest=sha256:9b43adbee907cf8cfa06e772c574b3c539c144c51e34d71b65e62d979a5f03e1

Observation 553d2b69-b567-4d94-9e80-6517d6fa6cf0 · inbound

TIIF-Bench: How Does Your T2I Model Follow Your Instructions? cites this paper.

TIIF-Bench: How Does Your T2I Model Follow Your Instructions? Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:20.782033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:20.782033Z digest=sha256:61d485d952557656615f5a58694a2d13a918cd5f51ff32adfbd26e20d61c9622

Observation f932e840-318a-45e1-9f6f-875b17f5d0a0 · inbound

How Far Are We from Generating Missing Modalities with Foundation Models? cites this paper.

How Far Are We from Generating Missing Modalities with Foundation Models? Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.584502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T08:15:12.947854Z digest=sha256:e4a82b210083bb90cd52c139b61b11997c3c232f775f30e110ef978d583380f8

Observation d4c71618-6fcb-43ad-bb39-3eb3b7ef5313 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.738793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.738793Z digest=sha256:3ddc81612139997770a10968991fb6d7c3e9ebacecba3ad28d5feb4fe46af46d

Observation 23dca6e3-ebf9-46de-8494-cb092f44b9b5 · inbound

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment cites this paper.

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:44.051500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:44.051500Z digest=sha256:32f7c066f88f3f32890ed3d145139a43b6fcfc4ea6e96d0751ac6d2625de2b2d

Observation 96b75dea-f716-4ad6-8b6c-2b525eb43572 · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.511095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.511095Z digest=sha256:641363d9d7ba5294d1540b240315fc3097d3b2c9b639fa657cfe8f83d6e02899

Observation 294919e4-d9e3-4c01-9070-ad012af279f8 · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:20.308003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:20.308003Z digest=sha256:5aba597f3608798d44448786aa2ac51dbeb990a924626392082f115e27f01abb

Observation 26c2409b-046a-4e3e-87f2-a4338ad6850b · inbound

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought cites this paper.

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:33:02.188477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T08:32:20.566798Z digest=sha256:90399dba7a26a866b2c354a4170e51cf61de8e8cbe36dfbc47d2443aeacf7402

Observation b8fadb8b-3881-4257-9666-e71df3074dc1 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.744905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:8e5bf2dad8ae5e74cda77307d7e90df1018d6a6e210efb0a81aabd78c9c475d1

Observation 1255e509-c867-452e-a7d0-a6c4e527b6b2 · inbound

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation cites this paper.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:42.906185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:42.906185Z digest=sha256:529b492dd99cb74eecc94ab89fc49b53b76ef750b58d2afcc0ddc1c14e2dabf5

Observation e73c98f9-504a-442d-99f6-92ef727bdfd2 · inbound

MultiRef: Controllable Image Generation with Multiple Visual References cites this paper.

MultiRef: Controllable Image Generation with Multiple Visual References Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:32:51.804152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:32:51.804152Z digest=sha256:35f4f38b88ab5e909fce474fcb644f165c2c8fbac7ba0f82a7170e91d54ec3cf

Observation 1c41b22a-8423-4dd8-9ff9-b14b3a0b11d8 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.432203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.432203Z digest=sha256:b4af6d7c81d4a2c4d092d2ec8ab6bf82064aca9e879711cc598822cb809360f2

Observation c7d72a8c-5a3b-47d9-8018-1bc7d72d3564 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:07.963130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:07.963130Z digest=sha256:2e03a637f0ca5b826a337b478c9f9ad529baa10b30700313e13210afdf8e7bbb

Observation 23d37186-de7a-4292-9018-bf899153e531 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.343091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:b5740d3151873fbb0c5c0b264bb2bed6eab285ab85205d939455d294f812a071

Observation d3c4c922-b22d-4177-b0c6-590127eb25af · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:5d4eac5f977816cc5db7ad5eaecdb186933ce84f6050797d18d58579ce7d7748

Observation 17c05c7b-cb86-4bb6-a657-19a2bcbeb692 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:54.197765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:54.197765Z digest=sha256:5bd661460c180cc00b77cfeeb5e42df778c4d5025aaf9b292282aac7cff0a74b

Observation 53b29446-6ab5-4120-bea8-1451dcd9a279 · inbound

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation cites this paper.

From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.069714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T12:50:13.764159Z digest=sha256:67317bdb5ea01574f48b1bb8106d8d9d9512f379d042296512a15fac88677f5a

Observation aafb7341-2c9b-4aa3-a2c7-b6e392123764 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.520773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:3b85c84736a6e7f13ac46ed7b4c1b9c4a5b2ee539b20d1fe031359807f1541d5

Observation 1ea53590-1d6f-4963-b40a-5219ead2ee88 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.295363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:c7ddede32a4b1d095e3cf03924cc8f4fca33a81119cb1b3291f736cc9c0b7612

Observation e823c830-b45d-49d7-92a8-b36f8676a353 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.837081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:621089f77989a750f54f1bc80cd20ed623beb9222fa2ed8cb07b427f516bd592

Observation fdf1b1a6-3013-4842-8d7f-e210294a46c2 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.177090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:13c33c9bd114c0c2be6a663de5e670880bf57ad259f5623f01e2737af34ec5a2

Observation 1730a604-eb8c-49ea-9e3f-50043610d767 · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:c26227b23c2ab628273272f7895a3723a2f82af5454b8adef65921894f1696e5