Pith. sign in

Paper Citation Record · LEDGER

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.21391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21391 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:54:38.274595Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 056a1e6a-beb3-4817-b201-201c4fffe76b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.526329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.526329Z digest=sha256:29cc7cf9448043139affc2c171012d82631ddacd306145f9fe695d545168abb8

Observation 0eb2ab8a-9a25-4c24-aadd-539f4e16e3a0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.634149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.634149Z digest=sha256:2e016fc1cbff623615c64957ff896e8057bbcc2550e2bce3057099cb9d1eb53a

Observation 1414ac53-eb29-4751-a06b-8f3f69844a80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.751253Z digest=sha256:b7b45b6eed733e0b569d69f94361ee400a2a2cfc72fa0ae3ad224889663132d7

Observation b0e9eb8b-c2f4-4e4f-9905-3f8dce5d3266 · outbound

This paper cites A Note on the Inception Score.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A Note on the Inception Score

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.829262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.829262Z digest=sha256:78df2a1db438329471588e610c3f89d9827659af496daba9e98b6ea7f1757cd0

Observation b4616089-20a5-497c-8a4a-67aef162a28a · outbound

This paper cites Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.897882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.897882Z digest=sha256:fce2026e9784642b16269f4e58753ec3fdda62a81fc53a1d891c349372819221

Observation 01a57abf-d454-4edd-97d9-20a2f7402118 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training Diffusion Models with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.955610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.955610Z digest=sha256:9f3879dd51eac85778bc5280ede058959977d44a8e52fe936867d911d8a2541a

Observation 2c9bb625-c517-49c5-a40f-a8c3e4d0db53 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rank analysis of incomplete block designs: I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.693809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:33.077318Z digest=sha256:e8c4b6f530190fa85462c0a0bd3f315f7c0beb17194087de4336f3be69a2c5b7

Observation bb5a2d9f-8017-4cdd-8ae8-61688734c0a3 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.162969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.162969Z digest=sha256:5fe4e89793c8566be4a3f7f3471808c7b411e76f7c3da6eff9989627ccd74bde

Observation bed65cea-25c8-4a87-a108-28b39587f9d8 · outbound

This paper cites SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.250868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.250868Z digest=sha256:2e045c7b375bef586898f60c2571f6f8eec434aa722051ff0fe31da46d7ae8c0

Observation 02c55c24-4378-4579-8a74-f34efdfec3c2 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.346700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.346700Z digest=sha256:7bb38447a9ccf5c7cb937771368b8f780787ef5f7ef3c507398d5f5203444b19

Observation 8f92e954-ec8f-447b-a014-6cb2464831dd · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.452384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.452384Z digest=sha256:357a4a34f04baeb7b324a20263b8fb07c2857fa18ce95c24a7c852bf0b3f75da

Observation fe32db23-e6ac-4c04-a95c-a2c6e99d02ca · outbound

This paper cites The socio-moral image database (smid),.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation The socio-moral image database (smid),

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.450201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:33.551927Z digest=sha256:6cd61eb95bdbc0b99f098425d01a6bc69dfb1fc7354bab8c9358f8014df07539

Observation 62b1b304-f88a-48e3-a114-a693b7d5c173 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.643754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.643754Z digest=sha256:2659905dd3e89547fed5219f14ab8b6729ba68e70c8b831c968a54eb8af56eda

Observation 9f9e2686-3aa2-4027-9f15-d3df3ff8be4f · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.744321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.744321Z digest=sha256:eb9c669604f30760880c2dc116fc2072a274d753b6c5d927f0334df8f340d0fe

Observation 78af6481-780c-46fc-b0f7-791cee9478e1 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Diffusion models beat gans on image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.829967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.829967Z digest=sha256:4bfd2c52adaacdb5ec9552b6b0fec4aeb6dc2b472dc61df1d6a1054702796b88

Observation c1490d63-88ed-496b-aa9e-9014014c8769 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.159261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:33.924185Z digest=sha256:96511a125b7a8124ff2c7e6adb4447a21c2697b128e30a9ae414893ce17ed117

Observation 52916587-4e53-4a68-838e-944fbb9d03f4 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.900852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:34.027720Z digest=sha256:feb4edd5866a4a543e6675022445c205fc4ba7281a47dd04a68ee79c2fe8562c

Observation 34de7a70-025e-44a6-9eb6-2c9eec8e2305 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.638002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:34.106894Z digest=sha256:8757e0fe03a353ed4f853d114f1ee92b5b44e6201e8e9727586ac2092b8c01f5

Observation 60ca7fe7-16c2-4966-bc17-8305b4f4f500 · outbound

This paper cites Generative adversarial networks.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Generative adversarial networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.359411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:34.174919Z digest=sha256:ff521e7e077f1581c5a7174129c40535faeee6035faafc028117ba82fef46d0d

Observation d8979830-9d36-44d0-be4c-f66261f68e20 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.226626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.226626Z digest=sha256:eec52fa5dc9497325e5f4f5ad39da1955c70a997d23310c7cd7095f1cacb74eb

Observation 93f16588-f141-461c-beb4-88fdcbfe6f57 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.274450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.274450Z digest=sha256:4ae5ab7360eebaccfad338fa0ef7841bacddec28a31a8b1ca4387140b0323d3f

Observation 221cbe1b-c39c-48fa-b1f5-12875ac86735 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.347454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.347454Z digest=sha256:f4862bcaaf88c8a13f672362856bd91f7671d2428b0676ed34668e367a795218

Observation 175a9c86-8bea-49fc-b5ff-ec2393d1ae23 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.455883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.455883Z digest=sha256:1ecd53ba0878de879de64c6323849f7f51efed28ea91f9bc07304b9c0f909feb

Observation c6bf3fea-02f6-44eb-904f-15421c7d9897 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.578707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.578707Z digest=sha256:660ef06098ada0b39c3cb82f61b2ef58d88ada36aa9b2f8606ea1c49aa1654a3

Observation 848fe526-4132-4bf4-9e22-842a315f02bf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.709673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.709673Z digest=sha256:bd24d4a5888754f17b32d144a976cd99af8e02b4ec053693034a26caf1a6f0b2

Observation 0bc91a6b-bbdc-4591-8105-4bfd5f909bf7 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:34.861279Z digest=sha256:7c73a7b67de1796791848fab0cf34af931d125a7ff8f0d00c171aa7731d08c54

Observation 4bf251e1-f89a-4031-8745-a2ac7ded0574 · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.950914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.950914Z digest=sha256:88762d58c906da7f4fa42814edda006df28dbf075912c52d892344a35a427d3b

Observation d72d80f3-ba3e-4c4d-9c51-5d707ddf2eb9 · outbound

This paper cites What matters when building vision-language models?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation What matters when building vision-language models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.047298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.047298Z digest=sha256:232cce1c3b2ca192475a71323b5977cebe4ffcd7a6f9c64b2233612d34bed3c9

Observation 71df43e3-76ce-4791-a801-b71882b18abe · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.807225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.145183Z digest=sha256:2c34515828f863d48e3ec98745419318d81bc228c985387e88520ac8545352aa

Observation bc9aae33-9a27-401b-b7c8-be632420679e · outbound

This paper cites T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.249132Z digest=sha256:ccd2602923e5d49bd8cb0ed7002a39369ad07a87550e66346823fd9f9bc0d6de

Observation 94664539-1edf-4155-acf0-6690ccf5666b · outbound

This paper cites Remov- ing distributional discrepancies in captions improves image- text alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Remov- ing distributional discrepancies in captions improves image- text alignment

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.539329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.362579Z digest=sha256:e2aae7865adc6b2a2ced474db6aedb7f1aa12096187853c0f94871c107763834

Observation 337ed284-26b4-4f43-b34a-095d129c1003 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.271746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.442644Z digest=sha256:eaf84ba649096be1074bf71043c82964f0d51ec26680eff2060e26039492b8f2

Observation 0b495463-d938-487a-836d-29917f2f9c97 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.526604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.526604Z digest=sha256:15e49b5ef5e3340763e4b8271facbd1ca92c3cf47889438a77ddcb044bea6681

Observation 485cd105-58bf-4858-b115-6c11111da2f0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Improved baselines with visual instruction tuning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.023960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.620233Z digest=sha256:ea50d10bff4d914b85f10dc99de179fd748f83bfff06c8315ec3a2aafd99c1f4

Observation 0750e59b-d266-4b5c-be87-abf663dab75c · outbound

This paper cites Visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Visual instruction tuning, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.790950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.676524Z digest=sha256:ea5d52ddce24c9edab2d40725a3c5234e67c52abf2e8e07ee9cd5af12842c298

Observation e18fadd7-3903-415d-8b9c-58d5e539e4c0 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.529889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.725893Z digest=sha256:961d87271a3ddc8cd29d599de648bc52d51e976e666aab37b9202d9a93301f22

Observation d4a9139f-d60a-4b32-a7b7-4b4aaa609763 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.785004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.785004Z digest=sha256:eddf74b5caf533f072264b01354394a47a4016670e370d864a5ec179c4c37beb

Observation af03050d-fe22-4c9c-86de-850fb1ef8d9a · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.880961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.880961Z digest=sha256:ba583bc548c3d5f5c62738ec1cc8922ba3ae58552d1cfa17508d372df79d546d

Observation 90215b59-d14e-4456-8c68-0e470e90ac48 · outbound

This paper cites Simpo: Sim- ple preference optimization with a reference-free reward.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Simpo: Sim- ple preference optimization with a reference-free reward

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.318338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:35.929519Z digest=sha256:10507878f42cb46a359fa71f4ae3476e940ec6d51958e975216ea276ec86f5b2

Observation 1d91967a-7d25-4f32-86be-571db2738094 · outbound

This paper cites Training language models to follow instructions with human feedback.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training language models to follow instructions with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.985916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.985916Z digest=sha256:6ae350808e53d5a516d9f114ef44b1ef2f185d8d3bf13c1b8718be9088411e58

Observation 7dbeb8a1-6a9a-4e17-bca3-1bd6e508fc90 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.083635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.083635Z digest=sha256:c0de04d534addcdf64d6a0a8d35bea2488a542bcb18d42b53f49fa0c2a28fe96

Observation 19a5864c-3e88-4d4d-8c3a-5565a4ec0ab4 · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.061472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.177621Z digest=sha256:4928f67e8161e505608b37cb12828bfb44868c0f0d7e049b187d53372d961ada

Observation 486054fb-1c29-405f-a1e9-74cbe132e51b · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.257427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.257427Z digest=sha256:cdcb2a86f4a042256f61162fe25d83aec97e0ee284585397ffe3f5b788a806a3

Observation 498d6646-c498-490b-95c6-c800bcb74c8e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.362963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.362963Z digest=sha256:1246b247d589956b57429947db5d7182934bdcc720380ef4b475c3a444189422

Observation 95dc7cc0-f7d0-40c6-b4f4-aca3578260c1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.824402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.445317Z digest=sha256:d74058dd03cf796b1954105e05852a4cd89286819d5dfee418f283791642c104

Observation 607b8cf6-d085-4e66-8f7e-74768b3fca54 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.561887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.562776Z digest=sha256:89719ace668a75bb8af8eccf328f0294cf147d2751f1d6f4a837701db1238df2

Observation dc7536f8-7fcf-463d-90d7-e39f42ce4053 · outbound

This paper cites Zero-shot text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Zero-shot text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.280839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.684650Z digest=sha256:252c5b943abb4b92a6d846d2c09342c185fff150b8384d4267b7c7c2da21e9c1

Observation c52fbfde-6f70-4c49-9cd8-36c5ef12c567 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.038981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.771953Z digest=sha256:08624a1064d3568efe9c065924f3ddfc27b91d0910b43d899ffda4c9315891a1

Observation 4d7867f7-2ff2-42e1-8576-59eb0668b5c0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.835962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:36.863239Z digest=sha256:e86b2e6f0fa2114188224e2a9dac5cc852d2ebd4b25af1ea72e12d15209e8bef

Observation f0c300b8-ddfb-44bf-b455-1d9f0393f153 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.914296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.914296Z digest=sha256:f40dc726de827cb84a4b28088156aa521e66b29f75b3f7779feb949ddf50f271

Observation 3ce9b861-b945-4995-ba66-1d41e41b7056 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Proximal Policy Optimization Algorithms

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.988698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.988698Z digest=sha256:be94a816bac068e68f5bf267e49845aba97ab77d17b7ce952f521c8bbf3ef691

Observation a326a054-deaf-4fbf-8f2f-962890284bd8 · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.089978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.089978Z digest=sha256:85605a274a2016e1b58dea442b4a0ef47cb5c3ab704d5160a7302b750cff7133

Observation 0fb5a6c6-185f-4428-8657-6b08bb29b32c · outbound

This paper cites an unresolved cited work.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:54:39.626966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.170906Z digest=sha256:d0ef41962d38d5e246bfd925561a70919c91bea2885422e28802dc2c7925f6d5

Observation 04f9475e-a2e7-4fcf-acdb-d64cd9abb94f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.484015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.288393Z digest=sha256:a3a26c6ff262a50433abce7bc9938b6620101031c16d3dc4198890a21975c0b3

Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · outbound

This paper cites Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.374599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.374599Z digest=sha256:d18150d0a7082f34feb4c744c5516b78ca45796b35d4c76f9ba9c7e3051b2199

Observation ee6f6be9-143b-4030-8a6e-1d5888fb3122 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.474204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.474204Z digest=sha256:1099a97062a7527012620445facd66569511a410082519fa59f70b6e4d62cc91

Observation cf014844-e467-460e-99a7-964de4cccaef · outbound

This paper cites Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:54:38.472533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.533704Z digest=sha256:7a3bac9971f6253b60c277ddcecdaff50c9b587bd1e74c35df638340a4fcb0ba

Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.619519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.619519Z digest=sha256:87dbdbfd273df1c5a5508a4313a721e574989467b6ba219b5573eec5d5d1e35e

Observation 0649d742-b68a-404e-b33f-82115b5ade71 · outbound

This paper cites MLLM-as-a-Judge for Image Safety without Human Labeling.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MLLM-as-a-Judge for Image Safety without Human Labeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.686456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.686456Z digest=sha256:c9028f2aa9fbc8b614f53f2d3235ba455ea029ae95b714e992ca7bc930aca5b8

Observation 47994f55-4f79-43be-96c8-6f0b8b3df90c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.787465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.787465Z digest=sha256:068bb9b9cb63539287a0bcecf04a9c644dd394ea385f34fc50cf6f1e7c63c79f

Observation 55c643a1-c5da-4050-9945-d9ba9e50d216 · outbound

This paper cites Human preference score: Better aligning text- to-image models with human preference.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Human preference score: Better aligning text- to-image models with human preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.344930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.859469Z digest=sha256:5fca24699cd223e8a08aa3d6eb120861837a224d7d1c24f04b2820cb32e5bd65

Observation 75506a12-6c51-49bb-ab43-87fb31752d09 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.261401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.927082Z digest=sha256:534312c30146ab0d5b9a5097c68f6f187cb699a1fe9bcf0dea366e7f89e1d77b

Observation 3eaeffba-521a-4480-85fb-6f093aab3080 · outbound

This paper cites A normalized levenshtein distance metric.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A normalized levenshtein distance metric

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.129080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:37.998771Z digest=sha256:5fe4ff832cdc3cbc13765f612202baff9a02c9f94313e60166088a70664b4dd2

Observation 3f3acc0d-51d8-422a-8f4f-e5647e7a1e6a · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.984919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:38.090385Z digest=sha256:2f53bb149e07f63340ee4b6d2c0c539fe98ea38561cfb617fd74822b48bb0bc9

Observation a5e53093-eb1b-4ff2-9218-0079c8c509a9 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.169486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.169486Z digest=sha256:09f82900c55e3c8c76d0e5b20126202018031495840150f0acae37e89c380b4b

Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.229729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.229729Z digest=sha256:bb08b390120e8f3bbfe5aced7a3b29aa92e8a08e5a0a3adce2be69beff69635a

Observation cb6d2b9e-4aab-478f-ba2d-91070b9da082 · outbound

This paper cites Towards language-free training for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Towards language-free training for text-to-image generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.832890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T12:54:38.274595Z digest=sha256:d4d360f4ae27f6d2dbf4c396cc7c001d2b6ed3409c395b5b660b498d851df7ed

Pith citing papers

No inbound Pith citation observations are available.