Pith. sign in

Paper Citation Record · LEDGER

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.21391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21391 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:54:38.274595Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 056a1e6a-beb3-4817-b201-201c4fffe76b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.526329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.526329Z digest=sha256:d5bcaaaa833552d051db584a79b4fe154e7d1e4260a240deeab8c446ae1eb2d2

Observation 0eb2ab8a-9a25-4c24-aadd-539f4e16e3a0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.634149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.634149Z digest=sha256:2f8d2886bfa82a1e513d661bd6744e9b16785e092cd2998f9864452a548d3f6a

Observation 1414ac53-eb29-4751-a06b-8f3f69844a80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.751253Z digest=sha256:aeea51b7122eef6ed14de055b757cbf43e4541bfe67fde051bba3638aeb70652

Observation b0e9eb8b-c2f4-4e4f-9905-3f8dce5d3266 · outbound

This paper cites A Note on the Inception Score.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A Note on the Inception Score

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.829262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.829262Z digest=sha256:f29b94fa9bf80b145f6e531e1bc282d108c8310f7d4f782e61f0efb75dc131f5

Observation b4616089-20a5-497c-8a4a-67aef162a28a · outbound

This paper cites Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.897882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.897882Z digest=sha256:6261745825655a7a12ae80ec0a44a058f936756df9bbb186a41bf29e377208e4

Observation 01a57abf-d454-4edd-97d9-20a2f7402118 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training Diffusion Models with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.955610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.955610Z digest=sha256:da1165ad44e53c0f5f9bc6a1592b40cf8b88dc5fd8ba14234ff47e039d1a4714

Observation 2c9bb625-c517-49c5-a40f-a8c3e4d0db53 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rank analysis of incomplete block designs: I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.693809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:33.077318Z digest=sha256:fea9d0191159eb7a41c95a0b839b9de41b9b5f84ac2874e3c50551141e53eff7

Observation bb5a2d9f-8017-4cdd-8ae8-61688734c0a3 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.162969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.162969Z digest=sha256:e2398baca7caf8d1f9607bc9c5e51e0a8984d0344358f138686ef3e88fecc780

Observation bed65cea-25c8-4a87-a108-28b39587f9d8 · outbound

This paper cites SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.250868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.250868Z digest=sha256:391241b0a1a8678c3055708433cfc7edeab7bc7a41d44bc19e9645d00a048600

Observation 02c55c24-4378-4579-8a74-f34efdfec3c2 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.346700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.346700Z digest=sha256:1c52795f70038c2d7941c315f8fda5b205310439d310f7dc9ebe7fbd58b5470e

Observation 8f92e954-ec8f-447b-a014-6cb2464831dd · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.452384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.452384Z digest=sha256:b463a1a78c7fe5ac90ec6a3a2999b32b89d11bc97c375bd868f4bb74333f5271

Observation fe32db23-e6ac-4c04-a95c-a2c6e99d02ca · outbound

This paper cites The socio-moral image database (smid),.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation The socio-moral image database (smid),

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.450201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:33.551927Z digest=sha256:6c398c2504b14bdbfd96276de9f1a87ec66904f9485f3f024a420cf5e747ec1d

Observation 62b1b304-f88a-48e3-a114-a693b7d5c173 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.643754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.643754Z digest=sha256:ee26b9157b662791166c18e391fdd46999c761118550388de5e89cbb2600ad59

Observation 9f9e2686-3aa2-4027-9f15-d3df3ff8be4f · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.744321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.744321Z digest=sha256:9322d058d3e9d211aa5feac45c1133e28c6380e5cae785dbb17a39c0acdc9445

Observation 78af6481-780c-46fc-b0f7-791cee9478e1 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Diffusion models beat gans on image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.829967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.829967Z digest=sha256:65d20226cb4e331c48a83eb31f13f00238aa0c5347762b33087269290e28b35c

Observation c1490d63-88ed-496b-aa9e-9014014c8769 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.159261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:33.924185Z digest=sha256:ae639e4b014fbf5494b5af365d2c552a99d2635db4eed1e6ea536f0c6262701b

Observation 52916587-4e53-4a68-838e-944fbb9d03f4 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.900852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:34.027720Z digest=sha256:a80783dfafe11ede921e38297f3378018897acfcb0299a587b3d6974f2645599

Observation 34de7a70-025e-44a6-9eb6-2c9eec8e2305 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.638002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:34.106894Z digest=sha256:669d880cb964c526f9a94786bcc7e5f4e37425ad124778cec4c8832dfbaa9269

Observation 60ca7fe7-16c2-4966-bc17-8305b4f4f500 · outbound

This paper cites Generative adversarial networks.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Generative adversarial networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.359411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:34.174919Z digest=sha256:8c31cae85c8b5b3e0576f961572a27bc8a588073e6e67b509a41f84b65d3c041

Observation d8979830-9d36-44d0-be4c-f66261f68e20 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.226626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.226626Z digest=sha256:24917478ab7e92dcc4c66c103cb2a468f23deb5873c9082e98beb272f2d7a1f8

Observation 93f16588-f141-461c-beb4-88fdcbfe6f57 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.274450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.274450Z digest=sha256:a27366b80e39ab798781bdf6d5d0d4a41c5ca8c303376a41b7e1fb43dbecf315

Observation 221cbe1b-c39c-48fa-b1f5-12875ac86735 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.347454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.347454Z digest=sha256:a8c10c8a7d1b428282c1a19fe0f77155b31e557cc526deb9e3bd8cf64508be7b

Observation 175a9c86-8bea-49fc-b5ff-ec2393d1ae23 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.455883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.455883Z digest=sha256:9df5d31b4cc0993e5c8a26ae5b3281cbcb591df3df12dff79720b48f4b3c1a1b

Observation c6bf3fea-02f6-44eb-904f-15421c7d9897 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.578707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.578707Z digest=sha256:e789159aea3ac95afc5cdcfd0e43ec8a065bf8139af13c90b203f3f8b989bc24

Observation 848fe526-4132-4bf4-9e22-842a315f02bf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.709673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.709673Z digest=sha256:c94866c978d17dcbb46d432dcec96dc89efebba9e506c5247e8dc0ae37a13a11

Observation 0bc91a6b-bbdc-4591-8105-4bfd5f909bf7 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:34.861279Z digest=sha256:9eaa7a353b135b850c5d1f4dc759e3c760a2b92aea6dea56cf811c7b5253f37b

Observation 4bf251e1-f89a-4031-8745-a2ac7ded0574 · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.950914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.950914Z digest=sha256:3bd3f274afbe68d48ed4fd91fdb158f8fa000b46a5d2e517c530c873b46d3270

Observation d72d80f3-ba3e-4c4d-9c51-5d707ddf2eb9 · outbound

This paper cites What matters when building vision-language models?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation What matters when building vision-language models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.047298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.047298Z digest=sha256:f16f8d68d2a1597fa987cf0119093be8499e29199f5907ebbb5ab959182ed82f

Observation 71df43e3-76ce-4791-a801-b71882b18abe · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.807225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.145183Z digest=sha256:f2cd87428de7f0148fbf150ac3babf5d831db6c9175b2375924a9239c7dfdda3

Observation bc9aae33-9a27-401b-b7c8-be632420679e · outbound

This paper cites T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.249132Z digest=sha256:c02a1bd1064c0afc3408a8f7ae53df7f7383724486052056566715b45b5536ef

Observation 94664539-1edf-4155-acf0-6690ccf5666b · outbound

This paper cites Remov- ing distributional discrepancies in captions improves image- text alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Remov- ing distributional discrepancies in captions improves image- text alignment

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.539329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.362579Z digest=sha256:9a0c3322423df37ca84d6aa3127748d83c877ecba53425d39f83c8ff34f3d955

Observation 337ed284-26b4-4f43-b34a-095d129c1003 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.271746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.442644Z digest=sha256:b1b67ca1e1a5e3339dd26e33808c44c6a275d2bd03743c955cddb1dee1338759

Observation 0b495463-d938-487a-836d-29917f2f9c97 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.526604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.526604Z digest=sha256:af831390f1b7bbe641f6c14aaee8f1c1dc362078d3b89f5424b89f83049afaec

Observation 485cd105-58bf-4858-b115-6c11111da2f0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Improved baselines with visual instruction tuning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.023960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.620233Z digest=sha256:71a3398f90a653abfaac02768cde1a6411e36765c63e10f20c08517b7c17d712

Observation 0750e59b-d266-4b5c-be87-abf663dab75c · outbound

This paper cites Visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Visual instruction tuning, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.790950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.676524Z digest=sha256:126ff6f878d5bdf16a72bded9bfc7b97681038768db3bfd229b58ab0af0a3093

Observation e18fadd7-3903-415d-8b9c-58d5e539e4c0 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.529889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.725893Z digest=sha256:b55b8fd5e60e54faefd4f873dfaf7c72abad429e424a73c45d36f18bfe8b17c5

Observation d4a9139f-d60a-4b32-a7b7-4b4aaa609763 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.785004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.785004Z digest=sha256:3627140e6f677d46233538e75f574360c901babdc8fb50cb8dfa90927da6a826

Observation af03050d-fe22-4c9c-86de-850fb1ef8d9a · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.880961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.880961Z digest=sha256:d05d8bd3a5d97d141c7c27aafec5371de1c5b6e9f37a50b6932b2b8a6853385b

Observation 90215b59-d14e-4456-8c68-0e470e90ac48 · outbound

This paper cites Simpo: Sim- ple preference optimization with a reference-free reward.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Simpo: Sim- ple preference optimization with a reference-free reward

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.318338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:35.929519Z digest=sha256:23acf395a8a4cdf146f8c5720a88c810dab562454d9f8f22b1e0bdcb295bfc1f

Observation 1d91967a-7d25-4f32-86be-571db2738094 · outbound

This paper cites Training language models to follow instructions with human feedback.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training language models to follow instructions with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.985916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.985916Z digest=sha256:53e7e9ddc370357fb0cca4afe90c2d943b755de0cd22e3d3c2f19f7b3d5d1c58

Observation 7dbeb8a1-6a9a-4e17-bca3-1bd6e508fc90 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.083635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.083635Z digest=sha256:529b98f876a5b595ae10abe0751ff47ff45b20b1a5d94b187a642bd79213041c

Observation 19a5864c-3e88-4d4d-8c3a-5565a4ec0ab4 · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.061472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.177621Z digest=sha256:44d31ff1c756449d350a29b41ad4bd6a7844090d8b7336fa3e518d5692976588

Observation 486054fb-1c29-405f-a1e9-74cbe132e51b · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.257427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.257427Z digest=sha256:4160cb9488050d1107e4ef017fc484aafe644b1079ed352cd887ccecfad7089d

Observation 498d6646-c498-490b-95c6-c800bcb74c8e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.362963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.362963Z digest=sha256:23ae03552f25f5dd798569fa2b22f2e6415c5cab54aa5c956c1fcee0712642cb

Observation 95dc7cc0-f7d0-40c6-b4f4-aca3578260c1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.824402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.445317Z digest=sha256:a6469e4907bd836688332061b4c58f45e9f4a07591e145072f013c5ac89e4ecc

Observation 607b8cf6-d085-4e66-8f7e-74768b3fca54 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.561887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.562776Z digest=sha256:2db292ff26c0d9719c9971789ac9cd8f78f74265c11bd2f24d85e0ca37c5c075

Observation dc7536f8-7fcf-463d-90d7-e39f42ce4053 · outbound

This paper cites Zero-shot text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Zero-shot text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.280839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.684650Z digest=sha256:b74b09e4259732e5462fb09556fa038c671246520e4af6153c964f6187f0a042

Observation c52fbfde-6f70-4c49-9cd8-36c5ef12c567 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.038981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.771953Z digest=sha256:89cb037c5cbbc7076a8b15c811c23b4d191115262d634af456bba455007f9621

Observation 4d7867f7-2ff2-42e1-8576-59eb0668b5c0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.835962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:36.863239Z digest=sha256:5c708c4332c9da08047a982d83f94e73c30afd643aec7820237ad13860b7a028

Observation f0c300b8-ddfb-44bf-b455-1d9f0393f153 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.914296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.914296Z digest=sha256:f87db68a74e8e3bd81d49e57d9d82b1374974037c2eeb0d9d6c3f192db53992f

Observation 3ce9b861-b945-4995-ba66-1d41e41b7056 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Proximal Policy Optimization Algorithms

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.988698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.988698Z digest=sha256:389e20a0f1e8fa725d1fc26eb7243940ff7a09b6271ce865798983acb9030bdc

Observation a326a054-deaf-4fbf-8f2f-962890284bd8 · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.089978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.089978Z digest=sha256:ef52c7da24a4792d06237ff82dbcf062660298edcd2e637d70039ffbbce43a32

Observation 0fb5a6c6-185f-4428-8657-6b08bb29b32c · outbound

This paper cites an unresolved cited work.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:54:39.626966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.170906Z digest=sha256:36684d64b36163add94eb94009d00efd3ccbee7b986c4caba20f9b817071f052

Observation 04f9475e-a2e7-4fcf-acdb-d64cd9abb94f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.484015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.288393Z digest=sha256:8f4d8c15d1d514711c574afc542f952b675a3ea691c598ccc993cf8ec20246eb

Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · outbound

This paper cites Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.374599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.374599Z digest=sha256:afbb23eef8e1f2eb60c0e2804e5566bbbff06936e0e66867707db1930fd0bbcf

Observation ee6f6be9-143b-4030-8a6e-1d5888fb3122 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.474204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.474204Z digest=sha256:353d962fed940891c4c22420b305bcf302a76057880cda54653c58b0595b003f

Observation cf014844-e467-460e-99a7-964de4cccaef · outbound

This paper cites Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:54:38.472533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.533704Z digest=sha256:73581ef75f4f47ff525a93046725548ed3abe1f47ae792a9a5fa53e66ea4737f

Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.619519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.619519Z digest=sha256:39906f1c047ad4acaeeb6edc8a23d3305d45e68d75a6ab42db802fe21b5473cc

Observation 0649d742-b68a-404e-b33f-82115b5ade71 · outbound

This paper cites MLLM-as-a-Judge for Image Safety without Human Labeling.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MLLM-as-a-Judge for Image Safety without Human Labeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.686456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.686456Z digest=sha256:596c2e52279d62b97e12e30923a89f2270d23503f61f51387daf85478071abd2

Observation 47994f55-4f79-43be-96c8-6f0b8b3df90c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.787465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.787465Z digest=sha256:a5bbc2edfce8feb30541d34883a11ddc075f11a0e82ef86d42d5cf384d6867d8

Observation 55c643a1-c5da-4050-9945-d9ba9e50d216 · outbound

This paper cites Human preference score: Better aligning text- to-image models with human preference.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Human preference score: Better aligning text- to-image models with human preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.344930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.859469Z digest=sha256:a731cc158bb5836c39db660101ca2508904e8717a01a17d354f13a5095a96ae8

Observation 75506a12-6c51-49bb-ab43-87fb31752d09 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.261401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.927082Z digest=sha256:c2d19a9053116a851786379340bc5902b40103f870b5eca8be10cfec1273ba29

Observation 3eaeffba-521a-4480-85fb-6f093aab3080 · outbound

This paper cites A normalized levenshtein distance metric.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A normalized levenshtein distance metric

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.129080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:37.998771Z digest=sha256:8a5a348d50b8b98dd5afad39027b8b5d58e38ee5e9e89c224d75f5a59e978883

Observation 3f3acc0d-51d8-422a-8f4f-e5647e7a1e6a · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.984919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:38.090385Z digest=sha256:066d442eed7e12c934cdf74268d664583b52b7132e014f7713aaf7bdc90ac9a1

Observation a5e53093-eb1b-4ff2-9218-0079c8c509a9 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.169486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.169486Z digest=sha256:f8d58292df60e08a10be7e10b24ca3c5a6e0f8630323ae5a21a123bc4dc04bb1

Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.229729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.229729Z digest=sha256:1d925ca09d97ee06bae5897c5c78cab6d3b894e9ee9c25255e4cd739b1e3e293

Observation cb6d2b9e-4aab-478f-ba2d-91070b9da082 · outbound

This paper cites Towards language-free training for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Towards language-free training for text-to-image generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.832890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:54:38.274595Z digest=sha256:64f5a3005c33e3b6d614d2674dd39bed21cfddf65bfa57efba7f109d6a971de5

Pith citing papers

No inbound Pith citation observations are available.