Pith. sign in

Paper Citation Record · LEDGER

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

As of 21 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2506.03234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03234 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:06.054981Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:26:34.918099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.522097Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c232ca51-e49b-4862-ba9d-ffcc20072146 · outbound

This paper cites Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.450932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.450932Z digest=sha256:fbbce498ad5c514dc9448bcac69eaa5bc477784df5a0f2c8cd5124445644dc56

Observation 4c5bc4b2-3cf6-451b-8f0e-fa0f87f8f7ff · outbound

This paper cites Poisoning Attacks against Support Vector Machines.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poisoning Attacks against Support Vector Machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.502246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.502246Z digest=sha256:d5f6bdfb71987077d4395252b4e7a6f09bc79177b55a2566388d59835b4ad3ba

Observation 9a7d5036-1e5b-40ac-86eb-e677ced3f55f · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Training Diffusion Models with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.581756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.581756Z digest=sha256:a0b5dc78a7f4f70df2e590a43d76c3e962cd4756f1d0e30b2ac9813bd56cfae1

Observation cc6c3d91-5ef5-4e50-bae9-873226a676fc · outbound

This paper cites A survey on generative diffusion models.IEEE Transactions on Knowledge and Data Engineering, 2024.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF A survey on generative diffusion models.IEEE Transactions on Knowledge and Data Engineering, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.657237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.657237Z digest=sha256:9f564ed02dc80771e0b33333608407cc6ad864d7bf44396d77bde4853fda7a8b

Observation 08bfd459-da94-4898-be9a-85e1de397f36 · outbound

This paper cites Trojdiff: Trojan attacks on diffusion models with diverse targets.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Trojdiff: Trojan attacks on diffusion models with diverse targets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.750828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.750828Z digest=sha256:13cf05cccc12b99d5e4f33e8bf8c4c9cfb144b6dae95a610e4d05b3a4d5878c1

Observation 161f05ef-a822-4d0d-9935-86d69a731fc9 · outbound

This paper cites Amplifying membership exposure via data poisoning.Advances in Neural Information Processing Systems, 35:29830– 29844, 2022.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Amplifying membership exposure via data poisoning.Advances in Neural Information Processing Systems, 35:29830– 29844, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:08.561155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:03.836456Z digest=sha256:c825586f46b276ef709fd9dff92bb0f3b0610bb09e68ec8180f41060fa98574e

Observation e758bf7b-1171-43c7-83a9-775a9b3d7d14 · outbound

This paper cites Villandiffusion: A unified backdoor attack framework for diffusion models.Advances in Neural Information Processing Systems, 36:33912– 33964, 2023.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Villandiffusion: A unified backdoor attack framework for diffusion models.Advances in Neural Information Processing Systems, 36:33912– 33964, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.929697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.929697Z digest=sha256:16a0da84b904fa68f6dfaa60be49bdb3ad22257d521fbf90cefb8aec18f29c15

Observation d73902d0-311e-4dd5-b6c3-8fe604230f35 · outbound

This paper cites Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9):10850–10869, 2023.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9):10850–10869, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.996111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.996111Z digest=sha256:cfbcf575490e47d0b53c7e0013fcb74d3d1d472d7079b64e5c72529a92c087f5

Observation 37ae54e1-c492-41d2-9d81-98f72add11b5 · outbound

This paper cites A survey on data poisoning attacks and defenses.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF A survey on data poisoning attacks and defenses

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:08.342924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.089124Z digest=sha256:41eaecc2869acd2ef1f2c363e2589f6f4fbe01eac25a6036dd5e5f9b3664a25b

Observation 154ec68f-8638-40af-9e34-1f2dc9129a4a · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Systems, 36:79858–79885, 2023.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Systems, 36:79858–79885, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.152983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.152983Z digest=sha256:3832d9c96ab8c5f4b8d261138b700c69f6b2548830b7f0e3391d4c22323904e4

Observation 8d7a39fa-b174-45d3-9f5e-7a4c4ddc7b8f · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:36652–36663, 2023.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:36652–36663, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.220875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.220875Z digest=sha256:9cace7196e0c20624a529dcd7b6abc8d3fdfbada9f1a9035404728b4563127b4

Observation b37bad84-ff1b-432d-a2b4-af5336bffd74 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Aligning Text-to-Image Models using Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.275836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.275836Z digest=sha256:020d794b8677c55845601251286daf49ba5d59dd6b45db4852fb27fcc3ed8edf

Observation ae670811-d078-4079-a6b9-4f6066a54382 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.336470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.336470Z digest=sha256:fe25797a03def497cfedfecbafcef282a58f65519ccab4a4f83bd4fe14f5e2e7

Observation 48f08532-d743-47ae-a892-36c3f5f6222f · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:08.117549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.397499Z digest=sha256:b45f8c00ba5e1139239c58a7957fc7c412da5559888f139ee42a0e28ac95ae54

Observation e2d3d101-2de6-41d1-ba22-f1a52ee1920b · outbound

This paper cites Backdooring Bias ($B^2$) into Stable Diffusion Models.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Backdooring Bias ($B^2$) into Stable Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.459824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.459824Z digest=sha256:7b86827261d486f653ae8bced89a53bbe43ef2290a741a4fba085aed4510728d

Observation 7f949655-be9a-4032-8c9f-09a7e5a7ad61 · outbound

This paper cites From trojan horses to castle walls: Unveiling bilateral data poisoning effects in diffusion models.Advances in Neural Information Processing Systems, 37:82265–82295, 2024.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF From trojan horses to castle walls: Unveiling bilateral data poisoning effects in diffusion models.Advances in Neural Information Processing Systems, 37:82265–82295, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.974486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.508075Z digest=sha256:3b92ee2f2d7ac4720d7e9e4202a8d4d8a3566079b854d11fe35a780415a5a492

Observation 59a7cbec-dd05-4dad-bc13-825bbfec08d6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Learning transferable visual models from natural language supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.576132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.576132Z digest=sha256:cd7bd8da3e40929a30d0cd43f54eafd89ad3edadfa63594a853c379e03c184c1

Observation b6007439-ac16-4601-bf79-aefa8c707234 · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.647947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.647947Z digest=sha256:92a86f4aca208b8c85d0ab813356c72da4fcde7df81944f7e537e1fd1f43cc73

Observation 50d6e421-ac7a-49fb-9868-0ca8442a5a76 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.739251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.739251Z digest=sha256:80e6277b6622e70e6a3d8b976b5cf55f0a6cb98b862ca014566cc4c07793ff00

Observation 5e05d140-3491-48fb-814f-a6e74ae761a1 · outbound

This paper cites Poison frogs! targeted clean-label poisoning attacks on neural networks.Advances in neural information processing systems, 31, 2018.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poison frogs! targeted clean-label poisoning attacks on neural networks.Advances in neural information processing systems, 31, 2018

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.823975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.821469Z digest=sha256:8de936d76bf0c192f1ab2936cf2330e321a9b73a818a0327dff227c87a3d64e4

Observation 54f7eae3-a740-4c7a-ad9f-5aac9f7f5c47 · outbound

This paper cites Nightshade: Prompt-specific poisoning attacks on text-to-image generative models.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Nightshade: Prompt-specific poisoning attacks on text-to-image generative models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.666359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.886395Z digest=sha256:21235fa178d6cc66cf6abebd52b73afb3a98cb32919b7ff3be9d426de5faced5

Observation 9380d697-33e0-47d7-8aa8-f39f48fe80e4 · outbound

This paper cites Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35:9460– 9471, 2022.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35:9460– 9471, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.937273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.937273Z digest=sha256:20a191a06b4c1c41b2efd193a4d07319863be5c9f3ec0eb8c6cb35999171b087

Observation ff8d6a03-ecc2-4516-9809-9da4766a8a65 · outbound

This paper cites Attacks and defenses for generative diffusion models: A comprehensive survey.ACM Computing Surveys, 57(8):1–44, 2025.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Attacks and defenses for generative diffusion models: A comprehensive survey.ACM Computing Surveys, 57(8):1–44, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.493436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:04.989332Z digest=sha256:6a7c33275e3446f69df9bb2b80dc5229df2516267dc4942cdbf0d3e8cba0eef1

Observation c35545e0-c05a-4570-9d41-3a9605c58095 · outbound

This paper cites Rlhfpoi- son: Reward poisoning attack for reinforcement learning with human feedback in large language models.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Rlhfpoi- son: Reward poisoning attack for reinforcement learning with human feedback in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.285820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:05.061206Z digest=sha256:6b9bccb477c9e4bcdc4e54edce03476cd3ec211942a7fefe9ea866e97b80a99c

Observation d7e337ca-da47-4492-90f3-2a6be172ac8c · outbound

This paper cites Preference Poisoning Attacks on Reward Model Learning.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Preference Poisoning Attacks on Reward Model Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.136617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.136617Z digest=sha256:b06c7660e55d11a4247538925af8e7c79b0b24ce80411c2870f97a6d7a7f0a03

Observation 6c07524e-4b2d-4435-b510-47ce1c7e16e8 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.209176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.209176Z digest=sha256:5f8276c105a1f4e307ea164482673c853b76147307a3120c80c69cbba58e5481

Observation 5955355f-aeb3-4f18-a634-ce49693845f4 · outbound

This paper cites Human preference score: Better aligning text-to-image models with human preference.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Human preference score: Better aligning text-to-image models with human preference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.279581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.279581Z digest=sha256:ae8c122ee3047b066e345f2bf68b6ce39f32df55caaa8acbb322b76fb6c597a3

Observation 180fd271-a4cc-4700-9b6f-c017fbcedf9d · outbound

This paper cites Adversarial label flips attack on support vector machines.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Adversarial label flips attack on support vector machines

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:07.088382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:05.358874Z digest=sha256:c2b9c523880429c5395d207acbd4ee9037290d750dfaa78051efa48edca5d0d4

Observation fd29907a-9f14-49c8-b5f2-e187b32ce0d1 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.433869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.433869Z digest=sha256:3f5a4b550138c65c0728b70b52c15676c3fccb2bf3746eba5a266b4a8c6a0870

Observation 859bd800-3a36-4ae6-9149-195f24f1cdce · outbound

This paper cites Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.514154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.514154Z digest=sha256:87b6b11b47a595eda2c7649b9d87c93945438d8ef1b72248e8b46b6c974ec5b1

Observation 3e6ebbea-9cba-46fe-a5b7-8c55c5a571cc · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Using human feedback to fine-tune diffusion models without any reward model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:06.916979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:05.550741Z digest=sha256:0473af0912814ab5d3ec8cebf75572f93f2d09237f14c10d692ddddaa51d5745

Observation 65970768-5df4-4d28-8b38-2231aaf7b64e · outbound

This paper cites Diffusion models: A comprehensive survey of methods and applications.ACM Computing Surveys, 56(4):1–39, 2023.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion models: A comprehensive survey of methods and applications.ACM Computing Surveys, 56(4):1–39, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.607651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.607651Z digest=sha256:4e517da0ba351f4ddcd2f63364b628bb6003f8cf455bb4f7b5dc45c146248591

Observation 3442542c-30c4-4352-af91-c631d8821e31 · outbound

This paper cites Poisonprompt: Backdoor attack on prompt-based large language models.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Poisonprompt: Backdoor attack on prompt-based large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:06.732475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:05.679279Z digest=sha256:ee69fd5e0895c0e0a78263e8c5d08c882bed2d1e7f252c58e70b42fc5bd456a0

Observation 0341263c-6b8a-4661-8ce4-95d8a6a8b82d · outbound

This paper cites Text- to-image diffusion models can be easily backdoored through multimodal data poisoning.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Text- to-image diffusion models can be easily backdoored through multimodal data poisoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.729802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.729802Z digest=sha256:bb8b5d2fd095d09be04d14a148329175e7ea36678a31c710fabf19b983e08abd

Observation e7e926ed-4c2b-437d-bc93-608708a72727 · outbound

This paper cites Text-to-image Diffusion Models in Generative AI: A Survey.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Text-to-image Diffusion Models in Generative AI: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.800563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.800563Z digest=sha256:838e153239d6748465cfac7e8c2505c0db54cf26c72a757e125a4784003f65ff

Observation 2b07a20c-d990-48b9-a212-dd7b6110fd25 · outbound

This paper cites Aligning few-step diffusion models with dense reward difference learning.arXiv preprint arXiv:2411.11727, 2024.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Aligning few-step diffusion models with dense reward difference learning.arXiv preprint arXiv:2411.11727, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:05.875543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.875543Z digest=sha256:ba1c8cb8c89aec1f72e150d86d1be6bb0756c670b7e9bbcc40de8f59364d8ebe

Observation e09bfbe4-338b-43b7-8659-9d26de7da513 · outbound

This paper cites Shielding collaborative learning: Mitigating poisoning attacks through client-side detection.IEEE Transactions on Dependable and Secure Computing, 18(5):2029–2041, 2020.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Shielding collaborative learning: Mitigating poisoning attacks through client-side detection.IEEE Transactions on Dependable and Secure Computing, 18(5):2029–2041, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:06.433447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:05.936387Z digest=sha256:df304ea93372f121a7394a9f3d47f9e798af84020e545857a7b16a25f51f2a67

Observation e9268e5d-894a-41da-860e-3e09c4e33196 · outbound

This paper cites Diffusion Models for Reinforcement Learning: A Survey.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Diffusion Models for Reinforcement Learning: A Survey

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:18:05.968907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:05.968907Z digest=sha256:4ffe2fa4955eb97c5f1be25d73567de0c7c8ede009c44cdc390f87554316a245

Observation c89c4aad-343c-4484-8241-cfbb1765f46e · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:18:06.292128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:18:06.054981Z digest=sha256:f334d9b65667c0f241cc23b1917c62b5c6c6403b033fab2f98f76244187afd89

Pith citing papers

Observation 4b6439e1-8155-4c28-8686-855b5bee67a9 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

Reference 221

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.455819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:f3c73e2e37418bdf9696a6e3c1a52421b2d99ffa2e6d96ee6845750f3f6fe661

Observation b58c985f-04cf-476f-ab99-e55e1c7a5c69 · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:29.437150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:a9641453b056c3cbc8f519fa996746bb81ff06d3d7721b43f5f4d2d6c49feb91

Observation 2607aec9-0022-4ee2-95b0-b7697c054fc0 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.523466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:7dad7949b4edbf0c126802318a1fb21e41a6593c8c75e09d55dcec4201c6a495