Pith. sign in

Paper Citation Record · LEDGER

Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2404.05530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05530 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.434136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:35:34.297370Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 537888dd-05cd-4539-a3fe-422b4264b677 · inbound

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs cites this paper.

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:18.667825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:18.667825Z digest=sha256:8c0b671168498f2a106640f49b1a88af55b284774db9dba95c683bdacddd04bd

Observation 05655a15-77a3-4cf5-8d64-d4b422561197 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.434136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.434136Z digest=sha256:04b6f9fedc91bcb0b6a8f81ab57643067f7cc122be29e8809afd45703c8af180

Observation c232ca51-e49b-4862-ba9d-ffcc20072146 · inbound

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF cites this paper.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:03.450932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:03.450932Z digest=sha256:1a398e96ea5eb39a002eea5f3ff3ddcd4204d628e35bd6c5ce131e309803faec

Observation 7ea8583f-01ce-43e0-8962-3952464ac255 · inbound

A Systematic Review of Poisoning Attacks Against Large Language Models cites this paper.

A Systematic Review of Poisoning Attacks Against Large Language Models Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:31.659055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:31.659055Z digest=sha256:1a9f1c3008d6b12fa30ba6ec8c9a718677b9e9318ee4b84b1ce1bc95dec002da

Observation a7a1e449-3bf3-4fce-81c0-d98dbe07515a · inbound

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users cites this paper.

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:02:07.889604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T05:58:17.452837Z digest=sha256:8a1890069a22f1ea287e7562c4aa997dcd80022652cb826d1009780ba93bfef1

Observation 2d117101-72e5-4df1-a85c-1bed623829a0 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:29.010728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:4a1f521ed17b86f3bac4ad69ed1348aa043ff39ba2376dd77c671e40138fbec6

Observation 6d7cedd1-8dc4-434b-a735-c9b982f3d33c · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:50:27.078067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:d47f098e5fb6a7c8dc4e8b2c30660327e1cef641ebd720e4a1a6e98418e60551

Observation 07fd624c-d5f8-4775-981a-df5f100701ad · inbound

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models cites this paper.

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.876995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:01:10.756844Z digest=sha256:bfb38f83fb76af6ee32931119ddef9b5226c9c5231d9bfb3f2ac5f49bef71ad6

Observation 755586fc-8e5b-48df-b2e1-fb846640574c · inbound

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks cites this paper.

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:58:10.414114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T08:53:52.698758Z digest=sha256:fe516a6314dcaec89e793cd93ee656408f73f592409e37115a1e500b5aca64e0

Observation a5cfad75-ff3f-4040-8ca4-6fdb3083bd27 · inbound

Reframing AGI Confrontation with Off Earth Autonomy cites this paper.

Reframing AGI Confrontation with Off Earth Autonomy Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:35:34.298617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T07:18:55.451342Z digest=sha256:b1ac073ce37a7bc158116bf898d2e2170f5887c5f6ed9b5a328554ba83fb7fa9