Pith. sign in

Paper Citation Record · LEDGER

BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2304.12298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.12298 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:50:00.382122Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T17:33:45.482173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10b9aa62-1d6c-4765-96a2-d45dcc45ac36 · inbound

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions cites this paper.

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:23:52.841681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:21:49.775278Z digest=sha256:36a7c9e28187dc48ba114b138500f65e762c4de3ca2588a50aa55f1548ddc60f

Observation 6f146300-3526-4da1-910a-6964f6435ddf · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:25.838070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:f2751f36cc07af6169b1cb7be2057f773b329065e0f84f10991aab29e0f6b89b

Observation 3b91c2c9-b116-465c-80a1-abf66d2bc7a7 · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.382122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.382122Z digest=sha256:13a30669b61373fcff804e1e34e01e5dc83e973b3376871dfc4cbc0c9bfcade1

Observation 24375860-c72e-4ad5-b9c4-f90a96155f81 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.635415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.635415Z digest=sha256:e65e11c627e11e2e7deb0091e8e74414638328487e1d2112a799a1d644f25091

Observation efce6e4a-a467-4b81-ba43-4873f940c9e4 · inbound

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models cites this paper.

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:16.463357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:16.463357Z digest=sha256:6be8807b2e5d4d93a158e4319f29a907f4450fa5cb6d81fd785646a9142f4a89

Observation 4f3f77f8-538b-4ef9-bee2-232d43f03249 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:29.029474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:30e46a3c047e6f5d550aaad39769615bf65b99c499e973b5ace2e2e981fb4d1d

Observation 07f40ae0-af66-4c3a-a03a-cff0146af70c · inbound

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers cites this paper.

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:07.637226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:37:43.177370Z digest=sha256:48f76b3e473478c163ae9cc2ecc46674eb9020db3701fbfe938fb274ac7a18e1

Observation a2dc3a4d-c4f1-407e-aa02-bcff82090b77 · inbound

GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning cites this paper.

GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:33:45.483983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:25:22.733770Z digest=sha256:eb9cf65f39f4f87c7aba259500e4436c4547c41a5b715352402079d5c7caee25

Observation 290c6378-426c-429c-903a-9f76fb213550 · inbound

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks cites this paper.

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T03:06:37.057046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:06:37.057046Z digest=sha256:52a23e3c3f0c163bc2439885c72fe24efc9b5cbb1b25dfbdf8d2ac856685b4d9