Pith. sign in

Paper Citation Record · LEDGER

BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2304.12298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.12298 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:07:40.541758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T17:33:45.482173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10b9aa62-1d6c-4765-96a2-d45dcc45ac36 · inbound

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions cites this paper.

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:23:52.841681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T04:21:49.775278Z digest=sha256:5269ce41bc33728728bd00a36dbdefd97a2e8a36cc2e1fc1d9eda6e19bd7c04e

Observation 6f146300-3526-4da1-910a-6964f6435ddf · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:25.838070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:09f0b0954605d272969e0a953645163e50227d1b747f828dcf4501eca3c64d44

Observation e91b2d41-1ae5-4f59-8db7-1a952704305b · inbound

BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks cites this paper.

BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T05:07:40.541758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:07:40.541758Z digest=sha256:8d062c52f19ab2dca8815a3b012ecd526826b1841a16ef6d21bfce947a781397

Observation 88360933-99a7-4541-be0a-1230fb97414d · inbound

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage cites this paper.

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:30:10.297459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:30:10.297459Z digest=sha256:c04a0b0ccfb56a86e9b8330a30f52e4956533ff516ae652098d6fac4fa5eda0e

Observation d09b3dc6-bdc9-460d-b39d-4eb54325c82e · inbound

SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation cites this paper.

SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generation BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:33.716635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:33.716635Z digest=sha256:fcb526ee2c8fb74511b4243f7fe26e3e3c3aaafbdd531ebc790cbf8e447d61bc

Observation 1bd56c6d-bff4-4179-be4a-39a74ddbfb0f · inbound

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models cites this paper.

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:06.021273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:06.021273Z digest=sha256:a8ccaf1a8aaa6015b3a1c927387c4ea6c7a78d49874bfdaa67fb20d52cf0a50d

Observation 3b91c2c9-b116-465c-80a1-abf66d2bc7a7 · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.382122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.382122Z digest=sha256:13a30669b61373fcff804e1e34e01e5dc83e973b3376871dfc4cbc0c9bfcade1

Observation 24375860-c72e-4ad5-b9c4-f90a96155f81 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.635415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.635415Z digest=sha256:562af5458c48084f68e92e9dcd42e1a0ebf3d9aa292b617dcf7dbac535c12c4c

Observation efce6e4a-a467-4b81-ba43-4873f940c9e4 · inbound

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models cites this paper.

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:16.463357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:16.463357Z digest=sha256:ad00648270ca3af90aaa25a4c9e2ad933cd849e082eca2607b319ed465374183

Observation 4f3f77f8-538b-4ef9-bee2-232d43f03249 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:29.029474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:ce4a2a4741a6bf6647e3d2a7b68dce5735b3eb1c809fbf39d79f7dfdbf88ad31

Observation 07f40ae0-af66-4c3a-a03a-cff0146af70c · inbound

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers cites this paper.

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:07.637226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T21:37:43.177370Z digest=sha256:c130de5aa4f50ac8808c29ef40856f80ee2a6caba6a84ed582e9ee58d2bff548

Observation a2dc3a4d-c4f1-407e-aa02-bcff82090b77 · inbound

GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning cites this paper.

GradSentry: Gradient Spectral Entropy for Backdoor Sample Filtering in Large Language Model Fine-Tuning BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:33:45.483983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T17:25:22.733770Z digest=sha256:85d9fd7497a6aff4d18d50648cd56ac10eecaaea937d70f56d5efae82e473baf

Observation 290c6378-426c-429c-903a-9f76fb213550 · inbound

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks cites this paper.

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T03:06:37.057046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:06:37.057046Z digest=sha256:52a23e3c3f0c163bc2439885c72fe24efc9b5cbb1b25dfbdf8d2ac856685b4d9