Pith. sign in

Paper Citation Record · LEDGER

Explore, Establish, Exploit: Red Teaming Language Models from Scratch

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2306.09442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.09442 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:38:36.990324Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6f0ac5f-7df8-4bf4-a62d-954b350d54a1 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.570784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:4db57c8176221ffffecf1f6ce4d46a89cd2e5ebdad5387eeb07bfd5cd6872538

Observation 21b70197-ee2c-45d3-97f5-10495b97384d · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:24:40.070684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:b4b3ded784c9b9ccf52a2a0d8ad770ea9863f80bc6ba0f11f85ad65377c8e12d

Observation 7f4390b7-62ee-445a-8216-340d244f8201 · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:00:51.519583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:fba1a5b44d73f76c856e5d23e6b46a5af57c208bb9df7f7f256bbf2ffde16033

Observation d6ab93bd-c519-45f4-a798-b834acb4f06f · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.177956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:013ef11e67efa40efc5da5fe1c44a69e9767bc1552257a4cca5f2b507a5343c8

Observation febf8415-c6bf-42ff-b85c-77200f92c9da · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:17:08.665940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:f2a334865d08c33f2bfe629089863d690443e16718e0ebe9faec060a25048a0e

Observation 3eb0970a-203c-4127-8961-9e16f4df2c85 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.567150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:aa2ea98484d059d1a76d6d8f2c29d2b9c77409e82ff7b781cdaf6c7819ec6eb8

Observation e5b8c7e7-81f2-4b82-aa35-284e15c98033 · inbound

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning cites this paper.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.990324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.990324Z digest=sha256:088a36e49146775b347ba83f9792266843c8df8173079692c66363a976dc1568

Observation 73ca601b-d2f1-4d1e-94e7-bad9260152b8 · inbound

CALM: Curiosity-Driven Auditing for Large Language Models cites this paper.

CALM: Curiosity-Driven Auditing for Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:52.515074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:03:52.515074Z digest=sha256:7a4e80ca395240e1413acdc1bcc66e94b84dac5ccd2a2b084831cf678c10e2c3

Observation 5340d82a-18b8-4cef-977d-3eee2b8506de · inbound

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints cites this paper.

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:36.468720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:36.468720Z digest=sha256:a82120b8eae2b8304e5960fac214eea871fda8afa60d28621fa36156c858a78f

Observation c2dc83aa-689e-41e5-8127-bf4046b52290 · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.208623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.208623Z digest=sha256:19b315ec1a70564d000449e7730b13500becc782f28608e9301f41e9f1470eac

Observation 9b32b33d-7756-4ec4-913e-7c6d1be463bd · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.649185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.649185Z digest=sha256:8a765e2933ab4063d68a62912010d917a7b158d7deb0ea1339335bf3a1f73d32

Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.057855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.057855Z digest=sha256:c670d53d783024d7c05a2214704111a8317d737244681bc6fbe22d2617acc765

Observation cfba1cf5-35fd-49e4-81d5-182a7d296090 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:59.966022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:59.966022Z digest=sha256:1767875901f914a25638a1a90aec4f55f2b88576ee269f4802c8bb6c1f1b424f

Observation d1c03a4e-92e4-4a51-9c7c-61118fc31784 · inbound

Kaleidoscopic Teaming in Multi Agent Simulations cites this paper.

Kaleidoscopic Teaming in Multi Agent Simulations Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:47.454816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:47.454816Z digest=sha256:c222d5e76e55ba100ef92e2ae0b2e618990c1228b3dd4ab919dcf54b2a733a59

Observation 54a66195-b915-43e4-9959-84fd95786d5c · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:55.144189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:55.144189Z digest=sha256:076f5aa1f43c1005092c75b283baca5b5f8843f5735f8f59982f97d2d7a5e2de

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:1673627be3bb85834e86678d60981d42a7984b1a9cdf66251434799ef00451fa

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:ec61fe362c8cc3f95b4e5b38a2337b069c8cf46c41796cbc5a08b1b3e48182a1

Observation 532f346b-bf62-476d-9d95-24477cb4d118 · inbound

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs cites this paper.

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.872573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:52:44.872573Z digest=sha256:0d4278d3b5c3816769f4a96950ba0fdbde636d62418c26a639757b2eaf17d4ad

Observation eb7228e8-a91f-4b61-ae2a-d1c55a42c3b9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.273056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.273056Z digest=sha256:397361a186f6a1de431c64920bc0e3463f675a3936d5bd9b6f9bc93cb80b1514

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:5ae6e144dd334d8f368d3138f5d8b7afbc53bb647bca7e9c2f5ed4ffa321c0c8

Observation 0cf155fc-6461-4c27-a3bc-1c990bfdb00c · inbound

Tailored untruths: How personalisation challenges LLM safeguards cites this paper.

Tailored untruths: How personalisation challenges LLM safeguards Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:55:53.355255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:55:53.355255Z digest=sha256:40372e508780618156093a07209631e38fe7e95957344a938e9ecd04a021c4b1

Observation ab94034e-cb58-4d6f-acc7-7f9ffdad0f68 · inbound

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models cites this paper.

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.413805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T08:36:39.242766Z digest=sha256:5d4812c7ccfcf675a068b62f64a11a1361fcfad841e27b3cf65cd72c4edbfc9f

Observation d8955db2-bcdd-499c-b8ec-3b9aad7f8b3e · inbound

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF cites this paper.

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.405159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T04:35:51.223025Z digest=sha256:6322fa88ac193d300f1d4ddcae93b33fa35d002cd11ed56ef53cdbf14a619900

Observation 27e374e1-026f-4884-9734-44dca0ac509e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:40:52.596590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:32:42.151644Z digest=sha256:5c3590648c8f3845569d3766cf30b82e1c21ac8abc4f76951aac804a86c56e9b

Observation 1c8e15fe-8f07-42ae-8d0c-98f83597482b · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:35:46.533799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:59:07.861941Z digest=sha256:32674b9e0fe4663d392fe70d455c1a76677fe01cf7519bf272deeb6d2ddeb282

Observation 013aef7e-a486-4a92-b0c6-1e03faf6af9e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T14:41:05.649595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:41:05.649595Z digest=sha256:ad25faf74434312d65207e4b8034d0ee228cc8b37b9b45ed62237a9362697ffc

Observation b326c0f5-09b8-470d-bb57-d1c0b650a771 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.348080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:7422aeaf0cbe831c579db545dbf36d3dfc63540c4b475d69a95f6f0834c4df7b

Observation d7bd52b0-f8a0-451f-9ea4-8ab3a91f5838 · inbound

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems cites this paper.

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:28:48.020543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T17:26:02.346222Z digest=sha256:fd797ed95b6e4e85b2f8aaf7b922196dca979b16949256e0009d085990f31515

Observation 622045b0-0c14-4d76-a948-691036c06a9d · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.267655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:439181ae9cbcb154a0c97a96b3eb52b72bcd88017f54e65e0a1136a8ef5a974f

Observation c911afea-adfd-4ed9-aa69-f2d0abe6abf6 · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:49:44.864881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:595f62996eec8dc381ffcdd6a842283864925fbd9f89fe6e9511692957196adb

Observation 1c5b64cb-10fe-4708-8cb7-f0e89a5b8e2a · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.519430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:c25ce40779f4714699244baad15780f57c431deb16e077eefccf69ab7afc0cc0

Observation c257477e-7e72-4329-a983-e8ba226c1af0 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:08.962913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:08.962913Z digest=sha256:90221e79fdcb9fad05af3f7c42a80d98cffdc68bbf68a6eb975941e7a4b63d4b

Observation 40e14c18-6f9e-4d12-aa77-8bfc6d0d6d39 · inbound

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs cites this paper.

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T15:07:17.112776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:07:17.112776Z digest=sha256:a5aa28ea3e7bca73ef69f531ba2c6172153307473e476cc49dd2dc33e35ae966