Pith. sign in

Paper Citation Record · LEDGER

Explore, Establish, Exploit: Red Teaming Language Models from Scratch

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2306.09442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.09442 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:38:36.990324Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6f0ac5f-7df8-4bf4-a62d-954b350d54a1 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.570784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:b23ec819f22d27123455b866df38a969f895266ed886a551b3b3c0e5b2cd33c3

Observation 21b70197-ee2c-45d3-97f5-10495b97384d · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:24:40.070684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:a5bcf4d9a2689e5fc8e586c8a853731a1e940a3b3afaa01d33feb316e625585e

Observation 7f4390b7-62ee-445a-8216-340d244f8201 · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:00:51.519583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:edfb3b64fe08a822cf33d4bc1030472c8efe6af66a395edb56cddbbfbc1cbb29

Observation d6ab93bd-c519-45f4-a798-b834acb4f06f · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.177956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:0ec85a89feb916e0bb05f374aa8bd5c3f27e10cb9593a1ede97df6d9ff8ae1c2

Observation febf8415-c6bf-42ff-b85c-77200f92c9da · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:17:08.665940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:0e8e43cb4aae86d4baf633c208287c01edaaa7089a432aaabcf9e62bd8ce121d

Observation 3eb0970a-203c-4127-8961-9e16f4df2c85 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.567150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:d33568f59a1d0051eca64dc4638613eb661b20b2d70e949248972648359c4134

Observation e5b8c7e7-81f2-4b82-aa35-284e15c98033 · inbound

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning cites this paper.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.990324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.990324Z digest=sha256:ab8e0591b2e0bdd420934e582b316c281a6c9ea880ceecdc73466c5531d05d45

Observation 73ca601b-d2f1-4d1e-94e7-bad9260152b8 · inbound

CALM: Curiosity-Driven Auditing for Large Language Models cites this paper.

CALM: Curiosity-Driven Auditing for Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:52.515074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:03:52.515074Z digest=sha256:7a4e80ca395240e1413acdc1bcc66e94b84dac5ccd2a2b084831cf678c10e2c3

Observation 5340d82a-18b8-4cef-977d-3eee2b8506de · inbound

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints cites this paper.

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:36.468720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:36.468720Z digest=sha256:a82120b8eae2b8304e5960fac214eea871fda8afa60d28621fa36156c858a78f

Observation c2dc83aa-689e-41e5-8127-bf4046b52290 · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.208623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.208623Z digest=sha256:19b315ec1a70564d000449e7730b13500becc782f28608e9301f41e9f1470eac

Observation 9b32b33d-7756-4ec4-913e-7c6d1be463bd · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.649185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.649185Z digest=sha256:31d3da64e62fd376ddb326aea99add4fa7083223c8d1cdfc9ca763fd6b947053

Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.057855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.057855Z digest=sha256:c670d53d783024d7c05a2214704111a8317d737244681bc6fbe22d2617acc765

Observation cfba1cf5-35fd-49e4-81d5-182a7d296090 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:59.966022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:59.966022Z digest=sha256:1767875901f914a25638a1a90aec4f55f2b88576ee269f4802c8bb6c1f1b424f

Observation d1c03a4e-92e4-4a51-9c7c-61118fc31784 · inbound

Kaleidoscopic Teaming in Multi Agent Simulations cites this paper.

Kaleidoscopic Teaming in Multi Agent Simulations Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:47.454816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:47.454816Z digest=sha256:c222d5e76e55ba100ef92e2ae0b2e618990c1228b3dd4ab919dcf54b2a733a59

Observation 54a66195-b915-43e4-9959-84fd95786d5c · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:55.144189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:55.144189Z digest=sha256:076f5aa1f43c1005092c75b283baca5b5f8843f5735f8f59982f97d2d7a5e2de

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:1673627be3bb85834e86678d60981d42a7984b1a9cdf66251434799ef00451fa

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:ec61fe362c8cc3f95b4e5b38a2337b069c8cf46c41796cbc5a08b1b3e48182a1

Observation 532f346b-bf62-476d-9d95-24477cb4d118 · inbound

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs cites this paper.

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.872573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:52:44.872573Z digest=sha256:0d4278d3b5c3816769f4a96950ba0fdbde636d62418c26a639757b2eaf17d4ad

Observation eb7228e8-a91f-4b61-ae2a-d1c55a42c3b9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.273056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.273056Z digest=sha256:397361a186f6a1de431c64920bc0e3463f675a3936d5bd9b6f9bc93cb80b1514

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:5ae6e144dd334d8f368d3138f5d8b7afbc53bb647bca7e9c2f5ed4ffa321c0c8

Observation 0cf155fc-6461-4c27-a3bc-1c990bfdb00c · inbound

Tailored untruths: How personalisation challenges LLM safeguards cites this paper.

Tailored untruths: How personalisation challenges LLM safeguards Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:55:53.355255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:55:53.355255Z digest=sha256:40372e508780618156093a07209631e38fe7e95957344a938e9ecd04a021c4b1

Observation ab94034e-cb58-4d6f-acc7-7f9ffdad0f68 · inbound

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models cites this paper.

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.413805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T08:36:39.242766Z digest=sha256:35e80a7a48baebb6d1cbc3096fe9caa0975885977028d4ec3b71049d8e9c116a

Observation d8955db2-bcdd-499c-b8ec-3b9aad7f8b3e · inbound

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF cites this paper.

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.405159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T04:35:51.223025Z digest=sha256:3615767a8f279c627a4e001d1527acd3da175420ff9d62a9cd7d3dcf4f993149

Observation 27e374e1-026f-4884-9734-44dca0ac509e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:40:52.596590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:32:42.151644Z digest=sha256:3a9f1c3bb9ce3a0c7f3f61ed1eca959bb6a0172dee84e4904a4dca3882c2751d

Observation 1c8e15fe-8f07-42ae-8d0c-98f83597482b · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:35:46.533799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T22:59:07.861941Z digest=sha256:7162b57c3c8d68c6f83c2b3c6bbf6b497f3da2087c937cf5424ace5461756f30

Observation 013aef7e-a486-4a92-b0c6-1e03faf6af9e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T14:41:05.649595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:41:05.649595Z digest=sha256:ad25faf74434312d65207e4b8034d0ee228cc8b37b9b45ed62237a9362697ffc

Observation b326c0f5-09b8-470d-bb57-d1c0b650a771 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.348080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:285058698a2c6eebf706f54726e44ab679de24ece5238eb41b6af9088a2a0da7

Observation d7bd52b0-f8a0-451f-9ea4-8ab3a91f5838 · inbound

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems cites this paper.

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:28:48.020543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T17:26:02.346222Z digest=sha256:9638d6c47ee2bcae2cb82068fe7bd4ba57e5f23c525d58f618af0ea6df1b7a7b

Observation 622045b0-0c14-4d76-a948-691036c06a9d · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.267655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:70123c0ca235c51e92c02c5aa6f79d0acf98bbc4bf1410c75036d706671a95a0

Observation c911afea-adfd-4ed9-aa69-f2d0abe6abf6 · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:49:44.864881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:59f835efe2a40fd8c6ea8f081e2dfe74368ee292d19b9ce38abcc994880bdb96

Observation 1c5b64cb-10fe-4708-8cb7-f0e89a5b8e2a · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.519430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:0b851a155dbd3f4bf5f9d60e6d66e702279feaf1331a8b88c72c8ca8a7a656ea

Observation c257477e-7e72-4329-a983-e8ba226c1af0 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:08.962913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:08.962913Z digest=sha256:90221e79fdcb9fad05af3f7c42a80d98cffdc68bbf68a6eb975941e7a4b63d4b

Observation 40e14c18-6f9e-4d12-aa77-8bfc6d0d6d39 · inbound

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs cites this paper.

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T15:07:17.112776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:07:17.112776Z digest=sha256:a5aa28ea3e7bca73ef69f531ba2c6172153307473e476cc49dd2dc33e35ae966