Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:32:06.668110Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.28959.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:32:06.668110Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2551cf87-a5b1-4484-a124-3b114d97f47c · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates B Additional Results B.1 More Results Table 5: Additional cross-dataset robustness results on Pythia-1.4B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3296c2b7-700a-42d3-b72a-66524f9d805d · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Transformer feed-forward layers are key-value memories
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab517157-5720-4a3c-b2b4-cab81dd2413f · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dcae857-56cf-4d3b-b6a1-080cb97ec7d4 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Impact of positional encoding: Clean and adversarial rademacher complexity for transformers under in-context regression.arXiv preprint arXiv:2512.09275,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b5869f-2af2-4ec9-b95d-8a3193b6c1b5 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Scaling Trends in Language Model Robustness
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848ec80b-7426-4632-bcef-803bd1d40a29 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Scaling Laws for Neural Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365eb461-c285-4511-9f18-379cb93a83d8 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613df603-5f04-4fa7-96fe-b57127c83eb5 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Towards understanding jailbreak attacks in llms: A representation space analysis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6c4c26-d4e2-410a-a026-854060ed12bf · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Alignment-constrained dynamic pruning for llms: Identifying and preserving alignment-critical circuits.arXiv preprint arXiv:2511.07482,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cc0165-c5b4-41c8-9ded-d4cba665f6ca · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Attention sinks and compression valleys in llms are two sides of the same coin.arXiv preprint arXiv:2510.06477,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db00739-e0ec-4a77-90a9-3e8c425eda61 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates A general framework to enhance fine-tuning-based llm unlearning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02be9e3-24d2-422e-8a95-88d5772d846a · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Open Problems in Mechanistic Interpretability
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8291210-311e-4da8-86bc-6484b77d4e1e · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb6ffc55-2fdb-4cc1-b425-57bc77fd86ea · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b339f4c-fdf9-465f-bbfb-7c52b605216c · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c8dac68-8c8c-4713-8032-17643f6529db · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Steering Language Models With Activation Engineering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46aaab2-d5e1-44f5-83a5-5aed2fab45e8 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34c49ad-2ba7-4c9d-b8f7-0132ba070559 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Fast is better than free: Revisiting adversarial training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35456b1c-74db-4559-a505-52e9e6749cf2 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Qwen2.5 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf703032-cc22-4496-aa42-6183a8021c9c · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad144c06-cf9a-4a3f-bdb6-68b24ef6127c · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Adversarial Training: A Survey
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1d80d9-b084-414a-9a8f-51ecc7ee23d1 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates On Prompt-Driven Safeguarding for Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc43658d-ade5-4f40-ac50-deb247ede3c3 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Representation Engineering: A Top-Down Approach to AI Transparency
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9accf4-b6e4-42d9-80ee-6f69908a6df3 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Hence ∥(I−P S)rt∥ ≤1 m0 Mt
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9d2b97-93d2-4982-bbd5-09161d5f1032 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Pruning Convolutional Neural Networks for Resource Efficient Inference
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55167202-7ded-402f-a8cb-7e661079240d · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a7b00d-903e-46e1-96cd-43356c8c70fb · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ff00e6-bf42-4af4-a79f-c8ca2466c3d6 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Softmax is 1/2-lipschitz: A tight bound across all ℓp norms.arXiv preprint arXiv:2510.23012,
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eed1946-ed36-448a-8bc9-922f65ba0adf · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc1f7c2-9e55-40b0-adbf-cf44297a0133 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates The Llama 3 Herd of Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f5e25e-dfc7-4f77-a259-9fca08aa7157 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Toy Models of Superposition
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d6cbe8-68ab-46d9-8c3e-f35e0b6373a3 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a624ce2a-ef17-492b-88e7-954ac420c852 · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Mixat: Combining continuous and discrete adversarial training for llms.arXiv preprint arXiv:2505.16947,
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea2f2d8-3de5-4382-8e84-a4fd76f9be3c · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Towards understanding safety alignment: A mechanistic perspective from safety neurons.arXiv preprint arXiv:2406.14144,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c23aac0-fd84-4bc8-8368-9b8bd121bffc · outbound
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.