Pith. sign in

Paper Citation Record · LEDGER

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

As of 17 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2509.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08000 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.794271Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ad25e298-6573-45c4-85c6-f3496b163bae · outbound

This paper cites , " * write output.state after.block = add.period write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.382342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.382342Z digest=sha256:5f88dde9ac197a0cd4f5a005b6ae5bdefb1c4774c324707aa5d1abe406439fa2

Observation aa8a2ef4-2be3-40f2-a0b0-5e1d45d5e73c · outbound

This paper cites write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.387959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.387959Z digest=sha256:6867c65270c788a08fe096c67ba7f151a74ab0c9ab121024d6ddbb71f33745e0

Observation 5cb1a105-d857-45f5-bd98-f79924b9a913 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Yi: Open Foundation Models by 01.AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.392417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.392417Z digest=sha256:0bf7cbdae091e5c7c43e227de57e2ed54a8efc820a931a8e405e9f3327282f42

Observation 63e14614-66d4-4149-b35d-a0c099c42125 · outbound

This paper cites The Falcon Series of Open Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Falcon Series of Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.397446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.397446Z digest=sha256:7ef16f8bb2fa8e224166d637b69609d645513402f4b36455e74f0a1889fcba0b

Observation 00948fef-d163-4498-bb1e-14fb8d997059 · outbound

This paper cites E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.401443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.401443Z digest=sha256:6add0536a427a8a60c5241c10cc7d132131dbf536df5b4c1b7fc60c15aa40ed5

Observation 8f4c4d3c-4ac2-414d-a932-56dd9d53052f · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.405667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.405667Z digest=sha256:24b6e15a07bca992fea84a5f6d3d9079646d9f3281c0a18ba536fb1e936e966c

Observation af39a24f-0b0c-427b-9445-15d5df9c7f85 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.409950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.409950Z digest=sha256:ba1b658e0f69a15a4ff002e75cf23e37f4b2dae4622bfde2b8540ce58b4ca6f3

Observation a601b59f-ea31-4e22-a836-625c583ea314 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.413726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.413726Z digest=sha256:562d4303b688c855e4c0472180c891ab4246ba3812a5b89216de716cd4ac00b9

Observation da90c372-54be-4d7a-a2d3-25ba07ff3003 · outbound

This paper cites A.; and Hill, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A.; and Hill, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.418714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.418714Z digest=sha256:b159d9c022b04f599a6ccd5836303b7ab23c3f2d5a4aa9c83dde537503492c97

Observation 19928a12-01f7-45d0-af60-33a7fee39670 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.423138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.423138Z digest=sha256:f538c9637e9436eb96eb5d009eb29c76db6467b6f06a68ec47c12d2cd7d2e0a2

Observation 9179db64-dab0-43ef-b687-c3e2d825e606 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.429822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.429822Z digest=sha256:aa83dab8636672fb156c87ef576b8e53548bd7487b04347d200b61831e39edb8

Observation 37912514-015c-4863-89a5-7890860e2bbb · outbound

This paper cites J.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; and Wong, E

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.438338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.438338Z digest=sha256:163a0069e1ed99616a7de51f6d01d3d06b2495bdafe681e9ded13db543db9220

Observation d4017d78-8c14-4e7f-a732-46170f55fb0a · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.442838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.442838Z digest=sha256:b1c0a5522ac9bd62f924a22bb7268f3ee86930151627dc95ba0d30aec4694789

Observation 4859decf-14a6-450c-af55-fa9e40cf246a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.447097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.447097Z digest=sha256:e6f49d09eda7379ed2444c2e75c54c29023c5eeb28e0c3fa82bb023976823d1c

Observation 3d6f2e35-02f0-4520-9bab-70d7fe82028a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.450756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.450756Z digest=sha256:3ea494e87f46115cb3e8d09aa68e9f9320936479e9f7bc00bb5edfe882751a31

Observation c9f372f3-b706-4e65-b29e-3c151c9d96a2 · outbound

This paper cites Multi-Head Attention: Collaborate Instead of Concatenate.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-Head Attention: Collaborate Instead of Concatenate

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.454813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.454813Z digest=sha256:d72160d46a021abf1e9f0a9e2f5af5a515a9ada70912de9e66bbc6908a8c23c0

Observation 029c671a-408e-49ce-810b-7cc95ca2f8b1 · outbound

This paper cites DeepSeek-V3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.459056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.459056Z digest=sha256:d52a73b5cf4cae2167846f88f12aa6395fa5434fbaeceb3c9b5d2a3474dbea79

Observation e86dd41e-677f-408d-91a9-fb22a54ff9e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.463314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.463314Z digest=sha256:d48a0b620d779872de363fdfa790675341f2c0a6cb39f14bc22ec87ef6cb3d66

Observation 3647b421-971a-4cfb-9afb-4c2ce60a1fb3 · outbound

This paper cites C.; Allen, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs C.; Allen, E

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.466791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.466791Z digest=sha256:c423acfa45c939c01f5d8f6fef30604412a9c0ab84b4fdafc4caa2bc6b9b69eb

Observation 40ca8320-99af-4c93-ae36-9989f479325e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.470899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.470899Z digest=sha256:3f4c1213ca6300ebe286a595ca82ad3eb05555109c72ae5a444dcd2884a747d6

Observation c35178b1-326b-458b-af3d-87ab63237d04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.246678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.475467Z digest=sha256:af277a10c71bbccd17d4a9afdc6cc5a20756cfb5e5fb380c2a854758b9e06bc2

Observation 45c5d733-ba60-419d-916d-8e2335e4d032 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.231653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.479137Z digest=sha256:1de8c8de06f638ed7e6db9e96a46ed5562b50683acf0752290fa22a14f73d54b

Observation 4c40d6e0-a570-4127-b746-4696d7e71c9f · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.482899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.482899Z digest=sha256:67039b5766830b7b7f02f96c18b018a0fee623aab7aa66471e6c9168cb8fb95b

Observation 9009f364-c983-445c-afc9-abd99ed4df17 · outbound

This paper cites The Llama 3 Herd of Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.486772Z digest=sha256:5cd48b4a07f99604bc76d6b372f50a463cdcb76e21352456786902e54c6decdd

Observation 31fbdf9e-ce32-419c-bb86-c8bc84d5287a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.218168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.491734Z digest=sha256:8ea1c2f2093ad76a69482f2894df3441e0ce867a4d0b130f184a15fe10bb2d06

Observation 4114cf85-cd5d-4040-aece-c1df0a48ff81 · outbound

This paper cites Aloe: A Family of Fine-tuned Open Healthcare LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aloe: A Family of Fine-tuned Open Healthcare LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.495639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.495639Z digest=sha256:4dac832e405648393946483ba5f5a88d59a486297bc343ad75b51a97229f3836

Observation d062791c-dbb2-4b6b-b6a7-21ede5e85ef2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.500101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.500101Z digest=sha256:ffef80f2ad1895cdd9610bbd8b8a9a76c773e06622e5bd47d83813739f8aec75

Observation 48e79834-71bf-48de-b8f7-b22018b76567 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.504361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.504361Z digest=sha256:6bd50227f2e7e86582410cfcc3a6a2b97192cf36ed129dc06a9d92873da659d5

Observation 9980664b-b9ad-4173-a854-2838d6d1b1b1 · outbound

This paper cites Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.508322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.508322Z digest=sha256:65f8950dd747327c3dce612fab560f0a0e5aa9b2a67f0b3125ba651efd0b4d85

Observation bc112cc0-46f6-4001-9b2d-bd1e020d062c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.512022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.512022Z digest=sha256:5886bcf3ea47ba952a5903abd46ba596a609999bc8ecc08f1740c8aa3ac44cc2

Observation 468945ff-c394-4a8b-aae3-214d259cdeb7 · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.520496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.520496Z digest=sha256:bda5c0aace8c484fed36acfaeabee0691b84b96ecb37c83ec3e3056774836280

Observation f5f1cb5b-891b-44da-950b-d6cac8a6d69c · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.524521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.524521Z digest=sha256:c29334ddcf2cd499c6f0c3671d04330cff9e5b5765ea408e17bb76ad186b5c63

Observation 18be176f-584e-40d0-86c2-d944de26f4f9 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.533528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.533528Z digest=sha256:12ccaf4e5d1f744e679294c2e6ddfc2ef6f4918e0580017c4f1bf1655cccdcf0

Observation d04578d6-677e-4791-b157-4fac8425e1a4 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.205244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.537145Z digest=sha256:89a0374dd8ef5ffd93dd815060d4a94e135f555678c497b8452e1c85aa82e303

Observation 8dbbecfb-e8b7-4477-a8ea-9051a1b85c0a · outbound

This paper cites Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.540991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.540991Z digest=sha256:0b2855df1c7543758f1ebaa2cf5c5c43799f6f92862615426cdd0e372e417902

Observation 10c4157a-e12b-43ce-80dc-ec8054f35342 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.545273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.545273Z digest=sha256:42c6d77d87bd61dbac4b362b5b1934cc39408b285a09f2f52d54f64f66347ef3

Observation 96d1e257-d944-45cf-9a65-8e2a6c5847ff · outbound

This paper cites Mistral 7B.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Mistral 7B

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.548883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.548883Z digest=sha256:d348191ccee3055893e14dc4718a6b0f11fca50302fe57202d92186b9812af04

Observation 70274a89-1af6-4411-8775-a947129265ee · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.192162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.553625Z digest=sha256:66e0a57ce3cd0002dccb2459c810a32d6e57d4a9f4ce3e82bfc2e5547cf3e29b

Observation e6ae978d-4772-4f3f-a900-1d9ec49f6752 · outbound

This paper cites LLM Embeddings for Deep Learning on Tabular Data.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings for Deep Learning on Tabular Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.558732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.558732Z digest=sha256:d04b2390996e29d51a8fa262bccdf960310fb8cf2271972749b55b36152e606a

Observation 3dde08bb-842c-49fb-869d-762eabf1e434 · outbound

This paper cites Robust Distortion-free Watermarks for Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Robust Distortion-free Watermarks for Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.562902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.562902Z digest=sha256:98542c3b6b7248962a896bdcd6dbec21690c07529e0ffbea2ee28c293bb2e0d1

Observation 46e0b858-00b3-4787-9db8-d7fe7ad30e9e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.178690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.566932Z digest=sha256:cf610a20d97078ff6bf993a2c9ccb5b882c348f693af5a104f7c757424492034

Observation 926a9815-7dac-4d80-a119-e24de3292678 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.570589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.570589Z digest=sha256:ae6cd73d1ae3b389a899458ef5ae62d0fefdb58b32f4e4e68c7589ffc13e01b3

Observation c75d8f18-27c6-478c-aa7e-ace93bf76c7e · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.574328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.574328Z digest=sha256:4d99a53b83726ace01a6c04f445c846fd5c60bd1d0f160446468213c92802f98

Observation b91a5ea8-9ede-4d8a-87b6-4166cec6f27f · outbound

This paper cites TeleLoRA: Teleporting Model-Specific Alignment Across LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs TeleLoRA: Teleporting Model-Specific Alignment Across LLMs

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.529377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.582429Z digest=sha256:4a0bc3818d3fdf66e639d0c3998481d3a62a4890f61a7c06502ef60b47a2735f

Observation 2d16c294-5216-4998-9779-f7dbae85971a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:25:36.509939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.586017Z digest=sha256:10247dd5f348a6c9f8f1ac0d19ada3b9ab47e78b62165723c67a295947f49dfe

Observation 0e368ff1-ba39-4568-88d8-30414caf00aa · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.589876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.589876Z digest=sha256:32729d1a8411afaf898b3115f4ba5fe9d406190def256e11a602171f26d590d4

Observation c47c0b94-a64a-4edc-8d32-104d6ebbdb54 · outbound

This paper cites SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.597800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.597800Z digest=sha256:75f078338b855ed2032398c5435c2eb63dfd8ee1a90c72969d7dc3947ac92d05

Observation ca700848-885d-4c03-8a6f-32e5254d0726 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.601622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.601622Z digest=sha256:df981fa208037ff926806f079f25a3962728c3b768fbdc98bed51553315efceb

Observation 98430e03-cf23-443b-9007-c53b22089ba1 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.612838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.612838Z digest=sha256:329e312ba2f3664a1d8a3f07643d99011a19b5a4cbac5ec7b315ac0b3daeed16

Observation ca845d39-ef48-4bbc-9588-f5a40d6a5c14 · outbound

This paper cites R.; and Papernot, N.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs R.; and Papernot, N

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.616456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.616456Z digest=sha256:6398a2a78eed7f0fec86a8279d5f50143d2f33119270124df0f3c232e36d460e

Observation 833a1260-eedb-4242-bd8a-bccd6a290aab · outbound

This paper cites Topic-Based Watermarks for Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Topic-Based Watermarks for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.620427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.620427Z digest=sha256:735370222ab8a41197574f2ac9073c8fefc1d5beb1802e7ee0acd1c3f1b3ef41

Observation eb29432f-9822-466a-b1fa-8fa53f9fb69f · outbound

This paper cites J.; Hassani, H.; Robey, A.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; Hassani, H.; Robey, A.; and Wong, E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.625006Z digest=sha256:edbc06d054abe867ea22b479db865e8aae89bb3ed0bccb3f223e55a212b254f3

Observation 0edacc3f-f76f-4b44-b7a7-038ce3b123f5 · outbound

This paper cites In-Context Unlearning: Language Models as Few Shot Unlearners.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs In-Context Unlearning: Language Models as Few Shot Unlearners

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.629057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.629057Z digest=sha256:d53644c3d16b1e8b987d4eb51a5d308e34f631ecf7ff2ab05db3e0d01a4d6ba6

Observation c3ff32b1-5882-4710-a2ed-0495bdeabd70 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.148810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.633450Z digest=sha256:e1e1e6cd3c618f0a0ff20c5446c21bb32afd20a45daf6a4070ff9bc997e30196

Observation ddceaa23-7b54-4660-a7df-0af32dfa13e8 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.637678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.637678Z digest=sha256:ad8a6d7eb5f380f667477e94d72176727cff33004e0be97fd87624b14d16aaa2

Observation 53b212db-a8ec-46fa-84a7-f7fea271f2ea · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.641624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.641624Z digest=sha256:f1d990c2cdfe86b7407cc63dc4f1af70d27d468967b11743d50793c03d8b0624

Observation 3ce97506-2eff-4deb-b347-a0fcd6367b25 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.134159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.645915Z digest=sha256:559bc1b0943f4179d0f1046ca1deeca98832bc29ae1849e744a6198fdd6b8a21

Observation e76398f4-5be7-4001-bc36-72747e6bb4ad · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.649342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.649342Z digest=sha256:7e1feb2014589cd8e46bac43eaa9b6a3ebf95960d6c4d706f51b560fce08b754

Observation 0cf543b4-961c-4339-a167-43334dd940df · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.653241Z digest=sha256:b480e6b33852f73206f72da8c8eada425ce1dba6023aad4e2db0c8ec2450b403

Observation 529d0413-b7ce-4b21-969d-7682e60c144b · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.657150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.657150Z digest=sha256:2a957f2c564069870dec3d6ec2eb8d88bab0abc376b8c1622c66c7d4664cf4f7

Observation 0054103c-e053-4a9b-87c5-b95df03abe64 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.661660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.661660Z digest=sha256:00b30720365c051dfcf606c8902a31dc9e4f601be2f0d53428191625c0ac6b10

Observation 399fe1e5-aac0-4f27-a586-77d9448e3b04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.110048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.666208Z digest=sha256:e3e93e446a533706c417bda520b0cbdc4af7f91e4337ddb95df58304c1c71e57

Observation 36314a39-cde8-4e78-9130-60869cbd74c4 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.669919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.669919Z digest=sha256:adeb3e5ef1a61aadef72f46ec32e45a2d3189004f6a931c77af8d2d5a484f8f1

Observation 9ac3813f-9ba9-4c50-a3b5-403f03f45be1 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.097746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.673744Z digest=sha256:81aadf0c8f3d335720d64286ab98e3ef973b3c1d7d04f8cb9cf7ede6126f0f9e

Observation d1af0da0-2d5e-4a04-80e3-98009e17e73c · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.677147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.677147Z digest=sha256:5826ee6b25180a6850cd2a92a6b8f10b9689621ebd4dcb3484260033c05afce6

Observation 0df81bd0-f189-401c-b4da-e05e884a8bfa · outbound

This paper cites Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.160505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.681487Z digest=sha256:068785ffbaebe0bf4a4db70dc27ea6181794a9a0073b97c61e3a44d51824a1ad

Observation 6d520293-7269-4cab-88c1-5d92ee86b72c · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A StrongREJECT for Empty Jailbreaks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.685837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.685837Z digest=sha256:5306e64d8407bfb83e78bf6abaa0172363c5bfc5156493931cd0e9479677ea2a

Observation b5dffa9a-35b4-4ccc-af4a-d3635da1ee57 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.689623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.689623Z digest=sha256:2732ba98a501e6c17ac25dcd10bd200b736330bc8980edff07c9dd0f23031a89

Observation 8e6337e5-00e5-4a69-9bae-f25b6ef9dfa9 · outbound

This paper cites Gemma 3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Gemma 3 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.694213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.694213Z digest=sha256:491543e051b3ea6c99ef22ecf761a507bdd3b3cc1b490ea365b9e3fd905b9069

Observation 9af2336c-692a-4b0d-a3fa-f783b16507e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.698672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.698672Z digest=sha256:636b2e42812a500ed5a6d95de6bf84d20575de87f7e7ebd1e00c782f685c50ac

Observation 5655b3f3-d956-4b39-99a4-a1fa20e609ec · outbound

This paper cites Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.702135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.702135Z digest=sha256:e5a843532c34c52fbe11fe4617ed3c1d298cb10e23dcd87956608f86ea57b9eb

Observation 874015b7-cb02-40d5-924a-cdcdf96b9abf · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.706234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.706234Z digest=sha256:3a78d17c815ff507536319996937803dce829e07d2c28b5390999a21c0ea4994

Observation 51df8078-f20f-4b92-8d42-72156df20f57 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.713916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.713916Z digest=sha256:68a60c8ed9fa0d4cc06c89843a1a3fe64f5f961cc815c5541e233da5a3a4fa68

Observation 54f0fe5f-a4cd-4e81-974e-d7c8efcbbd7b · outbound

This paper cites SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.717645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.717645Z digest=sha256:8305a3cc8eddd1096b93e35ee85a2fb33eaa64360e62c735c5e68fb9b55424a1

Observation e8441d64-02c5-48ec-880f-9169d9c2e1e5 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.085522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.721358Z digest=sha256:10b2c2a05c2cf8b0f7decf6dc9033b61e44416557388b9fc89a7390ca791939f

Observation b33b4b48-4e00-4557-b1d1-80b2dd061767 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.069441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.724987Z digest=sha256:8291a705ce13940a0caa4642b325fda99534342311d99fe95e60805c0e330781

Observation 7744cce7-13fd-4e4c-a439-05382bd6f77b · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.728642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.728642Z digest=sha256:d9bfd26a39f8c79da5278104acaa8d8d44fe516a1e1fd23ad29b418e649c149e

Observation 6600dd76-3429-4677-9637-7463b1b07cf3 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.053075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.732524Z digest=sha256:c33c709b2b91c769943d7ada3c18e7b93e61b6d1b2ac3bff043a98e4b6e148d4

Observation 602c5ab8-683e-42e2-b488-b0ba8192c37d · outbound

This paper cites Qwen3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Qwen3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.736230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.736230Z digest=sha256:37cd009d00c3d2b1babf1297a5dea79059e56fd456d5cb686f268ad23fa46d4a

Observation 6737df63-0dd1-4039-b648-154aac2f6535 · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.740179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.740179Z digest=sha256:3e80fedfab1c28460d5a5563157f5d6571b072493ad79d9fbc9698a15a03d015

Observation 7bd9f140-c927-4bba-9625-210afbdebf6d · outbound

This paper cites The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.744416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.744416Z digest=sha256:a12b1993103d941788c4cb9c3cfb72765b7ffd5e993874f4e5757d89b75b3ad0

Observation a969486c-572c-44f7-aaab-60fbad43883b · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.748405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.748405Z digest=sha256:d6ac64ebf66bb436cad964495460dfcf1dc1e38a1630da8566b8c5517f8ddd5e

Observation 6b76a273-6f9b-4777-8418-4253fb003343 · outbound

This paper cites Position: Editing Large Language Models Poses Serious Safety Risks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Position: Editing Large Language Models Poses Serious Safety Risks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.752237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.752237Z digest=sha256:7e1384a18f95f309c5c00987be38b8ef35cbae1f7d9c45c676b7ab36a7432e90

Observation c7c025be-ad8a-434b-ab31-d3acc0e091cb · outbound

This paper cites Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.760832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.760832Z digest=sha256:df2351a52eb234285150dd82a00696703a25d28738e0216eec16ce475c185c70

Observation 2d1cdad1-23db-49bf-973b-2304ef7d8878 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.765564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.765564Z digest=sha256:6829866434ced753cc265ae9c5f767a3bd8fdfd2823c0f4821928296d23a16dc

Observation f7125f4b-9524-4f2f-bb51-c7472aa3e70c · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.773425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.773425Z digest=sha256:90afa7c01baeda8eec54e07e2b184e37843ba8e1bc61527f3c414552b6d2fffb

Observation 51ede104-63ed-4a65-95df-a7068fb2edba · outbound

This paper cites LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:35.868451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.777182Z digest=sha256:6cf4a08c410cbd0a3745882ca9283b0fe7684616ca3d85c811bc6536542f8de2

Observation 7e84904e-ce8f-4cf2-b123-5d300a9c3e9d · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.038890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.781938Z digest=sha256:29eb4b2d30e1d860c4e5a46ffcd32f1bbdfe5d92659530722711a8c196461475

Observation 7d03a866-8998-4115-942f-ca7cbea3c008 · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.786739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.786739Z digest=sha256:30fd57e89042d13ebe428d99f73971b7fee1f629a70c6bdf61dc4fad427ce00b

Observation 2511efdb-5ff7-41fc-acc9-8bd3f5727caf · outbound

This paper cites LIMA: Less Is More for Alignment.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LIMA: Less Is More for Alignment

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.790232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.790232Z digest=sha256:0bf96bb3335e426edb91658541a0a9c825939ada11967a9b761a2d80e2f0cbe2

Observation 82f303b1-65e0-4869-8233-fc05e3500b2e · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.794271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.794271Z digest=sha256:49c912fe26b41f73885185776ee6f310a820bb16a2e2c737163add4883bd5092

Pith citing papers

Observation 9d9a0935-e55e-49d7-b954-564f6fade08c · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.978682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:c12317763bb994e4ed7a0d6f27fdb508536478e9e68940db18fd0187c972091f