Pith. sign in

Paper Citation Record · LEDGER

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

As of 18 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2509.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08000 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.794271Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ad25e298-6573-45c4-85c6-f3496b163bae · outbound

This paper cites , " * write output.state after.block = add.period write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.382342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.382342Z digest=sha256:5f88dde9ac197a0cd4f5a005b6ae5bdefb1c4774c324707aa5d1abe406439fa2

Observation aa8a2ef4-2be3-40f2-a0b0-5e1d45d5e73c · outbound

This paper cites write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.387959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.387959Z digest=sha256:6867c65270c788a08fe096c67ba7f151a74ab0c9ab121024d6ddbb71f33745e0

Observation 5cb1a105-d857-45f5-bd98-f79924b9a913 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Yi: Open Foundation Models by 01.AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.392417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.392417Z digest=sha256:ba416d0d6bb1e58826488c3566f805ebdcbf29d3afd8d6e09bfc12269f816249

Observation 63e14614-66d4-4149-b35d-a0c099c42125 · outbound

This paper cites The Falcon Series of Open Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Falcon Series of Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.397446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.397446Z digest=sha256:7ef16f8bb2fa8e224166d637b69609d645513402f4b36455e74f0a1889fcba0b

Observation 00948fef-d163-4498-bb1e-14fb8d997059 · outbound

This paper cites E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.401443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.401443Z digest=sha256:6add0536a427a8a60c5241c10cc7d132131dbf536df5b4c1b7fc60c15aa40ed5

Observation 8f4c4d3c-4ac2-414d-a932-56dd9d53052f · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.405667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.405667Z digest=sha256:5a2ac2bce2ed46c378ad80bedd8899b305b1bbf3bc7e5f69c6eb8bce8cabd0fc

Observation af39a24f-0b0c-427b-9445-15d5df9c7f85 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.409950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.409950Z digest=sha256:ba1b658e0f69a15a4ff002e75cf23e37f4b2dae4622bfde2b8540ce58b4ca6f3

Observation a601b59f-ea31-4e22-a836-625c583ea314 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.413726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.413726Z digest=sha256:562d4303b688c855e4c0472180c891ab4246ba3812a5b89216de716cd4ac00b9

Observation da90c372-54be-4d7a-a2d3-25ba07ff3003 · outbound

This paper cites A.; and Hill, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A.; and Hill, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.418714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.418714Z digest=sha256:b159d9c022b04f599a6ccd5836303b7ab23c3f2d5a4aa9c83dde537503492c97

Observation 19928a12-01f7-45d0-af60-33a7fee39670 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.423138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.423138Z digest=sha256:f538c9637e9436eb96eb5d009eb29c76db6467b6f06a68ec47c12d2cd7d2e0a2

Observation 9179db64-dab0-43ef-b687-c3e2d825e606 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.429822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.429822Z digest=sha256:aa83dab8636672fb156c87ef576b8e53548bd7487b04347d200b61831e39edb8

Observation 37912514-015c-4863-89a5-7890860e2bbb · outbound

This paper cites J.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; and Wong, E

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.438338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.438338Z digest=sha256:163a0069e1ed99616a7de51f6d01d3d06b2495bdafe681e9ded13db543db9220

Observation d4017d78-8c14-4e7f-a732-46170f55fb0a · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.442838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.442838Z digest=sha256:0de7874a28e07882d1c6f8b85c7822ca2c3e2b56dcfb3d11f741b39741279dbe

Observation 4859decf-14a6-450c-af55-fa9e40cf246a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.447097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.447097Z digest=sha256:e6f49d09eda7379ed2444c2e75c54c29023c5eeb28e0c3fa82bb023976823d1c

Observation 3d6f2e35-02f0-4520-9bab-70d7fe82028a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.450756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.450756Z digest=sha256:3ea494e87f46115cb3e8d09aa68e9f9320936479e9f7bc00bb5edfe882751a31

Observation c9f372f3-b706-4e65-b29e-3c151c9d96a2 · outbound

This paper cites Multi-Head Attention: Collaborate Instead of Concatenate.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-Head Attention: Collaborate Instead of Concatenate

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.454813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.454813Z digest=sha256:d72160d46a021abf1e9f0a9e2f5af5a515a9ada70912de9e66bbc6908a8c23c0

Observation 029c671a-408e-49ce-810b-7cc95ca2f8b1 · outbound

This paper cites DeepSeek-V3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.459056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.459056Z digest=sha256:d52a73b5cf4cae2167846f88f12aa6395fa5434fbaeceb3c9b5d2a3474dbea79

Observation e86dd41e-677f-408d-91a9-fb22a54ff9e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.463314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.463314Z digest=sha256:d48a0b620d779872de363fdfa790675341f2c0a6cb39f14bc22ec87ef6cb3d66

Observation 3647b421-971a-4cfb-9afb-4c2ce60a1fb3 · outbound

This paper cites C.; Allen, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs C.; Allen, E

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.466791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.466791Z digest=sha256:c423acfa45c939c01f5d8f6fef30604412a9c0ab84b4fdafc4caa2bc6b9b69eb

Observation 40ca8320-99af-4c93-ae36-9989f479325e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.470899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.470899Z digest=sha256:3f4c1213ca6300ebe286a595ca82ad3eb05555109c72ae5a444dcd2884a747d6

Observation c35178b1-326b-458b-af3d-87ab63237d04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.246678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.475467Z digest=sha256:8239eef5a3cbcc566389902e0f9728656c7ccf4959d42308039a85badf1b0747

Observation 45c5d733-ba60-419d-916d-8e2335e4d032 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.231653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.479137Z digest=sha256:ae9baeb3bdfe92a758d226d0aa62463af53c3cc84f63fe9c3fc0ad22be518c0b

Observation 4c40d6e0-a570-4127-b746-4696d7e71c9f · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.482899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.482899Z digest=sha256:67039b5766830b7b7f02f96c18b018a0fee623aab7aa66471e6c9168cb8fb95b

Observation 9009f364-c983-445c-afc9-abd99ed4df17 · outbound

This paper cites The Llama 3 Herd of Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.486772Z digest=sha256:5cd48b4a07f99604bc76d6b372f50a463cdcb76e21352456786902e54c6decdd

Observation 31fbdf9e-ce32-419c-bb86-c8bc84d5287a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.218168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.491734Z digest=sha256:70684a3545004c691407d99bfce6539da9f8cc71eff2787555bf245033bd2084

Observation 4114cf85-cd5d-4040-aece-c1df0a48ff81 · outbound

This paper cites Aloe: A Family of Fine-tuned Open Healthcare LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aloe: A Family of Fine-tuned Open Healthcare LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.495639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.495639Z digest=sha256:5dd0f43ece8e6fbfd12129b1ba6a3f74fa57561bacdec2c0fe25235da0577e71

Observation d062791c-dbb2-4b6b-b6a7-21ede5e85ef2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.500101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.500101Z digest=sha256:ffef80f2ad1895cdd9610bbd8b8a9a76c773e06622e5bd47d83813739f8aec75

Observation 48e79834-71bf-48de-b8f7-b22018b76567 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.504361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.504361Z digest=sha256:6bd50227f2e7e86582410cfcc3a6a2b97192cf36ed129dc06a9d92873da659d5

Observation 9980664b-b9ad-4173-a854-2838d6d1b1b1 · outbound

This paper cites Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.508322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.508322Z digest=sha256:a01f2f8b27b9f355f8731d8ad801c475fc0ae2590e94b18ddbd1f204c87b7833

Observation bc112cc0-46f6-4001-9b2d-bd1e020d062c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.512022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.512022Z digest=sha256:e471c78bb0da72c77d1f85534083f8fc319d690aef4eca146c07b1cfad2bc66e

Observation 468945ff-c394-4a8b-aae3-214d259cdeb7 · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.520496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.520496Z digest=sha256:0fb14974d96d110c3f13853013246a3d7aaa3cbbe1c1674b085300bc54909756

Observation f5f1cb5b-891b-44da-950b-d6cac8a6d69c · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.524521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.524521Z digest=sha256:c29334ddcf2cd499c6f0c3671d04330cff9e5b5765ea408e17bb76ad186b5c63

Observation 18be176f-584e-40d0-86c2-d944de26f4f9 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.533528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.533528Z digest=sha256:12ccaf4e5d1f744e679294c2e6ddfc2ef6f4918e0580017c4f1bf1655cccdcf0

Observation d04578d6-677e-4791-b157-4fac8425e1a4 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.205244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.537145Z digest=sha256:1022da61b77e5df1e9d4d0ee20849e03389fa93284aab5e9cd487b515bc9d383

Observation 8dbbecfb-e8b7-4477-a8ea-9051a1b85c0a · outbound

This paper cites Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.540991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.540991Z digest=sha256:0b2855df1c7543758f1ebaa2cf5c5c43799f6f92862615426cdd0e372e417902

Observation 10c4157a-e12b-43ce-80dc-ec8054f35342 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.545273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.545273Z digest=sha256:42c6d77d87bd61dbac4b362b5b1934cc39408b285a09f2f52d54f64f66347ef3

Observation 96d1e257-d944-45cf-9a65-8e2a6c5847ff · outbound

This paper cites Mistral 7B.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Mistral 7B

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.548883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.548883Z digest=sha256:ab222caf7f17a2ff4efecfaceb81f723b81d5080334690a650a3163f8c38a008

Observation 70274a89-1af6-4411-8775-a947129265ee · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.192162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.553625Z digest=sha256:e2367a344cfd42374b6025546de4f3f5cd289e075767830d44779b84bca4d476

Observation e6ae978d-4772-4f3f-a900-1d9ec49f6752 · outbound

This paper cites LLM Embeddings for Deep Learning on Tabular Data.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings for Deep Learning on Tabular Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.558732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.558732Z digest=sha256:d04b2390996e29d51a8fa262bccdf960310fb8cf2271972749b55b36152e606a

Observation 3dde08bb-842c-49fb-869d-762eabf1e434 · outbound

This paper cites Robust Distortion-free Watermarks for Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Robust Distortion-free Watermarks for Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.562902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.562902Z digest=sha256:98542c3b6b7248962a896bdcd6dbec21690c07529e0ffbea2ee28c293bb2e0d1

Observation 46e0b858-00b3-4787-9db8-d7fe7ad30e9e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.178690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.566932Z digest=sha256:8a7bec34a4dc5f63880aae19d054111c03cf1c876a7e5a3d60d3eb5b173ff5a3

Observation 926a9815-7dac-4d80-a119-e24de3292678 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.570589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.570589Z digest=sha256:c5da2c7cacdc3ab701bf10a1b0f558b3ac8d2b083b3cfbd19b9d5cd68ceff0a7

Observation c75d8f18-27c6-478c-aa7e-ace93bf76c7e · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.574328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.574328Z digest=sha256:4d99a53b83726ace01a6c04f445c846fd5c60bd1d0f160446468213c92802f98

Observation b91a5ea8-9ede-4d8a-87b6-4166cec6f27f · outbound

This paper cites TeleLoRA: Teleporting Model-Specific Alignment Across LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs TeleLoRA: Teleporting Model-Specific Alignment Across LLMs

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.529377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.582429Z digest=sha256:a60aeda53017fa54378c63806d3bf05f490e29299d60ccdb8e688988d34b3c8d

Observation 2d16c294-5216-4998-9779-f7dbae85971a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:25:36.509939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.586017Z digest=sha256:0fa3352963196dac33f9a1968f03733c165d90446830f9830c7097022e31bdf4

Observation 0e368ff1-ba39-4568-88d8-30414caf00aa · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.589876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.589876Z digest=sha256:6cdd0601f37c0fa73af31211f1626d58f22695ed8394f068bd301264a327c0d6

Observation c47c0b94-a64a-4edc-8d32-104d6ebbdb54 · outbound

This paper cites SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.597800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.597800Z digest=sha256:75f078338b855ed2032398c5435c2eb63dfd8ee1a90c72969d7dc3947ac92d05

Observation ca700848-885d-4c03-8a6f-32e5254d0726 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.601622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.601622Z digest=sha256:df981fa208037ff926806f079f25a3962728c3b768fbdc98bed51553315efceb

Observation 98430e03-cf23-443b-9007-c53b22089ba1 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.612838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.612838Z digest=sha256:329e312ba2f3664a1d8a3f07643d99011a19b5a4cbac5ec7b315ac0b3daeed16

Observation ca845d39-ef48-4bbc-9588-f5a40d6a5c14 · outbound

This paper cites R.; and Papernot, N.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs R.; and Papernot, N

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.616456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.616456Z digest=sha256:6398a2a78eed7f0fec86a8279d5f50143d2f33119270124df0f3c232e36d460e

Observation 833a1260-eedb-4242-bd8a-bccd6a290aab · outbound

This paper cites Topic-Based Watermarks for Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Topic-Based Watermarks for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.620427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.620427Z digest=sha256:735370222ab8a41197574f2ac9073c8fefc1d5beb1802e7ee0acd1c3f1b3ef41

Observation eb29432f-9822-466a-b1fa-8fa53f9fb69f · outbound

This paper cites J.; Hassani, H.; Robey, A.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; Hassani, H.; Robey, A.; and Wong, E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.625006Z digest=sha256:d67528bbd2f40779f0179b60bf075159b3c03c70788c345dbec9604b667b797b

Observation 0edacc3f-f76f-4b44-b7a7-038ce3b123f5 · outbound

This paper cites In-Context Unlearning: Language Models as Few Shot Unlearners.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs In-Context Unlearning: Language Models as Few Shot Unlearners

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.629057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.629057Z digest=sha256:752b47084af82c4bfb21f09e2e7b3c1d489c14990476d31314c3ed53b9419534

Observation c3ff32b1-5882-4710-a2ed-0495bdeabd70 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.148810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.633450Z digest=sha256:f8916ff624d7bb2901ea2060f0e00bcf7448d3412eedc3d4297cc509ca16fdf9

Observation ddceaa23-7b54-4660-a7df-0af32dfa13e8 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.637678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.637678Z digest=sha256:ad8a6d7eb5f380f667477e94d72176727cff33004e0be97fd87624b14d16aaa2

Observation 53b212db-a8ec-46fa-84a7-f7fea271f2ea · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.641624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.641624Z digest=sha256:f1d990c2cdfe86b7407cc63dc4f1af70d27d468967b11743d50793c03d8b0624

Observation 3ce97506-2eff-4deb-b347-a0fcd6367b25 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.134159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.645915Z digest=sha256:e4c685a052d9b286724f7f76ba9eef4da17c34ad51199d4ede29688a09a9c250

Observation e76398f4-5be7-4001-bc36-72747e6bb4ad · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.649342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.649342Z digest=sha256:7e1feb2014589cd8e46bac43eaa9b6a3ebf95960d6c4d706f51b560fce08b754

Observation 0cf543b4-961c-4339-a167-43334dd940df · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.653241Z digest=sha256:e15f8ef33816f95d9f34afc54565cfaff3b696e1c9051e8aa9195ce048cbddd8

Observation 529d0413-b7ce-4b21-969d-7682e60c144b · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.657150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.657150Z digest=sha256:2a957f2c564069870dec3d6ec2eb8d88bab0abc376b8c1622c66c7d4664cf4f7

Observation 0054103c-e053-4a9b-87c5-b95df03abe64 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.661660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.661660Z digest=sha256:00b30720365c051dfcf606c8902a31dc9e4f601be2f0d53428191625c0ac6b10

Observation 399fe1e5-aac0-4f27-a586-77d9448e3b04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.110048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.666208Z digest=sha256:3ac3faa9e04f7d153cfe3a2da6afd849322ad5cd6a31fba1b6657ea369a6bc13

Observation 36314a39-cde8-4e78-9130-60869cbd74c4 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.669919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.669919Z digest=sha256:adeb3e5ef1a61aadef72f46ec32e45a2d3189004f6a931c77af8d2d5a484f8f1

Observation 9ac3813f-9ba9-4c50-a3b5-403f03f45be1 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.097746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.673744Z digest=sha256:0f16cfd1866b8f811edace4c205c0442e407b64ee59dde7e1d7880cdc4dc735b

Observation d1af0da0-2d5e-4a04-80e3-98009e17e73c · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.677147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.677147Z digest=sha256:5826ee6b25180a6850cd2a92a6b8f10b9689621ebd4dcb3484260033c05afce6

Observation 0df81bd0-f189-401c-b4da-e05e884a8bfa · outbound

This paper cites Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.160505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.681487Z digest=sha256:4c0b40518f3a5eec42299fa1942c0369f7c737af281ae03a0eaa5b5e4856733d

Observation 6d520293-7269-4cab-88c1-5d92ee86b72c · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A StrongREJECT for Empty Jailbreaks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.685837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.685837Z digest=sha256:5306e64d8407bfb83e78bf6abaa0172363c5bfc5156493931cd0e9479677ea2a

Observation b5dffa9a-35b4-4ccc-af4a-d3635da1ee57 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.689623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.689623Z digest=sha256:2732ba98a501e6c17ac25dcd10bd200b736330bc8980edff07c9dd0f23031a89

Observation 8e6337e5-00e5-4a69-9bae-f25b6ef9dfa9 · outbound

This paper cites Gemma 3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Gemma 3 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.694213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.694213Z digest=sha256:491543e051b3ea6c99ef22ecf761a507bdd3b3cc1b490ea365b9e3fd905b9069

Observation 9af2336c-692a-4b0d-a3fa-f783b16507e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.698672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.698672Z digest=sha256:636b2e42812a500ed5a6d95de6bf84d20575de87f7e7ebd1e00c782f685c50ac

Observation 5655b3f3-d956-4b39-99a4-a1fa20e609ec · outbound

This paper cites Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.702135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.702135Z digest=sha256:e5a843532c34c52fbe11fe4617ed3c1d298cb10e23dcd87956608f86ea57b9eb

Observation 874015b7-cb02-40d5-924a-cdcdf96b9abf · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.706234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.706234Z digest=sha256:3a78d17c815ff507536319996937803dce829e07d2c28b5390999a21c0ea4994

Observation 51df8078-f20f-4b92-8d42-72156df20f57 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.713916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.713916Z digest=sha256:68a60c8ed9fa0d4cc06c89843a1a3fe64f5f961cc815c5541e233da5a3a4fa68

Observation 54f0fe5f-a4cd-4e81-974e-d7c8efcbbd7b · outbound

This paper cites SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.717645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.717645Z digest=sha256:86cff4deddd177b69bdda1e4f5e9eeffff31fc88ad16bd58076ab6673bf09145

Observation e8441d64-02c5-48ec-880f-9169d9c2e1e5 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.085522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.721358Z digest=sha256:23efc6fadc980e6f37838bbc1ae30390a3e83faa0a9f7b8b1fab6be9c6d442b4

Observation b33b4b48-4e00-4557-b1d1-80b2dd061767 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.069441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.724987Z digest=sha256:9914c465242245471a7855ba1c9e5738ef8e237344b3e8f7ffb060740ab68dcc

Observation 7744cce7-13fd-4e4c-a439-05382bd6f77b · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.728642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.728642Z digest=sha256:f16491819cba619954adc62aaff85b10cc8ad6837c9a6d3c97e31ff4552d7f2a

Observation 6600dd76-3429-4677-9637-7463b1b07cf3 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.053075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.732524Z digest=sha256:1ecd39dc1dfe6801fd92437e5265f5f4caa49c570cf767272345db0868a298dd

Observation 602c5ab8-683e-42e2-b488-b0ba8192c37d · outbound

This paper cites Qwen3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Qwen3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.736230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.736230Z digest=sha256:37cd009d00c3d2b1babf1297a5dea79059e56fd456d5cb686f268ad23fa46d4a

Observation 6737df63-0dd1-4039-b648-154aac2f6535 · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.740179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.740179Z digest=sha256:3e80fedfab1c28460d5a5563157f5d6571b072493ad79d9fbc9698a15a03d015

Observation 7bd9f140-c927-4bba-9625-210afbdebf6d · outbound

This paper cites The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.744416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.744416Z digest=sha256:a12b1993103d941788c4cb9c3cfb72765b7ffd5e993874f4e5757d89b75b3ad0

Observation a969486c-572c-44f7-aaab-60fbad43883b · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.748405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.748405Z digest=sha256:d6ac64ebf66bb436cad964495460dfcf1dc1e38a1630da8566b8c5517f8ddd5e

Observation 6b76a273-6f9b-4777-8418-4253fb003343 · outbound

This paper cites Position: Editing Large Language Models Poses Serious Safety Risks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Position: Editing Large Language Models Poses Serious Safety Risks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.752237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.752237Z digest=sha256:0b9ae853b20c445e1828fc7f253c3854907c7ef85a3fe723eda88eb6af5a5908

Observation c7c025be-ad8a-434b-ab31-d3acc0e091cb · outbound

This paper cites Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.760832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.760832Z digest=sha256:df2351a52eb234285150dd82a00696703a25d28738e0216eec16ce475c185c70

Observation 2d1cdad1-23db-49bf-973b-2304ef7d8878 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.765564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.765564Z digest=sha256:6829866434ced753cc265ae9c5f767a3bd8fdfd2823c0f4821928296d23a16dc

Observation f7125f4b-9524-4f2f-bb51-c7472aa3e70c · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.773425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.773425Z digest=sha256:90afa7c01baeda8eec54e07e2b184e37843ba8e1bc61527f3c414552b6d2fffb

Observation 51ede104-63ed-4a65-95df-a7068fb2edba · outbound

This paper cites LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:35.868451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.777182Z digest=sha256:a63a589d35dfb96c994adb756bf988a7d7ef364cd1763815e01843c9c024a012

Observation 7e84904e-ce8f-4cf2-b123-5d300a9c3e9d · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.038890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.781938Z digest=sha256:3cb8f69d3f8a75d577e5d4e53e354e2f7e1caa7aae9fa5235d338d76b4622da0

Observation 7d03a866-8998-4115-942f-ca7cbea3c008 · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.786739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.786739Z digest=sha256:30fd57e89042d13ebe428d99f73971b7fee1f629a70c6bdf61dc4fad427ce00b

Observation 2511efdb-5ff7-41fc-acc9-8bd3f5727caf · outbound

This paper cites LIMA: Less Is More for Alignment.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LIMA: Less Is More for Alignment

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.790232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.790232Z digest=sha256:0bf96bb3335e426edb91658541a0a9c825939ada11967a9b761a2d80e2f0cbe2

Observation 82f303b1-65e0-4869-8233-fc05e3500b2e · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.794271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.794271Z digest=sha256:49c912fe26b41f73885185776ee6f310a820bb16a2e2c737163add4883bd5092

Pith citing papers

Observation 9d9a0935-e55e-49d7-b954-564f6fade08c · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.978682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:a01da152762eb47d131cc49b643bd1f1d3169ae737f61c23cf795bccd0636e03