Pith. sign in

Paper Citation Record · LEDGER

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

As of 21 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.08212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08212 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:22:49.865192Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30f2de63-4956-4b3a-8e2e-ddeb207cfd9c · outbound

This paper cites Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.530177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.530177Z digest=sha256:5520700727fcd8f5404fb4d588cc5c83b0626bea6ae5f802fcfe1e1e3d86ed39

Observation b3d0543b-02d0-4d6f-ab46-3fcea3315427 · outbound

This paper cites Many-shot in-context learning.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Many-shot in-context learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.059538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.536944Z digest=sha256:74eeed57aa625ef4f7b6b990b5a5ac2581941f15ca0fd99cb49dc62b47d947c3

Observation 34b672a5-4256-405a-87d6-ffe7704fd03d · outbound

This paper cites Bowman, Ethan Perez, Roger Grosse, and David Duvenaud.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Bowman, Ethan Perez, Roger Grosse, and David Duvenaud

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.541983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.541983Z digest=sha256:fbf72f9536b25593b7cb586dea8bacf2ded378c927322ce594cfa0a82fc6110b

Observation 9e017570-f999-433d-9489-8145da40862b · outbound

This paper cites Refusal in language models is mediated by a single direction.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Refusal in language models is mediated by a single direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.547731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.547731Z digest=sha256:92621a53de7ba2278647ce8558476503d413acf4d3a1adfccb7ab7045b0b579c

Observation aae8d46c-7a70-48f2-9e1b-da1be140e47c · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.552757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.552757Z digest=sha256:46f68ac4a02eb2714b1153c1c05097028cfb135a3392614e7fec5a13c390fe57

Observation 639a3367-cad2-4e19-9d0b-8c672c1cfc80 · outbound

This paper cites Tell me about yourself: Llms are aware of their learned behaviors.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Tell me about yourself: Llms are aware of their learned behaviors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.023208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.558154Z digest=sha256:5a7c17b5b66a0d931791794a214123d10b62a12052de505703f1ecce34e4d82b

Observation 7eec4d22-4c79-4d89-8749-33db422aa9f1 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLM s.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent misalignment: Narrow finetuning can produce broadly misaligned LLM s

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.007811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.563493Z digest=sha256:cfd5726880ee5af3edd40a352fd6f9982be3ba8477a8769cbbade3b8a42f20f0

Observation 6987c8c9-ac37-4a9f-81f8-310a38e9980e · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.568072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.568072Z digest=sha256:b876553849de0a1bb060bf330d71038e7fad30042adb61d289bc31c9c1ed07a3

Observation 621d6d3f-6678-4dac-ba54-539ce9a0e19f · outbound

This paper cites \ StruQ \ : Defending against prompt injection with structured queries.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment \ StruQ \ : Defending against prompt injection with structured queries

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.991287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.573255Z digest=sha256:46b8ca10a332abc660fbe79e23196dc1d07e0d647f1a08e3d4cd53ed8ccc6b76

Observation 555278eb-af8b-4711-8b0d-782113f92ac2 · outbound

This paper cites Secalign: Defending against prompt injection with preference optimization.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Secalign: Defending against prompt injection with preference optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.973490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.578105Z digest=sha256:a406917a8cd4ddea0efff35eb98618c7102d9ef68d2e5cca0b02fd4f8baa78c6

Observation 287aaccf-420c-43a8-ac28-a9fd9141759d · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:22:50.956682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.582662Z digest=sha256:e9f37aa407cd88c3754bcf82659c74ccacc7997db16ee77773b3fd06751f041d

Observation e6d33028-ef89-4f53-9770-0574c9340885 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.587294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.587294Z digest=sha256:c418f66434dac252a4ed7cd7fac4eedfa9955bfa690edcee4cad31a0f56003db

Observation 2a708400-657c-4e5d-8fd3-9a4a967e2a09 · outbound

This paper cites Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.592032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.592032Z digest=sha256:31b915b4d086a25b644005496881b5eb42376ded06c8ef6d24579a959443b09f

Observation c34faec4-3895-4a1b-915d-106b7bb08079 · outbound

This paper cites Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.941303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.597204Z digest=sha256:181ad6c5b1c2e11a1b3403a2c026b5e6f8abaa9ca0e8ead1bd3e70128f75344b

Observation 0b928a21-9e6c-4d1d-ab14-36878b3288e1 · outbound

This paper cites The Benchmark Lottery.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Benchmark Lottery

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.601913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.601913Z digest=sha256:9dfe160b03116a9fc90a8d9b35a7f5c97677f4be12c5471eda61e8c0f8906aa6

Observation f2dee1c2-19d9-41ec-9e01-97332837ef01 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.606931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.606931Z digest=sha256:847b1adf0cd888f86834e3fa6db2cf9b6700fa6108184a8675fba04dacc1b871

Observation fe8b9d1e-32ad-40d5-b753-bca601da0d05 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.612207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.612207Z digest=sha256:933873ccbc3a28f4fcadf2fb688d810b125b1c212f5fdde85a2a6aafd4127369

Observation 271749fd-1904-46b7-98f6-0486be589366 · outbound

This paper cites The hitchhiker ' s guide to testing statistical significance in natural language processing.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The hitchhiker ' s guide to testing statistical significance in natural language processing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.617371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.617371Z digest=sha256:39c55f4fe0c79ba8a14c7487e46dc2182827a8cce9c1bc7cf0f28905d8879e29

Observation 4cafb2b8-5361-417e-893a-f445a07474bd · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.622390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.622390Z digest=sha256:a2bf84acb958644129aacf9ec9c16f15c93f0187e44a6d0142216e83c1d87508

Observation 011eecbf-38bf-4d25-b9a0-115bb38a7fa3 · outbound

This paper cites Wasp: Benchmarking web agent security against prompt injection attacks.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Wasp: Benchmarking web agent security against prompt injection attacks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.925148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.627513Z digest=sha256:bf55bf72fcbe75f2161054df61e31a16f6a5ab03dc4ee77ec419461d9826bba1

Observation 7ed41a84-9f6d-476d-b3c1-03ad84828a3a · outbound

This paper cites Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.632374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.632374Z digest=sha256:97b1210a3ff17969b7559402edb204af8532d389a820250bd14d62b210b2a831

Observation f3b8397d-f8b9-4ca6-8209-559019f8cdd5 · outbound

This paper cites Position: Anthropomorphic misalignment research needs stronger evidence.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Position: Anthropomorphic misalignment research needs stronger evidence

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.908290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.637059Z digest=sha256:80790dc4d777c5ea6d5797e08719b5db73afde62f419878b9a4a6b5239525e34

Observation a7ddc839-af7e-409e-a758-618ac23869c1 · outbound

This paper cites Defending Against Indirect Prompt Injection Attacks With Spotlighting.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Defending Against Indirect Prompt Injection Attacks With Spotlighting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.641681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.641681Z digest=sha256:e6815f8ca130a7fad1ca68f543aedef18fbe7cf1601a1fb21734f2c17f4f1ba7

Observation 707eb619-952f-477e-9c6a-ad3ac84840bb · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.646499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.646499Z digest=sha256:fd085a95df3c9dc5dedf7d22c7f358209a4a2cd715611d4aec5f5a09273c0bf8

Observation 3af5cd8e-7a50-4e32-8e75-375ad1ea6fba · outbound

This paper cites Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.651555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.651555Z digest=sha256:c3b300f5a1168fadabd2551712a956612f79b92e90526e47732661d4ff086389

Observation 943297e2-9f8b-4204-809a-eb77597ee23e · outbound

This paper cites Holistic Evaluation of Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.656331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.656331Z digest=sha256:188f4224ad1b133951d483facf127ebc56ca026802ac3c46f78ea061d65ebe58

Observation a957594b-48bf-4dee-a815-a3bdbd0f359b · outbound

This paper cites T ruthful QA : Measuring how models mimic human falsehoods.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment T ruthful QA : Measuring how models mimic human falsehoods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.661445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.661445Z digest=sha256:f629d037aba300d443ff8cdb642d09b83d7dc980ce9a80eb767fc0aee5933286

Observation 6cc4a7ec-b8ed-41f7-b751-0b973dd54bc7 · outbound

This paper cites Troubling Trends in Machine Learning Scholarship.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Troubling Trends in Machine Learning Scholarship

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.666283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.666283Z digest=sha256:4453aef3d759a9e0ee55c03b3adaa260d85572d7ae45d1317826fe53119a063f

Observation b0bfcf0f-6a85-4219-a202-965e0d18b19c · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Formalizing and benchmarking prompt injection attacks and defenses

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.671474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.671474Z digest=sha256:2ae69299a5394a10ed126c469c4d641ad04cb324b21690bdb1f2cbb31674d4db

Observation 6b497d54-f1bd-41c8-a238-7675a2f36cde · outbound

This paper cites Datasentinel: A game-theoretic detection of prompt injection attacks.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Datasentinel: A game-theoretic detection of prompt injection attacks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.882879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.676173Z digest=sha256:6130df3d5299f6d73d20ed3f55ca3f1afe45a2a3bf86237f39b78c45a802e7cc

Observation 44779ca1-8a59-498b-9a96-739c8251ba73 · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.681164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.681164Z digest=sha256:a09912c06fcdf33c8ca7ab9110a7d0ca7057e6e30522a37cd94557be83c27c61

Observation 6097cff1-42b8-4d28-88d9-1490e17c7d2d · outbound

This paper cites Harmbench: a standardized evaluation framework for automated red teaming and robust refusal.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Harmbench: a standardized evaluation framework for automated red teaming and robust refusal

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.867771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.686546Z digest=sha256:9e0560cc73b4d126384c5134fc0268d24299b392f5374a92ae2be92fedaa5cae

Observation 06198c22-fefc-4138-8fc5-8c0afef9c2a1 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.691804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.691804Z digest=sha256:5a744b429c2684a1b1f011ebb82ac4565ca484b982243e38b1ac6254cffc35d9

Observation f8ca78fb-49b4-49c5-82ed-3ef9b3d422ba · outbound

This paper cites In-context Learning and Induction Heads.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context Learning and Induction Heads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.696753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.696753Z digest=sha256:414865397236c5444cef8062dbf93c7326519640a852254c8b876e0ed63e66ae

Observation 7b11e404-96c9-4ba4-903d-ad3ed9c08048 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.701739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.701739Z digest=sha256:ee0e09bcda5f7e26e9334f319441d99881a734bdf7a115f9343e7d2c27836c6c

Observation e62d3fa4-7db8-4f62-86f7-a53fcf23a849 · outbound

This paper cites Red teaming language models with language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Red teaming language models with language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.706357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.706357Z digest=sha256:711f4b97ff7da4daa657782d35be591bb43233bb624da868466447ecc8e784d9

Observation 8af15e0e-5c6a-4abe-9cd6-e04a16c2ecd4 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:22:50.852490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.711447Z digest=sha256:e3893430210f7831eb1df4c54698937ebdb3e684d5503b9d6a5f633c98995aac

Observation 6c9fde20-01e1-4d2f-ade2-8b20f8760903 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Safety alignment should be made more than just a few tokens deep

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.836554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.716887Z digest=sha256:4d942007b477124e7f96de394326aa3bae7b62f0ce68a53cb4e8f4376c312ae0

Observation 1d6c92fa-476f-4f9b-a90e-df61fb8bfc51 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering llama 2 via contrastive activation addition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.721543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.721543Z digest=sha256:7f408cb86339f63f002256448f47239f23d92e7b3d5e5be2718e164fa8f1cb80

Observation d797a939-59dc-4ca1-8447-11448dabf0ef · outbound

This paper cites In-context impersonation reveals large language models' strengths and biases.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context impersonation reveals large language models' strengths and biases

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.819833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.726219Z digest=sha256:d79bf247a76b50666dab991790e0ab533969ff37244a51ffb5beeb2253f4c659

Observation 08a2ffe9-6d5a-4180-add7-e0ecdb619067 · outbound

This paper cites Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.802724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.731100Z digest=sha256:cfd0683780643cbc997f25adc6b1888f6086955818a6ad4752619c22fda11aa4

Observation a32a04f3-3381-4b68-9e3d-45ef6dab5ac0 · outbound

This paper cites Role-Play with Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Role-Play with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.735823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.735823Z digest=sha256:8b5efeef5a1dc6c0ecc2ea57f86b18c9d655a8cd4529ff52d45a5679256a21e8

Observation f48c357c-f6d7-4264-9f4b-5f2b0b83337d · outbound

This paper cites Judging the judges: A systematic study of position bias in llm-as-a-judge.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Judging the judges: A systematic study of position bias in llm-as-a-judge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.784805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.741322Z digest=sha256:213ddf9afb06d4ab4b22bdd7b3364094c5d9d70663abf3bc19a0f63ce397b798

Observation e00827f3-94be-42f4-bccf-f8692ea9bdba · outbound

This paper cites Convergent Linear Representations of Emergent Misalignment.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Convergent Linear Representations of Emergent Misalignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.747675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.747675Z digest=sha256:2a2b3fb91d01100923ac97a43645847b2d845701abf91c46b7489a03c8bf4d90

Observation 094bf606-34ba-4f86-8738-0cc34bc864e2 · outbound

This paper cites Extracting latent steering vectors from pretrained language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Extracting latent steering vectors from pretrained language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.752615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.752615Z digest=sha256:de7934f00ce7ad6446de3c7c65b6109ec729c6994f152cedd57e8d5041d61cb6

Observation 66c2b156-72be-4896-8993-2f3a708ed2b0 · outbound

This paper cites Function vectors in large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Function vectors in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.768991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.757488Z digest=sha256:ae0d6fa41132f48140aa0d73cf121f57e4d5c88dda4ac5f28f129feabd456db0

Observation 1c5d85c5-f224-4b3b-a920-24e93a739553 · outbound

This paper cites Steering Language Models With Activation Engineering.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering Language Models With Activation Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.761954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.761954Z digest=sha256:cba8da7226cd063dcb7bf82d1e53ef94bf965cb321661c676bf1234e1273e59a

Observation 0cf35d45-3bd8-4b67-9c32-630c9aa32987 · outbound

This paper cites Model Organisms for Emergent Misalignment.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Model Organisms for Emergent Misalignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.766759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.766759Z digest=sha256:fdf676ff4873b8070030742294496ddba7de41dabe0f369f6ab7ad04555005aa

Observation 46490b72-1a8e-4619-94d6-c86174bfc50e · outbound

This paper cites Transformers learn in-context by gradient descent.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Transformers learn in-context by gradient descent

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.753509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.771799Z digest=sha256:a34a3b5ef66442b014d7314c998ddfc04e7ed4bb253d51344dc6db7308856811

Observation ec7a35d9-4bfe-4929-b435-e965a9bdbed7 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.776344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.776344Z digest=sha256:73a2a4a31ec9ddfbec9113735c57e0605864fbe00989293154c213676206de57

Observation 4b094c50-97c1-44ba-9ddd-1ed45fe2be86 · outbound

This paper cites Label words are anchors: An information flow perspective for understanding in-context learning.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Label words are anchors: An information flow perspective for understanding in-context learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.781324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.781324Z digest=sha256:83518836673a63eb5f0713933deb63c3a014616271239dfdd235b03041908a7f

Observation eacf70e9-5587-4eab-b86e-fe3c99366590 · outbound

This paper cites Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.786015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.786015Z digest=sha256:3d01b85559d42b39383790a826c226afd83110aba02382f97e6977590e6197bc

Observation 04ec2cc2-764e-4977-81da-a5b56ba5ba9c · outbound

This paper cites Large language models are not fair evaluators.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Large language models are not fair evaluators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.737327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.790803Z digest=sha256:8ba4e864da2b384b5f465a9d2b7c3d7b8faf85b0db9de23617138e50c36ec647

Observation bde5495e-709a-4086-a340-311f6d4c8a40 · outbound

This paper cites Evaluating general-purpose ai with psychometrics.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Evaluating general-purpose ai with psychometrics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.795389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.795389Z digest=sha256:a7006727797d39557e57f95017824eda971015f726e0aad94f81f1e299d3bfdf

Observation 9b8fbc86-2e1b-4f95-958e-c74bc8a148af · outbound

This paper cites Do-not-answer: Evaluating safeguards in LLM s.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Do-not-answer: Evaluating safeguards in LLM s

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.801300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.801300Z digest=sha256:e4663990e299948398a24593518dd5e3f7aad29f4e7b02ff225259426c464e92

Observation 776dc142-fc57-413a-9d1c-f245d7b12027 · outbound

This paper cites Larger language models do in-context learning differently.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Larger language models do in-context learning differently

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.806663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.806663Z digest=sha256:a9585bdf22a918702c091e00b0de53236bdc038ee45b1ae0ec588f12353172f6

Observation fabcc8f9-20fb-44f8-bf08-f210d668cb7c · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.812498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.812498Z digest=sha256:4c73eceb0ce988d91a5084990936c5fc97bc05fe9906ac3dd7aa8fe8d14f62cb

Observation 062e3273-9fc6-479b-b782-78e195b40aba · outbound

This paper cites Benchmarking and defending against indirect prompt injection attacks on large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Benchmarking and defending against indirect prompt injection attacks on large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.817510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.817510Z digest=sha256:d6c7ea0dcf713b26c39d6c9b261bb8d667024f1d260c4b00b577cd382af51aa0

Observation 81dc7557-ccf9-4bdd-a4ff-f3a5204d5ccf · outbound

This paper cites I njec A gent: Benchmarking indirect prompt injections in tool-integrated large language model agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment I njec A gent: Benchmarking indirect prompt injections in tool-integrated large language model agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.822368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.822368Z digest=sha256:3b61ddaba478337954648155d11c815ab953a9bf02be000b3ec1ac4cd89261d3

Observation 06574918-5741-401b-9df0-05fae534f09f · outbound

This paper cites Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.828046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.828046Z digest=sha256:517c298c41b038e8397baed247d6169b733246adeb4987b8e9956b98c048214e

Observation 1d1f2b9f-31a1-47ff-a9c4-0d1878712b65 · outbound

This paper cites Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.721180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.834320Z digest=sha256:3efde870a8796624cc3c221a0b4d5fc631491a31d4ff7d578e80474569e1432f

Observation a3b41a47-a730-45b7-a7cf-40bb8f09ad11 · outbound

This paper cites S afety B ench: Evaluating the safety of large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment S afety B ench: Evaluating the safety of large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.839228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.839228Z digest=sha256:5145d002bece3f81532cc29208cf594418447ff79d1020ecb1ba41e7c087151b

Observation 1462a989-03ce-4f85-97d9-9d494000fb1b · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Xing, Hao Zhang, Joseph E

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.844506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.844506Z digest=sha256:5f1456d2df10de927b8218be6a72b741bbcca21babd351d74f70c74e90f5c0c7

Observation 5c8999f8-653f-426f-80c7-79ea91de4597 · outbound

This paper cites Poisoning retrieval corpora by injecting adversarial passages.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisoning retrieval corpora by injecting adversarial passages

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.849554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.849554Z digest=sha256:00b6e05492b6fe730c8c4d951f0fac79940c91cbd9a911b958c45ab551dae34d

Observation 82694d33-960d-4e7a-9b63-4501ef836839 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.854619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.854619Z digest=sha256:7d19e1d1055da8675f3695ebbf973b83de009bb29da261bc16c98d266b21c105

Observation b92a93f7-de49-400c-bc3c-92b6e4798c5b · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.859617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.859617Z digest=sha256:373ce2c58c28fe900a448a0f9f09b874e4b4552127e146ab69d76f4ee9409e91

Observation 8a726419-49b6-41ea-9824-9c1fbdc9142f · outbound

This paper cites Poisonedrag: knowledge corruption attacks to retrieval-augmented generation of large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisonedrag: knowledge corruption attacks to retrieval-augmented generation of large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.694801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.865192Z digest=sha256:2bb74554d749c0b065d2245fb94b8715042a0c85207d3b453a2b172a07d22692

Pith citing papers

No inbound Pith citation observations are available.