Pith. sign in

Paper Citation Record · LEDGER

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2507.02977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02977 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:22:59.700209Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:48:43.262187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:45.914511Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14b7a6ad-30ac-4e50-a1a3-671cd2f4a38c · outbound

This paper cites Openai o1 system card (sep 2024), 2024.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Openai o1 system card (sep 2024), 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.888753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:58.393895Z digest=sha256:e9813e25b2642c9e457e0cb457164d9c6e86bf2907ee69b8c0aaaf05f6bed249

Observation c839a1b3-e8fc-4d37-bd23-900c89b86489 · outbound

This paper cites Troy, Stuart J.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Troy, Stuart J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.486686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:58.455386Z digest=sha256:c61374cc65cafd2d8cf10ec863d5c4600d38df44cbd270c8da12e4b923f39799

Observation b162460f-f4ce-4b8e-bf00-7117e17111b8 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Frontier Models are Capable of In-context Scheming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.636617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.636617Z digest=sha256:a3e162eeb651f2da80f6944b81eb14b601b95e9a35d4fc9d2dc053c680ed2142

Observation da5ee6ad-0100-470f-893a-0314c97069f4 · outbound

This paper cites Demonstrating specification gaming in reasoning models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Demonstrating specification gaming in reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.773703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.773703Z digest=sha256:6da3b48dfdb9b0eb2ebd1e16f8f84d79d1c78c5db6a58dfb9e5d7757093fbdc3

Observation b9a3c403-56a8-4381-a94f-b4874e0c9b15 · outbound

This paper cites Alignment faking in large language models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Alignment faking in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.816501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.816501Z digest=sha256:23c4f4f37add0fd87299931c52aa406314ce7cd6f8d3485e9581b6850141c812

Observation f84e754d-9041-4a9f-ab13-352b95172087 · outbound

This paper cites I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:58.923788Z digest=sha256:7292f843954028582b06da533ecae9a6b666ae30d4dc9d33f8732402fe99cfc1

Observation caddabb4-294c-48d2-b1e5-78e4f70d6b39 · outbound

This paper cites Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.013316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.013316Z digest=sha256:5e87a7edcb24da550ba17ed2c430d20aaf998aa55ed95536e6669a4902026cb3

Observation 948e0b54-c2a6-436d-bc62-e1fc37608a58 · outbound

This paper cites Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.018086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.018086Z digest=sha256:f227a87c06f9b8595643719bf37e94032631044a02bea51c467bb5fa73d414d4

Observation 0b532cbf-e85a-455d-b071-122038bcb9e8 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.085926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.085926Z digest=sha256:9b73e7db66e8aa4584228e2affa20fcc1fded6a1e3439e0698305369c344aa4b

Observation ac2932f3-7626-4f90-8d04-0933120e2807 · outbound

This paper cites LLM Agents Should Employ Security Principles.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents Should Employ Security Principles

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.187197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.187197Z digest=sha256:88a7148f9d5bbb408931ae2ac2b697895e5c25eee6bbab0d88e0166dbfc99bfb

Observation 864ad225-6ec8-4119-bf39-664983f92734 · outbound

This paper cites Gemini 2.5: Our most intelligent models are getting even better, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Gemini 2.5: Our most intelligent models are getting even better, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.764807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:59.246200Z digest=sha256:349f3cbbd4a94f1d53fb811a6290b23e4ce4a7d39f9066b1c4cb59b9dd8fbdfc

Observation f8715da1-640b-4c2a-b3d8-ff1984219aba · outbound

This paper cites o3 and o4-mini system card, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance o3 and o4-mini system card, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.588657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:59.363069Z digest=sha256:4df47f9acbb472ce0fef0ce625ddc31cf958b5cb373441fb0f3365edde5bfbf5

Observation a2429620-6956-4cc3-a50e-c31eb707128d · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance System card: Claude opus 4 & claude sonnet 4, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.394776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:59.442120Z digest=sha256:e050c37d9d2908c62413834fc84e98e482758e826533210e54ba8579bb3daba8

Observation 36056a68-c910-46cf-8ba1-87983f700e77 · outbound

This paper cites Deepseek-r1-0528 release, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Deepseek-r1-0528 release, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.214313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:59.518923Z digest=sha256:9054fece12ac53171571875109a47cd02310386405fdfaed03740bfdfdc84e38

Observation a75e8976-5b5f-4ce4-9f19-0589630e90d9 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.603290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.603290Z digest=sha256:f71aa6246f4dd274f393cddac22924a14af87947620fc2a181f4d8a9ef26f12c

Observation e0e97d17-b714-42de-897c-4af5f0637f63 · outbound

This paper cites reference.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance reference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.014538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:59.700209Z digest=sha256:d379c3a134ebba1def52157da175e216c88d73a4581695f33d04c8184aef489f

Observation 5497942e-b2b9-4158-b589-003931fee4ad · outbound

This paper cites an unresolved cited work.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:01.201836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:22:58.506406Z digest=sha256:93b413d2a759636f1ad2c4730a8fa7fe6e70ab70d0ee2c4a5f48cf684b18ac1d

Pith citing papers

Observation 5ab2a844-f601-40cf-b5ac-f63a4e8c73a8 · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.427998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:33:40.813795Z digest=sha256:f25a4ee4864a926f153525081c365fc8b7883edc124b72f68eeb478fc77f9682

Observation c237a8e7-8289-4fec-8ef3-ab1df0f3d69d · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.916404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:48:43.262187Z digest=sha256:0f3525257297bbf40f78d7496ef61263a11d6c281b2b93e33c45842536f43cf6