Pith. sign in

Paper Citation Record · LEDGER

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

As of 21 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2507.02977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02977 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:22:59.700209Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:48:43.262187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:45.914511Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14b7a6ad-30ac-4e50-a1a3-671cd2f4a38c · outbound

This paper cites Openai o1 system card (sep 2024), 2024.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Openai o1 system card (sep 2024), 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.888753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:58.393895Z digest=sha256:6686065f1648e9dfd99b1eed6e208225563b0d119528cb36d3ff6e606cbd652d

Observation c839a1b3-e8fc-4d37-bd23-900c89b86489 · outbound

This paper cites Troy, Stuart J.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Troy, Stuart J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.486686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:58.455386Z digest=sha256:bafe0784c401eba12b846ccf9c45067ef0e575ad87f72d3848d5d6216f77db51

Observation b162460f-f4ce-4b8e-bf00-7117e17111b8 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Frontier Models are Capable of In-context Scheming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.636617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.636617Z digest=sha256:4cab7efd402addc2bbec199fb71ab8e6aa73e12e636eb8d87a596cc032a9b8b6

Observation da5ee6ad-0100-470f-893a-0314c97069f4 · outbound

This paper cites Demonstrating specification gaming in reasoning models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Demonstrating specification gaming in reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.773703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.773703Z digest=sha256:f61143692952609aa4b3225ede2e38fd100be2ad92d8d541181acd3be19f56d6

Observation b9a3c403-56a8-4381-a94f-b4874e0c9b15 · outbound

This paper cites Alignment faking in large language models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Alignment faking in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.816501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.816501Z digest=sha256:9c00838172cce8240a27c5d859e303384e4c1fcd386801af4e2293df770f00e4

Observation f84e754d-9041-4a9f-ab13-352b95172087 · outbound

This paper cites I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:58.923788Z digest=sha256:73215dc18fd210e9a875c6274532f44b20e6600e900b0f2b268117bf6be2d943

Observation caddabb4-294c-48d2-b1e5-78e4f70d6b39 · outbound

This paper cites Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.013316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.013316Z digest=sha256:c2764dd4bf620ba9b727efab70bcf404b76d54d008ff1a65b2c86dd8f2461a79

Observation 948e0b54-c2a6-436d-bc62-e1fc37608a58 · outbound

This paper cites Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.018086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.018086Z digest=sha256:387f2ef86dbcc33f1d0cbb201ea08778e2c14fa6d351f89570d9c84ff7e4c205

Observation 0b532cbf-e85a-455d-b071-122038bcb9e8 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.085926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.085926Z digest=sha256:6ea4ec52296e095bdf7effa3346a4acff3d5dfd20d765def68ad98314cf6aa42

Observation ac2932f3-7626-4f90-8d04-0933120e2807 · outbound

This paper cites LLM Agents Should Employ Security Principles.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents Should Employ Security Principles

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.187197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.187197Z digest=sha256:54aed7aceb42f90ac504c98d28207f08f9a0bacc7fba98923335321919b32068

Observation 864ad225-6ec8-4119-bf39-664983f92734 · outbound

This paper cites Gemini 2.5: Our most intelligent models are getting even better, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Gemini 2.5: Our most intelligent models are getting even better, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.764807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:59.246200Z digest=sha256:205dfe6906a288ab376b1ec393e2a5a809a253222ee36d64c0c24779b51b4e76

Observation f8715da1-640b-4c2a-b3d8-ff1984219aba · outbound

This paper cites o3 and o4-mini system card, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance o3 and o4-mini system card, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.588657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:59.363069Z digest=sha256:2b099bcfff0bf716696ab5dae5b9e00cb25d4ee034edaf5bff17ada282f3ba1f

Observation a2429620-6956-4cc3-a50e-c31eb707128d · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance System card: Claude opus 4 & claude sonnet 4, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.394776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:59.442120Z digest=sha256:56f1e91772145811516d9443bdc0f8404f81cda22f0dc2d7b903fd182d4db9dc

Observation 36056a68-c910-46cf-8ba1-87983f700e77 · outbound

This paper cites Deepseek-r1-0528 release, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Deepseek-r1-0528 release, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.214313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:59.518923Z digest=sha256:62d1018ba5abb494471af241e2f9c0701e03f33aebcb1a5f7036cf1dbb6dfd33

Observation a75e8976-5b5f-4ce4-9f19-0589630e90d9 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.603290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.603290Z digest=sha256:e4e8e544f0a62d0a11bbfd36302ac2cef9dbf954bc96f463daa88ea3d39f4c8c

Observation e0e97d17-b714-42de-897c-4af5f0637f63 · outbound

This paper cites reference.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance reference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.014538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:59.700209Z digest=sha256:0d1a12c5f3581a75e9eeafb1877afc6ef1b7e58a7b5dcaef3766e198c0e3536f

Observation 5497942e-b2b9-4158-b589-003931fee4ad · outbound

This paper cites an unresolved cited work.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:01.201836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:22:58.506406Z digest=sha256:f2e4c888b118ddea45cec5b456639907baea887a9936b2ae410cbb01ce8f4190

Pith citing papers

Observation 5ab2a844-f601-40cf-b5ac-f63a4e8c73a8 · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.427998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:33:40.813795Z digest=sha256:41a86733220ff959691736e634c2952498588c3f74cde5671749900912045a38

Observation c237a8e7-8289-4fec-8ef3-ab1df0f3d69d · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.916404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:48:43.262187Z digest=sha256:d86e54e5df2759942e2350dcc0dca70390c2b8f0e82a9d918069d07f1f526583