Pith. sign in

Paper Citation Record · LEDGER

Fake Alignment: Are LLMs Really Aligned Well?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.05915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05915 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:52.256199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T07:37:04.295513Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e31f9482-b5a0-44ba-ab4e-7991955b234c · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Fake Alignment: Are LLMs Really Aligned Well?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.256199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.256199Z digest=sha256:092daeb670d8d81428b05440a0307227ec2618ee1f905307b94b8af8be29c8d5

Observation 623bc94d-6277-4646-93e3-e520db0653df · inbound

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies cites this paper.

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies Fake Alignment: Are LLMs Really Aligned Well?

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:35.411291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:35.411291Z digest=sha256:44e20f735ea760497b65446bfafae276974b33b0d94db21d0dfc3f4990e1b900

Observation a8efab55-a054-434d-8ef5-7b8c02e87b1a · inbound

Why do AI agents communicate in human language? cites this paper.

Why do AI agents communicate in human language? Fake Alignment: Are LLMs Really Aligned Well?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:49.757309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:49.757309Z digest=sha256:17431a407c5737e493381e101beb07a6c5827c4bc88233df0fddbadd007ff463

Observation 8ec9bd09-a478-4f06-8507-756bbe4b9125 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Fake Alignment: Are LLMs Really Aligned Well?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.023624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.023624Z digest=sha256:7ea531611aa289e7462d59b4dfdc31021c24c9ad607ecfef7faa6478fb1db782

Observation f3cd0a44-9e8a-4d43-a9a4-9279b82d2539 · inbound

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints cites this paper.

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints Fake Alignment: Are LLMs Really Aligned Well?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:37:04.297016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:36:10.174861Z digest=sha256:5bbcac46f9df3183319745c44a3ea5078b13b520565dd01c9b73ef396821cae3