Pith. sign in

Paper Citation Record · LEDGER

Fundamental Limitations of Alignment in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2304.11082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.11082 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:58:34.926982Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

44
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc298fbf-da94-4833-8e0d-3f5871613c44 · inbound

Jailbroken: How Does LLM Safety Training Fail? cites this paper.

Jailbroken: How Does LLM Safety Training Fail? Fundamental Limitations of Alignment in Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:17:42.828873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T18:17:42.752997Z digest=sha256:be494f773ff72054f03675035cf7dcd6a5c9c221f74a2905460775b9f11eded3

Observation 3efae1f4-54d7-4687-aafa-7b082e8c0467 · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.470910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:e7fd4af88e6a5d2439af85c449611ea10c7822bb7ceb89946041a344762655ae

Observation f6879a9a-010c-4b07-8abb-41580cb58f24 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Fundamental Limitations of Alignment in Large Language Models

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:34.926982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:34.926982Z digest=sha256:29f51daddf39ceaf170a86967c0720359969f8eecc57bf9cf04f648868fcee5c

Observation 5e43b040-e749-4864-b1e7-84ea0d259d51 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Fundamental Limitations of Alignment in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.496083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.496083Z digest=sha256:9a57b39e1d5d2f6d8a667f86bcde274048793aeb1eacef4d8dac00ec13024f54

Observation 1885a77b-acf8-4ed0-bd2d-6a19a4d708c2 · inbound

JavelinGuard: Low-Cost Transformer Architectures for LLM Security cites this paper.

JavelinGuard: Low-Cost Transformer Architectures for LLM Security Fundamental Limitations of Alignment in Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:25.526186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:25.526186Z digest=sha256:11ff71492a40c3685c86e38622e679f72de227e6b6c2421e76f8e5edaedb40af

Observation 6decc58e-6b91-4b12-af93-4bdc9bc1e436 · inbound

Hyperbolic Deep Learning for Foundation Models: A Survey cites this paper.

Hyperbolic Deep Learning for Foundation Models: A Survey Fundamental Limitations of Alignment in Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:36.557710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:36.557710Z digest=sha256:34ec934b4eb44cd3aff98b899d46e13d5fa81c3142b129656ede4a17a777702a

Observation 560af3c0-046c-40eb-90de-d44ddc6a6834 · inbound

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring cites this paper.

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring Fundamental Limitations of Alignment in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:51:03.796947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:51:03.796947Z digest=sha256:a86b81b578afd8920eb0a8053c04628c2425e4b9ae5dc29a0407c8a35535938a

Observation 5e0dd17a-3fe9-4cf7-97a3-67fc29cff162 · inbound

Robust AI Security and Alignment: A Sisyphean Endeavor? cites this paper.

Robust AI Security and Alignment: A Sisyphean Endeavor? Fundamental Limitations of Alignment in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:58:38.228561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:58:29.306435Z digest=sha256:dff77139b864d202960ec1e67efca1d75d060303048cea2048ed228dd898afb3

Observation 6a98c7df-c782-4da9-8ac3-581311a95014 · inbound

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations cites this paper.

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations Fundamental Limitations of Alignment in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:18.723757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:04:24.077276Z digest=sha256:db1616e7e1882b8065a45a2157550b7e94ce9cc73a1d3655b3b1ef34a847e105

Observation 5d7d489b-4ec9-4eb0-8d06-8dab12957bd0 · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Fundamental Limitations of Alignment in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:43.581200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:71d8dd9d926f02c933bb36c3446b619c7a70bbda11eef072f3abe26e508a02f5

Observation a9fe88b3-3c82-4db0-9281-54d53f7f3910 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Fundamental Limitations of Alignment in Large Language Models

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:55.556966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:d528a37bc28dcd63a24074dbaaeebe1d7bc64118ff274dc549fdee303162f828

Observation 9b99e65a-6ad6-4fdd-bcf4-65cbeffdebe5 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Fundamental Limitations of Alignment in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.130331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:8170bbc47a1e431cbe470041c48218f1bc4d6cd9ed76c831bf87d2c3da85dac1

Observation 3de5c74b-9fdd-42c4-8631-8a55023605de · inbound

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs cites this paper.

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Fundamental Limitations of Alignment in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:49:35.636729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T04:45:35.079192Z digest=sha256:750d088afd5a2f65b79f10f80cb8991abd90a69699d73d23904cbf1cf019e679

Observation e70bd3cf-98ea-48d2-84aa-478bb6a88973 · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Fundamental Limitations of Alignment in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.928075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:28:02.493124Z digest=sha256:a583a114c7d3d4d0ce910015a6f816716a89e1a75d48cd244324f155dc9c8379

Observation 669a96b2-01be-4161-880d-11824b753b97 · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Fundamental Limitations of Alignment in Large Language Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:57.695515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:17:57.695515Z digest=sha256:e4f322c852110af7ddd7bbb1c87fca0f971fdb5efe97d5b944dfb80e56169ccb

Observation 7fab3428-33eb-4396-bbc3-6788746fba32 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Fundamental Limitations of Alignment in Large Language Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.150041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:18baa55d1d92c93e546c2f3f8e25c64e82836cc02b8e0b9112b4682cdf5df97c

Observation 8fc14387-4f84-447c-8574-00d7ebe85e8d · inbound

Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs cites this paper.

Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs Fundamental Limitations of Alignment in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.659788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T18:55:45.465260Z digest=sha256:e05957fe8d923813685cff4d4aafc13db7e50a1df5ca8dd77bac938474069016

Observation def98b02-acfb-4a38-bbab-17645d140d20 · inbound

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy cites this paper.

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy Fundamental Limitations of Alignment in Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.953535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T18:36:44.265273Z digest=sha256:b59f9eaca2320e5a8a5d38336d329d35a4694f62d2ea781dcb70439d73105244

Observation 92bd060d-0b28-4cbd-9348-c6f9896922f7 · inbound

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study cites this paper.

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study Fundamental Limitations of Alignment in Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:18:03.179730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:40:48.736006Z digest=sha256:68958cef433dcf9afdb63ca520f02633c4b8bddb7db658877391b8112748c7c7

Observation d459a3d0-42eb-458d-81b9-10a3b9baff94 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Fundamental Limitations of Alignment in Large Language Models

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:31.376035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:31.376035Z digest=sha256:1637b8a79507cd443a1e376459ef3f31c7fdf7e572220a230d9c0a19948fa2d4