Pith. sign in

Paper Citation Record · LEDGER

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 64 inbound Pith citation observations for arXiv:2404.09932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09932 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:26:42.160078Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e47d2fa6-25b8-4557-8bcc-89e4f3513b11 · inbound

Scaling and renormalization in high-dimensional regression cites this paper.

Scaling and renormalization in high-dimensional regression Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T01:55:55.089386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-24T01:54:48.781227Z digest=sha256:5adcc012e7934f47fa743c85b5a51c4445e49a86f0ad6d527f18f6a1cc74a7a1

Observation 07a1da5b-b4f0-44e8-b987-1a99bf9035de · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T16:25:14.801617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:e4e2f3ffe69588fdab6cfd84189bd1555aa18a76b72be80e2c6a76b374438639

Observation abb05878-58e8-4ea9-935f-d3ea7f231ea4 · inbound

Can sparse autoencoders be used to decompose and interpret steering vectors? cites this paper.

Can sparse autoencoders be used to decompose and interpret steering vectors? Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.160078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.160078Z digest=sha256:da5c0436d8b80a633ceb840c3dd916a96eee08807e5042f93df0f158a7d5e62f

Observation 499b9844-f69f-4404-bd0a-06424cd433fd · inbound

DROJ: A Prompt-Driven Attack against Large Language Models cites this paper.

DROJ: A Prompt-Driven Attack against Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.257404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.257404Z digest=sha256:d017abd9a31f480284f2461c78516014dd5150925daf6ece112835b24135fb09

Observation 1ffa0eb6-c7c4-4deb-b108-8acbd635c507 · inbound

Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations cites this paper.

Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T19:41:19.617548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:41:19.617548Z digest=sha256:4bd24289c155567d78cd59f7e90e0e1a81f642672d86cc8d00ff3e24b83638f6

Observation 8bdcc17c-1756-49fd-b3e0-98dcec4f223b · inbound

Predicting Emergent Capabilities by Finetuning cites this paper.

Predicting Emergent Capabilities by Finetuning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:45.982526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:41:45.982526Z digest=sha256:df7a3cb5233a73acc505319b51ee0e033748e52ee7500d11e6ab2a3c564fc25a

Observation 4c84a848-ebed-4799-a5fc-50bcde6cd406 · inbound

Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need cites this paper.

Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:48.425423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:48.425423Z digest=sha256:460fac402d768c16dde2020dc276170b8944884192fc3c36d2a714b41525d99f

Observation c03342b6-9e6c-46da-ad0b-1fcb20765b59 · inbound

Neural Scaling Laws Rooted in the Data Distribution cites this paper.

Neural Scaling Laws Rooted in the Data Distribution Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:30:15.413126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:30:15.413126Z digest=sha256:702964a228f1e55805c5ee40c977a00d32aca00f8264caa6cf570ad74bae4cf0

Observation 97645858-2052-4b3a-9109-d47e732b2347 · inbound

Obfuscated Activations Bypass LLM Latent-Space Defenses cites this paper.

Obfuscated Activations Bypass LLM Latent-Space Defenses Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.424767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.424767Z digest=sha256:d61615befdcd071f4962ad2a4e2dc82a28a10dbe7ee0e0ebed4fd043f5de4c5f

Observation 9390ee86-a5d9-4451-bedf-148ca9ffb44d · inbound

Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates cites this paper.

Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:19.011861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:10:19.011861Z digest=sha256:b85148822750b6aab06573f63e96e478b15a56c50422bc6a57a45dc8c3ea041b

Observation 94a84c6a-a27e-4769-a43b-42f026ef23e9 · inbound

Towards Responsible Governing AI Proliferation cites this paper.

Towards Responsible Governing AI Proliferation Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:48.277131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:48.277131Z digest=sha256:c1f7d87ff11af806e82367354344b7c5898fa7ecdd307577a33bb65bf4460bc0

Observation e5d1ffb1-306f-4bb6-a7dd-2e27061568e2 · inbound

Social Science Is Necessary for Operationalizing Socially Responsible Foundation Models cites this paper.

Social Science Is Necessary for Operationalizing Socially Responsible Foundation Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:42:47.257659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:42:47.257659Z digest=sha256:d5fa4648612e8036b01783818df8c4acacf324ef18c6f86add023fa0a736d6e8

Observation 9c9485fd-3d9b-48a2-a535-f7d52dd85d64 · inbound

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense cites this paper.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.428520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.428520Z digest=sha256:b2f37b4944e913615e8b6433fd7b94f4f68e844e69d5de725a242b9ed8da838d

Observation b11812df-a75e-499f-a9e1-35055655b0f2 · inbound

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs cites this paper.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.875258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.875258Z digest=sha256:d0b5fdc635c596d2a213fb11ca9480a7026ddc175ff994f34c445bfd7b4edd92

Observation edaf3004-9b57-4970-a514-98992a9076e3 · inbound

Mechanistic understanding and validation of large AI models with SemanticLens cites this paper.

Mechanistic understanding and validation of large AI models with SemanticLens Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:11.730920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:11.730920Z digest=sha256:2bc49fcb4f3bcbe1cf546344d715bc5193640683347c00a6a951cc7b65a66eea

Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.714194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.714194Z digest=sha256:4a278c421a251ae25086098709800f989d2f0f8ab167b1814a81faf29c444910

Observation 4a0ec520-0d0d-45c1-97ad-35ebf5e33ec5 · inbound

Clone-Robust AI Alignment cites this paper.

Clone-Robust AI Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:11.764071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:11.764071Z digest=sha256:7db9d6e5309ff56069e84b97c5dcb591b680d40fa43c26dcad72dc2a4c9343b7

Observation e88b23e9-baaa-4345-a6a3-3101215e6698 · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.488108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.488108Z digest=sha256:fb7b91b40404f987cb49821f9e39a414bcd64745ed18891d39c438dd5bf1826a

Observation 9a148bf4-fa45-45ed-9888-668ed4c51d37 · inbound

Episodic memory in AI agents poses risks that should be studied and mitigated cites this paper.

Episodic memory in AI agents poses risks that should be studied and mitigated Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T17:58:17.922278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:58:17.922278Z digest=sha256:6444025d7c01ea552061763d92c4c2a9403265e8057a1f175ab7ae5092fd35e1

Observation 926cd96f-cc6f-4f4b-ae9a-489852782b6f · inbound

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models cites this paper.

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T14:51:58.699740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:51:58.699740Z digest=sha256:a5d529af201a15b166e9ab6b38dcdbcf72277291d1423d767f8e9a5411cec465

Observation 2e708e35-c5da-4e6d-831e-ba710e2645fa · inbound

The AI Agent Index cites this paper.

The AI Agent Index Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:34.657585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:34.657585Z digest=sha256:7a4e90f29a6d7944a67e5304a02ecf8a720320251446e72043315a1b9ad92926

Observation 0d821892-22fe-43fa-9cf9-a296911e69e9 · inbound

MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf cites this paper.

MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:09:16.102465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:09:16.102465Z digest=sha256:4a9e0c9eaf5827255604342271953a2fc81eecb3b39ab2f674dba02df884aea1

Observation 67090237-d39b-4ec6-85bb-953388fac2cc · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.511215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.511215Z digest=sha256:855128415e4bd14798a90bedf8ae413cb6eea226c4cf2d6d84b70d580b5d4284

Observation 276b0fa3-77a0-4277-a525-fee2aa77b9ef · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:20.088578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:20.088578Z digest=sha256:42cb40051a7a756f1da928025cbdb938ede31e8dda49b8de7bd06c9a85975ad9

Observation 0c2bc9d3-9b3a-4bd8-a2df-7f37c0a94918 · inbound

Mitigating Deceptive Alignment via Self-Monitoring cites this paper.

Mitigating Deceptive Alignment via Self-Monitoring Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.428012Z digest=sha256:08bd705f22292bf5c8db75cf04ca0fdd1f0943b6c75ee1edcc0d33ea44b6546f

Observation eaaae2ea-aa11-4887-9fc5-b7df25e0b9b0 · inbound

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment cites this paper.

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:10:51.554189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T01:06:19.756032Z digest=sha256:13ac7c4dc203ed502230797f142107e3e740bd39f831932db2075301fcaae372

Observation 6604e4c7-4876-4fbf-a6c5-4753b52ba655 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:30.354244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:30.354244Z digest=sha256:3283a1b827fe760b886d3533d4483679c9a7aef2bc9689acbc310e9bcefc9e31

Observation 0d1cbc67-e62a-4ab9-8b73-d7ef4d75dfb6 · inbound

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It cites this paper.

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:13.874731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:13.874731Z digest=sha256:12500d2ddb792e239f576283fe00ae972d170b0234d95f605372fe1d83632ba3

Observation 75c56f71-475a-42c5-8fca-4ab6831a08d1 · inbound

Risks of AI-driven product development and strategies for their mitigation cites this paper.

Risks of AI-driven product development and strategies for their mitigation Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:58.104897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:58.104897Z digest=sha256:ba7269f84e1c69d1cb1bb8327b5138cb66d6c1262105daa86192ac3cc679d709

Observation 6ffa341c-45b3-4e5d-9adc-c2fa387f1b01 · inbound

Linear Spatial World Models Emerge in Large Language Models cites this paper.

Linear Spatial World Models Emerge in Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:28.039432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:28.039432Z digest=sha256:affc8d54ce40d24628a888222cbb3f2e14588b2abd16c0c7d687c317c706296c

Observation e14fbc27-e4c3-4582-9af6-2e467cb1714f · inbound

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents cites this paper.

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:57.913128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:57.913128Z digest=sha256:aca03a0ce3a4a2bed836a44aa67b32397371eb1af0c7ba7ddd0b86d46ca69f65

Observation e62633f9-e83d-4a72-b74f-41265a3c4fd3 · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.451035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.451035Z digest=sha256:95a61b9596816b581f074bef1fa895ef4afc323b0cd5ed7a180c6817134591bb

Observation 28c35c6b-b441-4b2e-a413-138b5172816b · inbound

Probing the Robustness of Large Language Models Safety to Latent Perturbations cites this paper.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.094114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.094114Z digest=sha256:62cef1609d1eb76ea499607503dea24adc9ccb259f7e5fa90dc250e76173ff93

Observation 555ad9b2-e220-459a-b447-9abe1b5f31e0 · inbound

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models cites this paper.

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:22.466382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:22.466382Z digest=sha256:80c3f59b50279a39181efdd91251ce24cf83038746ca39374012534dcb975f67

Observation d718aa0a-df81-4905-89ce-89b731533a5d · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.377039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.377039Z digest=sha256:6a0e963dab60198e396f01cbb06653b247ab4e1d2f38a5fb231ea5221e235ca4

Observation 3e01fb38-ff2f-4f15-9450-db40c886c0bf · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 292

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.423464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:56e9202c619a6508505c70fb21a06c4686311360f0183abe491489245f168e28

Observation 2e208163-7343-40e9-b35a-595ddf22ab0f · inbound

Against racing to AGI: Cooperation, deterrence, and catastrophic risks cites this paper.

Against racing to AGI: Cooperation, deterrence, and catastrophic risks Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:22:27.592346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:22:27.592346Z digest=sha256:ed607e65466450f1db79b214bf3ed3ce953742e2cec04fecb8ad44fc1f99d6c9

Observation 372bda71-ed0d-4395-82e4-c3299076289b · inbound

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions cites this paper.

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:24.795223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:24.795223Z digest=sha256:99ba0ff0416c0fcdd5a18e1dab02024ad33dd666659c7928a834002de32edf44

Observation e0f6b72d-cb2b-44a1-a86b-39d5e21a3c38 · inbound

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information cites this paper.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.902944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.902944Z digest=sha256:8c1ec0b505b76a455df47cbcc3479ebb4264f9a08a1084f09d2ed6b230cc527e

Observation ecc3fec7-e818-46c6-b563-51e46ebf87d9 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:31.598124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:31.598124Z digest=sha256:ec02879a694544dd1390c58941d567432aac1017cbaaea5c9858f18592630b2f

Observation e8c68cc1-77b3-4e49-af8b-3d246cf85831 · inbound

Scheming Ability in LLM-to-LLM Strategic Interactions cites this paper.

Scheming Ability in LLM-to-LLM Strategic Interactions Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:51:03.750005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T07:50:30.597108Z digest=sha256:795cce14fc64a0e74ae95f0e40589ff4c36751e14d293b1084c5f2a86e8d2289

Observation c691439b-6c02-4607-97ce-bce0e5225742 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.304935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.304935Z digest=sha256:1c140a82e546c3ec46c15c9b5f886fdf171a4722e35c4c292db044f33368db09

Observation 6777e281-4e8f-406e-9a6a-71c8f5b6d78c · inbound

Phantom Transfer: Data Poisoning can Survive Data-Level Defences cites this paper.

Phantom Transfer: Data Poisoning can Survive Data-Level Defences Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T04:59:05.071079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:59:05.071079Z digest=sha256:3f2d07e2bf432681427ff899df6707517909816cb5d37e8551e8da88ccc7f4aa

Observation 41cc4a35-8e10-49c1-a665-dadfb8cdaed7 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T11:21:28.998865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:e0765211c1b162444afd742b0a7b30a9f3d9e5b76ddd726bffce68cebd19fc8f

Observation 4bc4ceda-8734-4625-addb-79961c6c573b · inbound

NEST: Nascent Encoded Steganographic Thoughts cites this paper.

NEST: Nascent Encoded Steganographic Thoughts Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:21:31.165458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:21:31.165458Z digest=sha256:8bf8187462ce8282b4dac546791a0a4c73c56b043ad5829d53e4720ff6c749b4

Observation 21fcd595-f29d-4fd0-8931-13f04dc59457 · inbound

BarrierSteer: LLM Safety via Learning Barrier Steering cites this paper.

BarrierSteer: LLM Safety via Learning Barrier Steering Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:05:26.712893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T07:02:03.058731Z digest=sha256:56524515eb3fc449de787f3e805ad01f81e73c480e25ba9ae971d49c15976df8

Observation f1f45249-5e71-4b83-be97-9d66264b82b8 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:16:06.543864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:dffca0123ad5d0ace19f6ae562653559618e81bfd9ad989c05b23a1588fe2258

Observation ffd801da-ca43-4ed6-b6f6-99f976febded · inbound

Belief or Circuitry? Causal Evidence for In-Context Graph Learning cites this paper.

Belief or Circuitry? Causal Evidence for In-Context Graph Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.962333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T00:52:07.392800Z digest=sha256:c48c5bca47da1190b5133f3ca735d812a3962b820917c7ac05a95941aef40dd9

Observation 2bcdda77-c8c6-4b91-8ba4-782e7a398082 · inbound

Interpretability Can Be Actionable cites this paper.

Interpretability Can Be Actionable Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:17:22.798229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T06:12:51.656452Z digest=sha256:16b08a3cc6041df2883bba03bbf67ce8312a20fa498458f506c205f33cde3e66

Observation cf4c3b87-5ed5-45d5-b964-48f9624777db · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 144

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.282765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:57212bd74df0c2887141108e17e0211f47fd55aadd5e0fd7f3c2187cdba7ced8

Observation e4284353-edd1-4007-8b40-4521a20c3b61 · inbound

Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research cites this paper.

Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T06:03:08.437075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T06:00:59.614364Z digest=sha256:168e7d78c69a694a6b090b6bfb569e8a96c6fbafb8569b65c0c8e0e56f595bea

Observation b7dfc7f5-81bb-491b-ae4d-bf9f74b8ffa8 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.047592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:fa8eeadd5e5a8b1b7d67b0ce9537d6a38199216498496a3e21114469f96dfa38

Observation ca044699-010f-4031-944a-97e2aeaea091 · inbound

The Surface You Test Is Not the Surface That Breaks cites this paper.

The Surface You Test Is Not the Surface That Breaks Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.190683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-29T06:37:19.674012Z digest=sha256:57010d9c557c4fa948992d26595c678b52aee00c69f9be21595848e33590843a

Observation ee43106b-6bdd-435c-b545-49b57857f71f · inbound

VET: A Framework for Analyzing AI Discourse cites this paper.

VET: A Framework for Analyzing AI Discourse Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.479301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T14:41:17.714421Z digest=sha256:96b1483c730029b2d6bdd3ffa149b3ef381510a55cabb9603e2c2e5eec30d0a8

Observation fe12e385-1113-407c-b812-b8418590317b · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.165223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:e41d40a897de9c5fe1bd8c1e69569d35209823575897ef3892329870166737a0

Observation 2e78d5e7-5ce1-4849-b17c-6a011716ecbf · inbound

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs cites this paper.

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:16:34.761998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T09:21:57.373862Z digest=sha256:2fe906638754e8776e5de1a0d7096d80dcc54b77554d59a574acf068c91cdd52

Observation 46a39709-f2b5-41aa-892b-53ffb6583773 · inbound

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders cites this paper.

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.934660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T22:22:50.474397Z digest=sha256:bb4f9e09e60d2d383a486183ed8f989e96d6b2d14f275860d1e052d9b9fd4099

Observation dcbdf9e3-9228-4c45-bb78-1a1947a62269 · inbound

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard cites this paper.

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.017226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T19:04:07.735560Z digest=sha256:b31c37b07658244c1c8768ba1d2d98046dfbae0ea4844f4062536d49cc77c2bb

Observation 5eb8546a-7b7b-4288-9804-0fd91738197e · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.222546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:819a670c5478857bfe458134342d3dabe1d2ae77d3e783d8c2aee31f60d15c35

Observation e2cc9341-e13b-4344-aebc-f97c301d251d · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:29:42.338587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:9edd72673791f8926e223662e56dcabeace3834d6d62439eaf1c724d22092d54

Observation 2811f24f-910a-4c3e-b60d-9ea0d54caf92 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T22:40:37.839133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:40:37.839133Z digest=sha256:3d7633330eae68abe09fd231da597a7cda251d0dfa935b22f728bdd2228d8ed3

Observation 580bb6d7-6b73-421a-b14f-02db6794d56d · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:01:49.222325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:01:49.222325Z digest=sha256:916220e1c41d9ad70e9970be4eb76db1a9f56317b9c25883ecd76b64b290dd29

Observation 078c6946-2dbb-4e66-bc77-84bc0ee45a14 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.096313Z digest=sha256:30f8cebaa6f191281f023983553f5e4c052523a9d9cbd75bb8ba2b1bad9ddcb9

Observation b30a7b2d-a07f-4801-bbb5-93ca6fd4e08b · inbound

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools cites this paper.

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:11.241994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:11.241994Z digest=sha256:776a3a23c7dac29d633052d252520397557ba7091e4cc4fd12d89db3a3826b32