Pith. sign in

Paper Citation Record · LEDGER

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

As of 21 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2506.13774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13774 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.337382Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03ffc49a-1a7b-429c-8020-54a96502feb0 · outbound

This paper cites Artificial Intelligence, Values, and Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Intelligence, Values, and Alignment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.261527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.102830Z digest=sha256:d530905387fbd26d4aea31f563a9f3bf0c213a96693c3cf58310955faac52feb

Observation 93026cbf-d874-413a-9bd0-b2651e5c8a3b · outbound

This paper cites Artificial Morality: Top -down, Bottom-up, and Hybrid Approaches.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Morality: Top -down, Bottom-up, and Hybrid Approaches

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.247705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.113520Z digest=sha256:4b57e9f91773ef394ec996093770915ac67b70dbf633919a717ac6012684b8be

Observation e4ec19f7-68a5-4123-90ff-cb5f9267b597 · outbound

This paper cites Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.232427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.119100Z digest=sha256:1e89e64fda13f81d598db88084c89b976201d46d204f9802705754363cc61f71

Observation cab35d18-a447-4352-ac03-fd37e5153ac7 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.123823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.123823Z digest=sha256:6bc7188c4a5ca6250aa50fa85c4d446f7191900b14b16431ab4ecc6cdba5a7b2

Observation c5354f90-fad6-4bb2-96cf-a66ca53ebf2d · outbound

This paper cites Personalized Large Language Models.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Personalized Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.128902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.128902Z digest=sha256:a78dc0b7d3b55c825c51f53ce6d6934e7056b9104032dfd22b9a844ffbfe3854

Observation 49e8007a-9ada-423f-8dbc-f1bc1335f848 · outbound

This paper cites Towards an End -to-End Personal Fine -Tuning Framework for AI Value Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Towards an End -to-End Personal Fine -Tuning Framework for AI Value Alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.218849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.133884Z digest=sha256:143363042cbd8d5a6e2636c54e456623dd5dc10c92fabea0f02f7c68a766ef25

Observation a1f78d6b-2a43-4697-b583-f52683d34c35 · outbound

This paper cites Safer Agentic AI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Safer Agentic AI

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.204958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.139059Z digest=sha256:d10c0158c858d1aaabe59f2dcc156ae9cb6f1355ccfe115799d2ce340c25e584

Observation ab4fc287-4bcb-4053-b2df-466b9746a33d · outbound

This paper cites Introducing the Model Context Protocol.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Introducing the Model Context Protocol

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.190940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.143540Z digest=sha256:8ce99e420026ca2ab5f1150306e569029e7d6d722d89f7707ed3f869c441dc44

Observation 0b0b6157-c207-4c1b-b451-66d573f3884b · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Deep Reinforcement Learning from Human Preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.176703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.147943Z digest=sha256:65156a1add6e0aba824751411853d81d5bbfd66fb2e3869aa91a6674a83beedf

Observation c1a6a662-4abe-47f4-b593-49c0d6ede551 · outbound

This paper cites Confabulation: The Surprising Value of Large Language Model Hallucinations.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Confabulation: The Surprising Value of Large Language Model Hallucinations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.152276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.152276Z digest=sha256:aae685f89bb51cb7e9b568920984520507e803967a775441c338b547b86e3382

Observation 75f881bc-d2a2-4cc7-a983-2f2387095612 · outbound

This paper cites Choice Vectors: Streamlining Personal AI Alignment Through Binary Selection.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Choice Vectors: Streamlining Personal AI Alignment Through Binary Selection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.161821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.157090Z digest=sha256:c41f7146459fd96d92de6ada49b2b70d9f1521fda9a3826445d02cf877efb6da

Observation f8843945-5b73-4fb6-8fdd-6eacf9e521a7 · outbound

This paper cites an unresolved cited work.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:43:08.147667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.161465Z digest=sha256:70b6415887754802497891c54418deb9717fa54359acfd8be10c8b85394e5746

Observation 232768fa-1ba1-48b4-b37f-24a77dad24c0 · outbound

This paper cites Universality of Representation in Biologic al and Artificial Neural Networks.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Universality of Representation in Biologic al and Artificial Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.165649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.165649Z digest=sha256:d46bdd15cdb2a28d9d796e7dd5c15690da8517edf6613577077e1935e867be32

Observation b5015ad3-26fb-48f4-8a6f-ff1489e4c7a9 · outbound

This paper cites The neural bases of cognitive conflict and control in mo ral judgment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural bases of cognitive conflict and control in mo ral judgment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.133544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.169921Z digest=sha256:c4cbfeea2c7153207d1c8be2dac7f53d4c1b0b09e12bc9cc6b2beab2d23a9e80

Observation 5f808b93-0ce8-402b-a21e-fc5f0480cb73 · outbound

This paper cites The neural basis of human social values: Evidence from functional MRI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural basis of human social values: Evidence from functional MRI

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.118460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.174334Z digest=sha256:cf3dd9d38bd066fed676ad32b1c96e29be6003a64837bf4541b59b65c06e7c19

Observation 2d831fb6-b9b9-40aa-af80-9648e02cea1c · outbound

This paper cites A Cognitive Theory of Consciousness: The Workspace of the Mind; Cambridge University Press: Cambridge, UK, 1988.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values A Cognitive Theory of Consciousness: The Workspace of the Mind; Cambridge University Press: Cambridge, UK, 1988

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.103481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.179006Z digest=sha256:b78d60c52e3f79c780a2a95c982d38c505b7f611e670cc964b1916e3264922c0

Observation 6886e3ee-5210-46b2-b28e-efef9e830027 · outbound

This paper cites Unified Theories of Cognition; Harvard University Press: Cambridge, MA, USA, 1990.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unified Theories of Cognition; Harvard University Press: Cambridge, MA, USA, 1990

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.088209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.183568Z digest=sha256:182cb9b01f5a3658ff0440f5cdaf166bd7ada83fdc7d481146b70d0132914af5

Observation bec20e91-f154-4787-a276-9d29e7bc19d6 · outbound

This paper cites Revealing economic facts: LLMs know more than they say.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Revealing economic facts: LLMs know more than they say

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:43:07.780589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.188289Z digest=sha256:ca4e65a00cd129ab474785d553c48ec70a07076fd76c3537d0f853da9295af59

Observation 565c71ff-9edc-481f-a63f-4b03b833470d · outbound

This paper cites ShieldGemma 2: Robust and Tractable Image Content Moderation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values ShieldGemma 2: Robust and Tractable Image Content Moderation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.193249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.193249Z digest=sha256:0661742b1f9ce9b479c9defe588fc0f9e77e3d74614a2092e7f8826cedbc8267

Observation 40c700b9-bd04-4265-a0d1-417a7d2802c3 · outbound

This paper cites Superego -Agent LGDemo (Branch: Fastapi_Mcp).

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Superego -Agent LGDemo (Branch: Fastapi_Mcp)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.073012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.198921Z digest=sha256:5962099522bf5ecac6cc390cdf5820d7998209d4178ca9b1c61f574b72e48219

Observation c55222a1-896d-459a-9f3e-dfd48bfa9033 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.203855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.203855Z digest=sha256:ae36650a9ecb50c62a8688d513b07d743b2fa974bed50b403b2e88777a61e5e3

Observation 218dadac-60a5-4f85-909e-cf5e6a25d19e · outbound

This paper cites AgentHarm: A benchmark for measuring harmfulness of LLM agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AgentHarm: A benchmark for measuring harmfulness of LLM agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.058305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.209304Z digest=sha256:15f37adb8d9d7044429eaef91f6b6cca8366e9ec23245c64934516063477ce17

Observation 0d21dcae-dcab-4951-b7a8-6e0cf673c985 · outbound

This paper cites Do the rewards justify the means? Measuring trade -offs between rewards and ethical behavior in the Machiavelli benchmark.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Do the rewards justify the means? Measuring trade -offs between rewards and ethical behavior in the Machiavelli benchmark

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.043769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.213631Z digest=sha256:36dab579679ead71dba220f06378c9d8bd3c30e1c1f390d4dcc08cc405c3b7bf

Observation 7d230248-3583-4771-9434-10d2f6bcc2e5 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.218428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.218428Z digest=sha256:783f756885f7f9955b4cfe7b122644cae1cda8fc30291c293ab307592f37049f

Observation b9a1bd80-b170-4952-ac58-bb35175f738c · outbound

This paper cites Vijil Test Library: Evaluating LLM Trustworthiness Across Eight Dimensions.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Vijil Test Library: Evaluating LLM Trustworthiness Across Eight Dimensions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.029110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.223161Z digest=sha256:7c81bd369d73c2c63ecdbe9ad3ec7f290b57fe70f56c02e99ab8be45798b70f7

Observation dd65809f-b519-4dd9-aca7-bc452cb6adae · outbound

This paper cites INSPECT: An Extensible Toolkit for AI Behavior Evaluation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values INSPECT: An Extensible Toolkit for AI Behavior Evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.013655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.227370Z digest=sha256:51eab60ba3537f4e36bfa53196c17ee905eaad4e2a8c69216745a602c72ddd85

Observation 5c5cdae8-08c2-4c38-a138-b11566074059 · outbound

This paper cites Governance in Agentic Workflows: Leveraging LLMs as Oversight Agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Governance in Agentic Workflows: Leveraging LLMs as Oversight Agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.999467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.231834Z digest=sha256:e787e564268272d85b378c110cc528540606acf460671a0e064946d86460bf55

Observation f34152c1-3bdc-4796-b437-1ecc6eaf0dc8 · outbound

This paper cites Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.235972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.235972Z digest=sha256:86e006a2c06b0288d2452d0e9fafc5c9e990bacf0d50e9a8cf546dc972ff757a

Observation c56c63fe-b2f0-40db-8b6f-319be09db8c4 · outbound

This paper cites Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.240615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.240615Z digest=sha256:d03177afb54f02765d6371439dea7e43f4f05913e829ee87297fc2f767b2d91f

Observation 01f226e5-346c-4313-bf58-63314a39d76c · outbound

This paper cites InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.245116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.245116Z digest=sha256:a3cdfd1ce1c832e41755d20d5cbab7d392899b20a9950aa4fc46d44aaff82b4f

Observation f47fd653-d640-4acc-b755-e93a1f270232 · outbound

This paper cites On Almost Surely Safe Alignment of Large Language Models at Inference-Time.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:07.630186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.249475Z digest=sha256:2541d561c658d1658d1494ff31095574768670d73639a1f21c5f4a243dadb955

Observation 535df10b-43a8-4d3e-afd6-35257bf5828b · outbound

This paper cites Dynamic Search for Inference-Time Alignment in Diffusion Models.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Dynamic Search for Inference-Time Alignment in Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.253673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.253673Z digest=sha256:e32c2e1e4ede796417990dd7aaf9a24f388fada8caf438daa58097fe9ac2e5e0

Observation ad083c49-e7a3-48d4-8c78-de4e9196656b · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.258162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.258162Z digest=sha256:d2e8e1b1f75f52aa81e8b056b76e17f403429c2aaefa0c0e81e612d6da38973b

Observation b7f7c262-ade5-4959-a0a1-84e014de2c97 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional AI: Harmlessness from AI Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.262772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.262772Z digest=sha256:906cffb72bc8dde7ba3321132490937618f1a339f29e2749d093c0788d9ff7f7

Observation 7ae553b9-9542-4f09-986a-e2bb36766dc6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.267381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.267381Z digest=sha256:978976252395d53004032478e33a5d9fb59cbe387346fb35c65d01f064f79bcc

Observation 0307194a-196c-4c9a-adca-c6c142cca4c6 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.984868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.272296Z digest=sha256:fd5613d81375b0ce250c6d2f9a192fd2d6cd469897544a719942e64f37ceb9cb

Observation 171bc1d9-efa4-4d45-85cc-b610962d7f10 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AI Control: Improving Safety Despite Intentional Subversion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.278246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.278246Z digest=sha256:845fd76ca7f68886e8386c1719b31e788e4b41daf0538e818137d50c09026412

Observation f222863a-22fb-447c-badb-4127e9498e77 · outbound

This paper cites OpenAI x DFT: The First Moral Graph.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OpenAI x DFT: The First Moral Graph

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.970374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.283191Z digest=sha256:b15f95657307b426575041d483b582785b7d9215c991f888d00cf5ac5e513edc

Observation 8ff4d278-26db-4e2f-a2be-693040d26acf · outbound

This paper cites Model Integrity.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Integrity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.956049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.287785Z digest=sha256:696838c797422f1b997164ccdc034de02a3e409b58c8eb9af1d6e30216c68b73

Observation 27bab82c-a0dc-4c0c-ad58-3d049b54f333 · outbound

This paper cites The Global Landscape of AI Ethics Guidelines.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The Global Landscape of AI Ethics Guidelines

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.941746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.292294Z digest=sha256:1c8c7887020b4bbadff06ac1ce762d3fe295febb682204aaae058392b8f367f3

Observation 5d78eeaf-2263-4a86-9a54-c9b9b492120d · outbound

This paper cites WhatsApp MCP Exploited: Exfiltrating Your Message History via MCP.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values WhatsApp MCP Exploited: Exfiltrating Your Message History via MCP

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.926146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.296584Z digest=sha256:3efb92c0a930d6662c56947fec4b3a2d19d7027e1f3e50885e01eee2e33c25ea

Observation 48e2fdca-9791-46d1-a005-6bdfaca00cf9 · outbound

This paper cites MCP Security Notification: Tool Poisoning Attacks.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values MCP Security Notification: Tool Poisoning Attacks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.895246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.306568Z digest=sha256:399e36ddcded1d8c42b382fbc613c3caa236b16dac5f39d8b5058459f5fe65d3

Observation 20244967-148d-46a2-bb57-cdfbdc0be7ec · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.311471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.311471Z digest=sha256:38a445150834f788da10f9d1d33d210f6d0f76f13c2b8a92b6b12058e51349e8

Observation 31272682-141c-4702-a791-551e22302342 · outbound

This paper cites On Emergent Misalignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Emergent Misalignment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.880226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.315617Z digest=sha256:d4f214612415d493035b609ce83d5cff6e86720886419d976c1704e7405029f2

Observation 5f40ad55-bb11-4b30-bc91-b86935cea64a · outbound

This paper cites Model Plurality.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.864809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.319836Z digest=sha256:4b21322fbbfdee0584998ca046b248389fbf19b41c6aecc011f6bcd1a5e9c7f6

Observation 37621a40-2612-4ca3-812b-e77b1cdf5562 · outbound

This paper cites Model Plurality: A Taxonomy for Pluralistic AI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality: A Taxonomy for Pluralistic AI

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.849749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.324191Z digest=sha256:ffbbc0bb9701c97511581cbe838a4e2adc1d2735ccb00b73fe1f7ffec7adfc8b

Observation 6f6e36b4-7098-4f6c-9730-c2637ec23bd0 · outbound

This paper cites Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.328340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.328340Z digest=sha256:f6a048069e4ba8a1d19c4d3515d438d67c4ca11e0d5ddb2e188ccb196f178c22

Observation dc7ee6ef-2dca-4126-922c-1c188a5c2c2e · outbound

This paper cites OASIS: Open Agent Social Interaction Simulations with One Million Agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OASIS: Open Agent Social Interaction Simulations with One Million Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.332394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.332394Z digest=sha256:ff2d7c2efe7a5ca3266372c4ae497f880f0bc238e4d786722cb77154ab26a55b

Observation 1b915a17-c0ea-45d4-93aa-e2fcfdfb4806 · outbound

This paper cites Project Sid: Many-agent simulations toward AI civilization.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Project Sid: Many-agent simulations toward AI civilization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.337382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.337382Z digest=sha256:011377c28fa1ec6b8cab47a3782eaa0927f12290a75daa24f2b50192f35bbf8f

Observation 106163d8-0eb8-450e-9b6f-081131884779 · outbound

This paper cites an unresolved cited work.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:43:07.910198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:43:07.301344Z digest=sha256:c380ef74697f51aae98f6a81f33089ad91d7ca2d37b5a05b5cec6961f200a2f7

Pith citing papers

No inbound Pith citation observations are available.