Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2304.03279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T12:46:26.589118Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d6e829ee-a8f4-4e19-ac1d-65563706e160 · inbound
A Roadmap to Pluralistic Alignment Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88056ed4-f8c0-44ed-bb6c-0d29398f998f · inbound
Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99247684-0c9e-447f-aafd-d8371dd15337 · inbound
The Odyssey of the Fittest: Can Agents Survive and Still Be Good? Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b307d4f7-9636-4544-bb80-90daabd2a075 · inbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1192a071-1ed1-46a6-a603-41d70e10d646 · inbound
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f55633a-4be4-43d8-adfc-2160525c1239 · inbound
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3249d592-9433-4f85-99ea-2d2ebb4485fe · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b8da89d-87e0-4d07-b336-25fb556d38f0 · inbound
Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff9e1600-92e5-4847-8f14-6e8ccc1694b7 · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce6bc9a0-d7d8-4f36-a07d-38740ddff390 · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 228
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · inbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0da8262-22b1-4148-8e7c-96665a13f2c1 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d56c348-3b2b-49cd-a2bb-bd8e7a2fa2cb · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3420d241-80b2-4375-a7a1-d63611d127b2 · inbound
Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14430cf5-520c-4fa0-b301-bf36fb5c54a2 · inbound
Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e3602deb-3912-42eb-ab3c-2899864aa92c · inbound
Engineering Trustworthy Agentic AI for Critical Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.