Pith. sign in

Paper Citation Record · LEDGER

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2304.03279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.03279 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T12:46:26.589118Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6e829ee-a8f4-4e19-ac1d-65563706e160 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:53.622855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:d60c4ef38e372979b698dd34186cb566369077b5ba36be1eea7ce8e8b90f0e25

Observation 88056ed4-f8c0-44ed-bb6c-0d29398f998f · inbound

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models cites this paper.

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T12:46:26.589118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:46:26.589118Z digest=sha256:fb1699fad1fdb5f49f5c6d2c7fb47e319577f18e44c6c3edbd96e18cb96c5d1a

Observation 99247684-0c9e-447f-aafd-d8371dd15337 · inbound

The Odyssey of the Fittest: Can Agents Survive and Still Be Good? cites this paper.

The Odyssey of the Fittest: Can Agents Survive and Still Be Good? Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:23:42.468613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:23:42.468613Z digest=sha256:aab737a5f750e8194320e8bb806fd7dc9303ccff0d499d691dc0a70568f3fa54

Observation b307d4f7-9636-4544-bb80-90daabd2a075 · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.435854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.435854Z digest=sha256:6232e4f729ab05745ef670e492c8583253ad4870910e2a8a2cb28d7be52e2810

Observation 1192a071-1ed1-46a6-a603-41d70e10d646 · inbound

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs cites this paper.

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T00:04:57.367167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:04:57.367167Z digest=sha256:62d3ccdc6a08d5816fb6e7803bdf32a65916b9b64f7a31cbc92a09fa28ac8ec5

Observation 7f55633a-4be4-43d8-adfc-2160525c1239 · inbound

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models cites this paper.

Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:20:27.811163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T23:16:56.957905Z digest=sha256:f21d0b1f5bd31bce9ce1f93e6534f5ac8243f543e8849362951366e16db42b7c

Observation 3249d592-9433-4f85-99ea-2d2ebb4485fe · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.244031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:c203433839e8ae0eb220fcb82be39a4f313c6bebc29d7b6139ae12fb39176ed8

Observation 3b8da89d-87e0-4d07-b336-25fb556d38f0 · inbound

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves cites this paper.

Restoration, Exploration and Transformation: How Youth Engage Character.AI Chatbots for Feels, Fun and Finding themselves Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 64

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T12:45:37.283306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T12:44:53.677818Z digest=sha256:2a00ff8e42db6dee364598a1260c0d83f23d4a8b27c4d7b0616ebecb87b4e5ef

Observation ff9e1600-92e5-4847-8f14-6e8ccc1694b7 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:47.953274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:e692f8ea2a37ad245f76b87f6f7b87af2816b27a3de1af2718364f6d62ecd20b

Observation ce6bc9a0-d7d8-4f36-a07d-38740ddff390 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:55.545244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:9b3b07167c9b43ac32d9216d724f865de1523ffd08df2aca5eca23297f0541d9

Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.709741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:e2ea0b33b24dd771af56814113d0ec3bb2823c87d5663c669519b774e1e38268

Observation d0da8262-22b1-4148-8e7c-96665a13f2c1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.014677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:11042b5c9e1000094983cc0bf231323050b534b2a9f109c2e6b73f2ff5c6c0dd

Observation 8d56c348-3b2b-49cd-a2bb-bd8e7a2fa2cb · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.199764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:f0e4605d7c2afbd824ca1ef3c6566c20b3f1849a23be548cd2e5433f3e9fd38e

Observation 3420d241-80b2-4375-a7a1-d63611d127b2 · inbound

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems cites this paper.

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.747048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:53:37.616447Z digest=sha256:ff7125bcdba03f26630bf9f16418cd60c69e765895349ebe904d715150ec2013

Observation 14430cf5-520c-4fa0-b301-bf36fb5c54a2 · inbound

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing cites this paper.

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.662641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:28:38.810889Z digest=sha256:588dfe96018a58b251583a2769f895f03c50a74c03548d5284ed73c106673a23

Observation e3602deb-3912-42eb-ab3c-2899864aa92c · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:30.716801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:30.716801Z digest=sha256:0d9536c73fd0ba6f6f9fa1df9f7b7a8e9f4b0cb1816a365dd240d95ff07335d4