Pith. sign in

Paper Citation Record · LEDGER

Misalignment from Treating Means as Ends

As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2507.10995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10995 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:34:15.255851Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88265400-60cf-49f2-811c-ecbbfbd15654 · outbound

This paper cites Faulty reward functions in the wild.

Misalignment from Treating Means as Ends Faulty reward functions in the wild

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.644297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.445657Z digest=sha256:ad74ddd238e4c3edb30e023953529ffd568407dd8bb3818aaf950c87704c0a8a

Observation 79375c00-6796-4b4a-af29-0df029ce93dc · outbound

This paper cites Potential-based shaping in model-based reinforcement learning.

Misalignment from Treating Means as Ends Potential-based shaping in model-based reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.634052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.564815Z digest=sha256:366dba5f4da3800b0007109fafcf50457082f4d5a29cb84ba2954cb87c7f5a20

Observation 9afe300b-8795-4d1f-a755-f6937a1eae89 · outbound

This paper cites Discrete dynamic programming.

Misalignment from Treating Means as Ends Discrete dynamic programming

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.624257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.683133Z digest=sha256:606c40be96c1dfebc1927f0b3034d688da256e81b2e177650267227a69c2a5e2

Observation 7fb69216-d876-4780-a825-12f87281853f · outbound

This paper cites The superintelligent will: Motivation and instrumental rationality in advanced artificial agents.

Misalignment from Treating Means as Ends The superintelligent will: Motivation and instrumental rationality in advanced artificial agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.614442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.845517Z digest=sha256:9cbe24cdec99b0b589e0206bfe5043ba3301ba93d84f4ee0fa2807cb30d3bde1

Observation 1dda2239-00c0-4048-afc8-c3eddabe3697 · outbound

This paper cites Deep reinforcement learning from human preferences.

Misalignment from Treating Means as Ends Deep reinforcement learning from human preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:14.007127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:14.007127Z digest=sha256:e0a00275bc4946013233642671e2d2a56530ca7b12cd71c4ca80dd39609be3b9

Observation f4e68bce-d562-4453-8159-dfc05d25b0db · outbound

This paper cites an unresolved cited work.

Misalignment from Treating Means as Ends Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:34:15.597861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.115875Z digest=sha256:d840433ed3192c54bc23f3a8a7c40e1b357f0372347d56eaadd01f8c74b12a5e

Observation ee0d7ebe-8c38-4f58-8b6e-871135d1b94f · outbound

This paper cites Exploration-guided reward shaping for reinforcement learning under sparse rewards.

Misalignment from Treating Means as Ends Exploration-guided reward shaping for reinforcement learning under sparse rewards

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.588281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.267435Z digest=sha256:5a36666d808db275c3c6832c46b17807be584959f360ff2c9d2fb373babe7b54

Observation 59a02da4-7372-4db5-be5e-a2cb6fa5455a · outbound

This paper cites Dynamic potential-based reward shaping.

Misalignment from Treating Means as Ends Dynamic potential-based reward shaping

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.577539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.381757Z digest=sha256:a2f9f964e2836a2baeecc13189a642be4e83c2b156237b3e3754ed233bb68070

Observation b7db942a-cc12-495c-bbd7-0c354d537e01 · outbound

This paper cites What is it you really want of me? generalized reward learning with biased beliefs about domain dynamics.

Misalignment from Treating Means as Ends What is it you really want of me? generalized reward learning with biased beliefs about domain dynamics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.567489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.590020Z digest=sha256:623732f20f3034ac35a3d85e29b8579fd5a84f4c2280d7df7368d1f647174f2a

Observation 3acafe36-0836-4621-9bd8-5d3d59776dc1 · outbound

This paper cites Reward shaping in episodic reinforcement learning.

Misalignment from Treating Means as Ends Reward shaping in episodic reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.557064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.759732Z digest=sha256:c725d49a2f71bf084225a2b89308e2ae5ff0672e95a31d7d31bee596a85240dd

Observation a4264b5e-a178-477f-8c19-fd9242abc934 · outbound

This paper cites The off-switch game.

Misalignment from Treating Means as Ends The off-switch game

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:14.925169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:14.925169Z digest=sha256:5ab0e632a5b2b811a68a3677b17a7417d5d55b958d0c8f514d39d1a220e0bcfe

Observation 13b04838-d379-48d7-a2a2-0906e7b465d3 · outbound

This paper cites Exposure and response prevention for obsessive-compulsive disorder: A review and new directions.

Misalignment from Treating Means as Ends Exposure and response prevention for obsessive-compulsive disorder: A review and new directions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.541323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.097323Z digest=sha256:e2820c13f0a239930270678094d2cf64b3a215aaabb25871a7b1617a36a9ef0b

Observation f96126d9-8445-4a64-b98d-6a9385688096 · outbound

This paper cites Teaching with rewards and punishments: Reinforcement or communication? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 37, 2015.

Misalignment from Treating Means as Ends Teaching with rewards and punishments: Reinforcement or communication? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 37, 2015

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.531224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.181926Z digest=sha256:e30b7ad1e724bc635e1e900ea18b190d8616f8eda075bd5d7a73342907a58ad8

Observation 385de4a5-9185-4930-a313-3f4e81c05f78 · outbound

This paper cites People teach with rewards and punishments as communication, not reinforcements.

Misalignment from Treating Means as Ends People teach with rewards and punishments as communication, not reinforcements

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.521675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.185020Z digest=sha256:98f2abe2459254eae5b320d943e503e8ab762f0bb0257c854861720860e800bf

Observation 9385e40f-fbf7-4088-b447-795c1d6b83c6 · outbound

This paper cites Horn and Charles R.

Misalignment from Treating Means as Ends Horn and Charles R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.512021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.188111Z digest=sha256:6ddf231fc78837634427af67c2dbb6864ac54ba78c6074fd47619416baecc592

Observation 09d72ca4-d9c5-4a43-8820-b8070f38aa5b · outbound

This paper cites Reward learning from human preferences and demonstrations in Atari.

Misalignment from Treating Means as Ends Reward learning from human preferences and demonstrations in Atari

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.502179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.191190Z digest=sha256:c604782576dd1bb72143ab0488d964a8a706a2f8a04339b3df2054da79c7917f

Observation 34e377bc-249d-4306-b0b4-3c007a454b7a · outbound

This paper cites Interactively shaping agents via human reinforcement: The tamer framework.

Misalignment from Treating Means as Ends Interactively shaping agents via human reinforcement: The tamer framework

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.492914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.193878Z digest=sha256:538feee5301e56993445df8a0377220a49d3a4a9cd9eba35dd935ac13aac70da

Observation 3942d4f3-4cca-4c13-a3ca-f86360ca32ce · outbound

This paper cites How humans teach agents: A new experimental perspective.

Misalignment from Treating Means as Ends How humans teach agents: A new experimental perspective

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.483473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.196610Z digest=sha256:aa74e19ba1db25bcaf807e764b1619fd95763b68b3c672f1156075f97c0bf2ba

Observation b05217fc-f97d-452d-9eea-6474cbc25f63 · outbound

This paper cites Models of human preference for learning reward functions.

Misalignment from Treating Means as Ends Models of human preference for learning reward functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.199300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.199300Z digest=sha256:fd72f0dc9aded911e1c04fdb499e9ce767282e99749a4c9650aa038bbe785f70

Observation 3db4851c-dadd-4dbc-b68b-79336cf351f3 · outbound

This paper cites Learning optimal advantage from preferences and mistaking it for reward.

Misalignment from Treating Means as Ends Learning optimal advantage from preferences and mistaking it for reward

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.473562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.202726Z digest=sha256:60a3b68a58e449e9987063b52fcda7a40bcd229c7faf165b76b5427e65d411bb

Observation cd331bdb-b089-4b25-9a86-9a209c9fd58f · outbound

This paper cites BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping.

Misalignment from Treating Means as Ends BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:34:15.299575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.205426Z digest=sha256:09acadbaf1ad28d331cd3a06ead0057b88c995a96d3484c400f0876ef79d0df3

Observation 2c0656be-52b7-4a18-9432-4a76b93da383 · outbound

This paper cites an unresolved cited work.

Misalignment from Treating Means as Ends Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:34:15.463949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.208350Z digest=sha256:14f7855690a6e0db0912aa2ee3aeace2ab493d2aebf1a466edcf572f92d4dca6

Observation 6383da11-7435-4a03-bc2a-2451044abaae · outbound

This paper cites Choice between partial trajectories: Disentangling goals from beliefs, 2024.

Misalignment from Treating Means as Ends Choice between partial trajectories: Disentangling goals from beliefs, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.454002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.211065Z digest=sha256:e7508707abb9882eccd222bb17ec4819622ca3038b358e6d4abef042f0d58366

Observation 0400189d-aec9-4f9d-8d28-b305e0bf115d · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Misalignment from Treating Means as Ends Policy invariance under reward transformations: Theory and application to reward shaping

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.444941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.214008Z digest=sha256:7ce2023b8b642b00849af1ace092b6776fbc09d9ef4d4acc35dd7999621e6f3f

Observation 457b866a-e7c5-48ae-bcb4-c4308c19a0ab · outbound

This paper cites Anthropic’s new AI model threatened to reveal engineer’s affair to avoid being shut down.

Misalignment from Treating Means as Ends Anthropic’s new AI model threatened to reveal engineer’s affair to avoid being shut down

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.435244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.217137Z digest=sha256:6c2849fe4e262b2062211e799142091a841ad21b1de522d7637cc144f12508a8

Observation 2b3f10f5-6b0a-408d-84b1-8e72cf5ce63c · outbound

This paper cites The basic AI drives.

Misalignment from Treating Means as Ends The basic AI drives

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.424662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.219920Z digest=sha256:1d4a1874e81f17e70bc755d98dd7f9c92a6ac228a3e01101e80747ad6d92982a

Observation 8d5bda20-19af-46a3-a813-f9933b54f384 · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

Misalignment from Treating Means as Ends Learning to drive a bicycle using reinforcement learning and shaping

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.413722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.222968Z digest=sha256:2a22b6cac0278bee27d810c1024276ddf134aa634840243b63aab55ff3134de1

Observation 45e6ca2a-9107-4ec4-9254-7dea463fec9a · outbound

This paper cites AI is learning to escape human control.

Misalignment from Treating Means as Ends AI is learning to escape human control

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.402944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.225939Z digest=sha256:d4a818368f37025777a1aab260a7438af0a3aa9c24065d810da8ba9f411f46c7

Observation 4e66a997-67d8-41d3-be94-1999cabb3f9e · outbound

This paper cites Human-compatible artificial intelligence, 2022.

Misalignment from Treating Means as Ends Human-compatible artificial intelligence, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.393424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.229124Z digest=sha256:ccd00803750ffbc055b7d440d0bfd2fd025b08310d80f0f978fa97736a34feb9

Observation ae0afa82-5d18-4aab-948f-d298f630f731 · outbound

This paper cites Artificial intelligence: a modern approach.

Misalignment from Treating Means as Ends Artificial intelligence: a modern approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.383047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.232151Z digest=sha256:bf4e0121b8955fc1e372c3bbfde8eec97377771a9e831071c48c65d53b27f7ef

Observation 3c1f24bb-19f9-4cfa-ba51-389bb9f34d54 · outbound

This paper cites Where do rewards come from.

Misalignment from Treating Means as Ends Where do rewards come from

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.372851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.235350Z digest=sha256:674b22d613c9e8d9869a81c01dae129a235c500e702ce62ed0a4bc0c02f8023a

Observation 87166d99-fb94-4a85-aa04-c247f9d643e1 · outbound

This paper cites Intrinsically motivated reinforcement learning: An evolutionary perspective.

Misalignment from Treating Means as Ends Intrinsically motivated reinforcement learning: An evolutionary perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.238058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.238058Z digest=sha256:b5ee31a89d4f984a54a712c8158e9ac92cd0143aaa1edc642a0800502fc408c2

Observation 9bff7598-8344-439e-b985-d32ea35a874f · outbound

This paper cites Corrigibility.

Misalignment from Treating Means as Ends Corrigibility

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.357143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.240721Z digest=sha256:0b9333f972e1dab6934b385d32e184c4ea93957aeccb21becd0dd45b4e5bb496

Observation 028f26ba-7548-48ff-958d-1dcd59cf1be7 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Misalignment from Treating Means as Ends Reinforcement learning: An introduction, volume 1

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.243611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.243611Z digest=sha256:1174ab0e3e77e2d3515420039672e45aead75fb459cc33b495f4505988a0ba4a

Observation 6f36d324-050b-43f8-8cc8-0a64e926a041 · outbound

This paper cites Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance.

Misalignment from Treating Means as Ends Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.340391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.246831Z digest=sha256:dbfbd1ef8bd9a5a18fe7e9197efa17ed0745dd3bed4951e3c82a7ac538737fd0

Observation 16481d20-58c4-41ae-a253-d4e787523bb0 · outbound

This paper cites a ngberg, Mikael B \.

Misalignment from Treating Means as Ends a ngberg, Mikael B \

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.329928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.250509Z digest=sha256:4db2e0f86c969b0b4961244f7c6be7da77ad930ba5b877054f7196590bb07426

Observation 63e2764a-c98a-46d2-a69c-2090a722c3fb · outbound

This paper cites Principled methods for advising reinforcement learning agents.

Misalignment from Treating Means as Ends Principled methods for advising reinforcement learning agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.319816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.253266Z digest=sha256:6fdbf97f09ce1805ece40211636069d45f7f3aae8d056e3108d4bbfea5565b48

Observation f12742d6-154c-4225-adfa-8cd93a2728e0 · outbound

This paper cites Reward Shaping via Meta-Learning.

Misalignment from Treating Means as Ends Reward Shaping via Meta-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.255851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.255851Z digest=sha256:19a0d90ce48224988672a05390719cde67f56fb4522d1abbdab9c43d9929ea98

Pith citing papers

No inbound Pith citation observations are available.