Pith. sign in

Paper Citation Record · LEDGER

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems

As of 14 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2604.04237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04237 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T10:37:39.719847Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40ce4411-dd40-41ac-b4df-e315310c104d · outbound

This paper cites write newline.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:6e7d3e1d64eb1ec9d3abcd11996e55d88a192416b8aa0a7075e3f825bbf963b8

Observation bf7b6815-5323-4bee-a3e3-aa3271b4bc8b · outbound

This paper cites , author Hostetter, J.W.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Hostetter, J.W

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:2301875307aec7c5b97898f9c6318ef562ca6c5ca2dc70840006072a14c51157

Observation 1c97b2d0-899b-42ca-b599-e410782cc893 · outbound

This paper cites , author Held, D.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Held, D

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:f93db929785815fef1f3eac4016db99903f53c441f12b119e4a668393a7fca00

Observation 29823a9e-f0ff-4227-b4f5-6a79b7472b38 · outbound

This paper cites , author Fazeli, K.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Fazeli, K

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:df1b0b68af68f783b67845e0549ef5eac353d9a384b40ab4340f0aa8151d6231

Observation 7d18bbdd-b03a-4e01-af64-b320989a3f56 · outbound

This paper cites , year 1999.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1999

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:f0a0cc60153dd97f2363dcb5f102ff4ffdb01d2cdf2f7e62e4b58ce087269721

Observation beb88147-301b-40d5-935b-5c50bdd0f39e · outbound

This paper cites Concrete Problems in AI Safety.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Concrete Problems in AI Safety

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:bfb79b6360c780cf3391daa053655c5c6f67584e730e7743a452c963bef5dccc

Observation a89e3796-5ad1-46fd-9f3e-3eebb906cfc0 · outbound

This paper cites , author Corbett, A.T.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Corbett, A.T

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:f0c19ef74cc4f6936414582fc351dd053d0726b908427fd6c3bd762d03d75aa4

Observation 4310bc6b-a66e-4226-a659-b37f8772d567 · outbound

This paper cites , author Krathwohl, D.R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Krathwohl, D.R

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:5dd77a12b781d001bf379d6c3c1e3df460f85db7bf26b2897ad0c8fa4cb1e836

Observation e3861d88-9779-40a3-b54f-2b624855c054 · outbound

This paper cites , author Corbett, A.T.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Corbett, A.T

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:1b183e2e40d50cccc1fcbdaa6c2e1e86aff132f920ea18f902f139a7ba975a0a

Observation 70273c54-9e71-4ece-9b82-126d49b1f45a · outbound

This paper cites , author Corbett, A.T.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Corbett, A.T

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:89e6738c3d8c6d18f78ee2e6008703bd40133591ca7131d003e1906db98137a4

Observation 093c9473-99a1-48c2-b773-c07682ef2255 · outbound

This paper cites , author Corbett, A.T.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Corbett, A.T

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:8ee9c47d67f5861034d870039488dfccae28591dab608007b63ca6b3ebd25259

Observation 9f58c92e-3c58-4cc4-a528-91ff57b84dce · outbound

This paper cites , author D'Mello, S.K.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author D'Mello, S.K

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a883dfe25c29f59833a11d1909df1649293ba78cab43cea0b8f8960e095ae4a1

Observation a661adf8-4519-4e5b-bdbc-365e88a2b1ac · outbound

This paper cites , author Hawn, A.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Hawn, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:b1357b3f99eeea29654db8eaa443438dd4abe6e4910059c5ede1d2096c7aeaed

Observation e37989a0-fbd2-47cd-a66c-03e2f9f46cb8 · outbound

This paper cites , author Woolf, B.P.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Woolf, B.P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:b81a30bfee65aa991cea5883e6dba313d31ed9e8bda745f77aab207579102134

Observation ab24a9bf-af5a-414a-b647-3a2d16efaa3e · outbound

This paper cites , year 1984.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1984

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:883d5b72d0c84917788ad54b557ef19023b0d1c90eddcf9e78e7eb733247eba4

Observation e5330711-0d08-4931-a069-0acc578ac256 · outbound

This paper cites , author Wylie, R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Wylie, R

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:54200a1b96ab50c0530b2bf1a28725ee62496ac3e1873df587671ed4e6b1e35d

Observation 318bf2ff-7c18-427f-91ce-6676dd1397e9 · outbound

This paper cites , author Leike, J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Leike, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:97f4cc7383fe6de602b62d3827d62094ca639eefc68e056ddcf49aa970c45d34

Observation 6afe64e1-d35b-43f3-93ee-9c714b8df478 · outbound

This paper cites , author Roy, D.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Roy, D

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:182657c55906c2577c49cb2b7a7e2de71159c8667455b38f8f08402d98b02507

Observation 36541357-8847-47f5-830c-1324558e10fe · outbound

This paper cites , year 1988.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1988

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a9cac67c92c07c9a448bc9ec717f231178d1a1cf39862730849a630139506dc6

Observation b14ee215-6cce-4ebd-8d7b-da1adb7d2940 · outbound

This paper cites , author Anderson, J.R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Anderson, J.R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:b655212cb5f958fc3e8831c4965bdfd2af02ebb1a48816242f16eb8437b6f2f1

Observation 0df90e53-46af-4820-8ad4-57c5c152bb10 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Safe Exploration in Continuous Action Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:5f35fd5ca940b99e9dc9e9d64a6f0652ec8f7bb6fcf6be581c5eb28fc717d5dd

Observation e36e242d-a1f1-4f89-b6df-3dc32c952020 · outbound

This paper cites , author Koestner, R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Koestner, R

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:15b57c538c109d138b55c948e2df362ec0d8d181e3e1958d7b15ac69d8b376c6

Observation a0b0a3ed-429f-45ba-8539-35cdf40ca9ce · outbound

This paper cites , author Graesser, A.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Graesser, A

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:90e8d63a9aae6afdae2082cf5b9c8f43ce5471e540e2d17a36cde9eeeb2135c8

Observation 409f4ac6-92cf-4c3a-9317-9cdff90a5c09 · outbound

This paper cites , author Aleven, V.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Aleven, V

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:030b373c733dfa9342b78c72a745ffb8a9c8e91a5c1f3124754aa765edc2cd13

Observation 2fc8c0f6-865f-4505-82a2-67754d8d0c58 · outbound

This paper cites Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a5f6d752e0249c768c5699150224b2c89049df11f60ac226a1c1c818effcbc21

Observation 7c3b945f-1423-4839-accd-40e24e191850 · outbound

This paper cites , author Blumenfeld, P.C.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Blumenfeld, P.C

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:437a4e50455805644f57bd134a255d5c600f07fcfa4b80f92fd50ce58a1de483

Observation d36d7c72-501c-4819-9880-c757c0608499 · outbound

This paper cites , et al., year 2024.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , et al., year 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:824af18b68e65592cc04bcfddd36e3bce99e3e62b19d6b8738fab25165a98e09

Observation 5577933d-b3b2-4c2e-870d-e415aa6e0733 · outbound

This paper cites , author Fern \'a ndez, F.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Fern \'a ndez, F

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:17066783be7b378fc6a350a1a4be076c6a3258f6cea4f65d3795e3b61f0e66a7

Observation 77bd082a-d9a4-4ef3-a765-020dd8ec057f · outbound

This paper cites , year 1984.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1984

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:e4bd84286fa6c71322fae289d41c6f31ba05f09b549205fdadd47e4ddfa227ba

Observation 72aee1ec-49b5-4eab-8b52-6afe0bcc686a · outbound

This paper cites , author Lu, S.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Lu, S

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:da59a1703ba8486abf123c696f6de8abdee8d23099e6e8a181b5032ba225efd0

Observation 3fbdc9cc-0036-4a49-8f90-65c78d1d27eb · outbound

This paper cites , year 2021.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:4e36f61537959ab66bb88d3977541c0050333f126a8a02c7328979de24b944cb

Observation 627646d9-3a16-478f-ae4a-90ae3bab464a · outbound

This paper cites , author Muthukrishna, M.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Muthukrishna, M

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:528ff905dd4ffc14c3b4ff57748d513b57f48b0d54c77d3a8cbe137098c71a9b

Observation 49551fca-5b30-4adb-9de6-be9135f44b64 · outbound

This paper cites a llstr \.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems a llstr \

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:03672725f98c1bb97256077e6e8aaa168a7110c75c21a92177e64d0b890e32b2

Observation f4588a95-0c03-49da-b2ec-71ab56404821 · outbound

This paper cites , author Heffernan, C.L.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Heffernan, C.L

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:139ae08bec0078064e02cc0e9bc818107cb086a18abdb32eac690471bd932384

Observation 8a849791-d96f-4588-8f90-2b0b549d6ce5 · outbound

This paper cites , author Porayska-Pomsta, K.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Porayska-Pomsta, K

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:fc768846533bb713291b826e8c7d53d74966e0c22e86c449f4b7bf6019a6c65a

Observation d0f97767-a952-4bbe-ab22-7679550f7d17 · outbound

This paper cites , author Wortman Vaughan, J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Wortman Vaughan, J

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:7557598c1eb72cf5fc1389be998228ab69c4336b054fb7136c3d9627393c53f4

Observation 3a3e2419-d20e-4273-b1ff-8ace4dbc054f · outbound

This paper cites , et al., year 2024.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , et al., year 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a9caa8e770636ee9638552b421f3d4d48f603fc70ae6d84b1a10b3953e81dc91

Observation eabe6261-a691-4b33-88da-1024e022b1bf · outbound

This paper cites , author Mart \' nez, P.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Mart \' nez, P

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:0847915d4b5a825e5173a4758345589ef934f07bb51285937c5ac3efa614940f

Observation a6d10bef-73c2-4556-80fd-894e1935e1c5 · outbound

This paper cites , author Yang, X.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Yang, X

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:d47131e3b731702642ec80d89d9395127988efa1f767850d28367a5d4865893a

Observation 972cbfe0-e413-49ba-ac3b-0b2076309cab · outbound

This paper cites u chemann, S. , author Bannert, M. , author Dementieva, D. , author Fischer, F. , author Gasser, U. , author Groh, G. , author G \.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems u chemann, S. , author Bannert, M. , author Dementieva, D. , author Fischer, F. , author Gasser, U. , author Groh, G. , author G \

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:ed1f0adbef522d503f945fefeffb123d16a0d902b19dc0451b8f40ea63f5792c

Observation 352bf568-0d82-4517-9844-ca07da039b0d · outbound

This paper cites , author Brunskill, E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Brunskill, E

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:d9b7b152defcd45bf69f2ee84d5c7756d6bd4434cb0ae8797c3aceebb9f42c57

Observation 2457f59b-54cf-4da2-9b07-4d3adaffa2f1 · outbound

This paper cites , author Uesato, J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Uesato, J

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a666949f4b81d04f6108ecafb46325a5f293a83440b4ca05620fa9036d006d9d

Observation 003bc79c-5e21-4224-8501-fd85569a91b2 · outbound

This paper cites , author Fletcher, J.D.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Fletcher, J.D

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:586b5b44680bff71199323ac86087fcca6fdad87ec0947785d16d8662fdd4963

Observation 4de42525-53e3-432a-87c6-93bd7eca7914 · outbound

This paper cites A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:9a9cacab0cce491a5eb56b66f380cdd1215b467e684f26d3f8aebf26fb31bd1c

Observation 6f694e5f-4080-4fd4-a56e-e8d684ac58a4 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Scalable agent alignment via reward modeling: a research direction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3c1252c066dbe516fd0dbb51512b41fc52958cdc7ed51d7f53eae9b8b648ce66

Observation f635b969-de24-45ad-813d-20ad0b299933 · outbound

This paper cites , year 2013.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2013

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:0fd099ac5c59c7b5078062f853eca691b477d6e74e972bd1c6a0fa304346441f

Observation cddea6ca-a817-402e-afd9-6a7ab18346db · outbound

This paper cites , author Harpstead, E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Harpstead, E

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3bde22a4c31a41173ef0c3a87993f638aaa8e9f707b3a2813d1adbf6cbd5eb44

Observation e3e65764-15c2-4520-bc12-ffc14349bc60 · outbound

This paper cites , author Liu, Y.E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Liu, Y.E

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3319980470115e41c20d368865e734f3c91d4b03acc70bad432ebc66c940eadf

Observation 5e9e9439-824f-4283-9413-b771e78885e0 · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Categorizing Variants of Goodhart's Law

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:213beeb87a07113c0a2bb2ae134c1a070255888fc0844636fb0053f9a5e86ff8

Observation 62581fe4-3073-4dd9-9a3b-97a28ea022c0 · outbound

This paper cites , author Kochmar, E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Kochmar, E

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:1ee9c7c1ae09b85eff95f733c760e820f4b3f358d2416708d07cbb925c4627f2

Observation a30b6491-4a57-470c-91ec-61d3a82b3023 · outbound

This paper cites , author Reuel, A.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Reuel, A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:4418dbdd60a1ad780c6986c7b30f84dd17e8e1b9e559c8a43c12825f8de91e27

Observation 2284b373-70b8-4eb9-bd8c-303e917ad204 · outbound

This paper cites , et al., year 2025.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , et al., year 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:f73d6d1988b70b590e8ba23d2ef50d5553e7257dc9dea91af288002d0ae7dd45

Observation 9f3feb5e-d990-443f-a6d9-79580d8ad5c6 · outbound

This paper cites , year 1990.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1990

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:8bd574b89e9a280e04256a8c3ffc95de4da7a65a245e449f3140d307ea39d931

Observation 872a174c-c94a-411d-937a-a6311285c6c1 · outbound

This paper cites , author Wu, J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Wu, J

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:7c1bc885b9724703221f57787f73e35eee6a7d821b868a8e41490850bd06deb9

Observation 38194d39-a07f-4490-9939-457cc00af421 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:92b66a2f703ddf6f108d7822334f6043edf17e9ab0fe141d879ee1261dd44d1b

Observation fe2aa1b6-14b1-455c-a9c3-87b94b054ad6 · outbound

This paper cites , author Jones, E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Jones, E

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:1e2ec049531653a3959d33fc3fbf215ca39a95e6d2b1ea9726e038beda9f39a8

Observation 9be1de51-d088-4b2d-8a2d-7e3dc2291e4f · outbound

This paper cites , author Cen, H.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Cen, H

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:230b384b63e21a27ea804770724147a14d925061dbbe5cfb1352b8c65359cf86

Observation 3a6f4e8b-fae2-4748-9bd2-26aed9b8fb84 · outbound

This paper cites , author Bassen, J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Bassen, J

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:9e3a06e11079df6d12a95a461c2f552b0499baa77148af284759b7b385def094

Observation 85e89902-b83e-40b7-87e0-86b0a5b358d3 · outbound

This paper cites , author Brunskill, E.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Brunskill, E

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:84082c5622a56d50a6f40c284bcdbe992845e74c629e39a2678d129d7bde850f

Observation 59f7449d-b563-4ec3-86cf-cffb7c9d3c12 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:c0584df13bb4d17fcb431cf15d45237dc5a00dcc198f9d1787107ee33e604718

Observation 05d13109-53bf-4484-baf3-52f9ebfbfdb2 · outbound

This paper cites , author Vamplew, P.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Vamplew, P

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:497f5175b245873b3b34f8683140779419e53f9f0881d923252c7fbc38a8369d

Observation 66c14b41-d1e8-49f2-8d95-b931dbd07a07 · outbound

This paper cites , year 2019.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2019

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:161dab8c53845861f29882bf6101a85d749cc2540c7f9bd62e3d7f9117368a6b

Observation d355fd72-03a0-4c78-ad06-28b6948d2780 · outbound

This paper cites , author Abdelshiheed, M.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Abdelshiheed, M

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3b0b1cb2714b5454d5ce8b9c00d3af03cd16789a8063f14fd25820fd49f8a3e8

Observation 4758e1c1-25f0-4db3-b908-e124141b6a39 · outbound

This paper cites , author Maniktala, M.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Maniktala, M

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:d70ba9d77e2e142f11652bff900d3d39d296a3bf51b37e460c8f3b08ee3034bb

Observation c47e2ce0-dd1e-4c4f-a98d-9d6dde1602c7 · outbound

This paper cites , et al., year 2025.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , et al., year 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:6500d8fbe38f0f2d3f1d59ae9f78c0d9f7075c22d9c059fae86495efc011d11b

Observation 924c8457-f61d-475e-a6bf-a20dc431ea7e · outbound

This paper cites , author Howe, N.H.R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Howe, N.H.R

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:c625953568ae49ca6bfba2e5383f4af9fb16afabdf8210e14ce6d30cf18c3a49

Observation 430f68f8-bab7-48b4-a99f-1c083fe61d6c · outbound

This paper cites , author Cooper, H.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Cooper, H

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:e5e47211546210bd348acaf6461f6fc29d3ebe3f34a51b03865cad06ea901282

Observation d02d38a7-6bda-46b9-af6d-d25059dd0a97 · outbound

This paper cites , author Barto, A.G.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Barto, A.G

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:7b075750fe421f43af99a6af1b211b4687a73d77dde605dd186debe41e33cae5

Observation 7632cf0b-af06-44db-b7e0-dad4569fd3fd · outbound

This paper cites , author Mankowitz, D.J.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Mankowitz, D.J

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:9f9b0311f3b199a51d4464446ae2f8fa011e9ece87bfab07e285211e89ad4f1d

Observation 79dea177-55d9-4601-aa0e-5868b5c6b910 · outbound

This paper cites , year 2006.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2006

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:73ac5037a2df15278a00aecdcc0fe6950cd53ed8e80db882aa528e4c05cb9c44

Observation 4e595aa0-915a-4158-8b8c-81a109a4d310 · outbound

This paper cites , year 2011.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2011

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:3517ca024cceb2dd5f133febfa045f8173d2434d51b0e18d2e47b46b9984d9f1

Observation d84c7e89-7b18-4ff8-9b75-0f86a3fa6165 · outbound

This paper cites , year 1978.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1978

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:0bff37605607bd25ea1e2282408db0edbd633014955d330f23dd0903f5ce0073

Observation 4756ffd4-44e7-426b-8f88-dd05be76e795 · outbound

This paper cites , author Sui, Y.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Sui, Y

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:1817cb9dcc858254bcc597c58408e38de8a30e91162d29a056cd225634993584

Observation dc6c0429-e7e7-4647-ae4b-84ee450a3a12 · outbound

This paper cites , year 1997.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 1997

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:cf1701ee60d0edf056f344dd3b5647599aff13eec2fcc86720ef7958d6806a30

Observation 37bc3a80-1658-432d-bb81-86d5f2ca0016 · outbound

This paper cites , year 2024.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:ab9a67c4a2e793712a31b0f5f16b2a87ec8d7073c1e8eb978db2f0b6a97c7f99

Observation b3b7f8e2-f38b-4985-b543-1c8ef20a4166 · outbound

This paper cites , year 2008.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , year 2008

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:5d016e0d656370f2240ea3cdf9b510b12a7440ebc7f42de6d2c691abba87af20

Observation a9aeb0de-af7a-4d2c-a090-ed9fc881a92c · outbound

This paper cites , author Sha, L.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Sha, L

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:dec8c7a2fc487156c99938fc760c522079e5dbbe53df6ce45a9a4d664f2a993b

Observation d4e15f4a-b28e-4199-85d8-30ccc8e0fc49 · outbound

This paper cites , author Koedinger, K.R.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Koedinger, K.R

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:a366ee940afdee1ed71bf71a80b2ae0e84857eb9c081e93798075693fd957d71

Observation c856f506-e62b-4dc0-9b01-71dd6eca9ef9 · outbound

This paper cites , author Azizsoltani, H.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Azizsoltani, H

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:292458d53be757e60881a3d04d5e6e68ec2298e7ca2866621794b312d9d2a1ab

Observation c66d9341-a015-4b1f-becf-283dd6153289 · outbound

This paper cites , author Hadfield-Menell, D.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems , author Hadfield-Menell, D

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:e7658288f5356b88f65bd3c0354d739f1c7c70b099f9349c6050b814547435f0

Observation 40bb3623-1326-41ec-81d1-d4f32bf8a991 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Fine-Tuning Language Models from Human Preferences

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:206e7819977183d73edfd6de14157a54227d4b79c0569fd03a840914a6b7df56

Pith citing papers

No inbound Pith citation observations are available.