Pith. sign in

Paper Citation Record · LEDGER

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2506.23626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23626 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:13.953628Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f13b1593-9979-4bf9-af18-bb1a1b7a84c6 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.571140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.179049Z digest=sha256:f4d67b10796a7041d855fb0dc87a5b5938cb159f51b8ce709f9ab1e35cbc6fce

Observation b3449306-4e61-4056-9538-1e58989b0715 · outbound

This paper cites Instructions.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:16.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.299088Z digest=sha256:9851f461c3995e42b40baebfe5d91993e6e84db4be547b929502c366db814994

Observation 3a6e5395-af65-4535-bcb8-925621b60483 · outbound

This paper cites • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.765168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.519280Z digest=sha256:fd717f17e6999ab79e0afe40f743a79afd0593f286628b0a2b89c62730aec2da

Observation 02429b1e-248e-4587-8aeb-b749060398cd · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:15.533487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.606616Z digest=sha256:fcda852198a451c43b3482b384367bd730e7f77e5aeaab04d3813476d5038163

Observation 8cd012bf-5014-47db-a29b-7e198e3d1ee5 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.179890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.394406Z digest=sha256:8fd0de101c8b0584df80cbbe5eb373215e0f58d17a5ad82855f888eb1dcdd7e3

Observation ea1367b6-e3a6-4383-8564-c12a9ee351f8 · outbound

This paper cites • These parameters control the agent’s learning and behavior.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • These parameters control the agent’s learning and behavior

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.991458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.462055Z digest=sha256:1bbb6ec9b030242fe7905b7a390f3b7812d5c6bb8b24e2d1705d31dfe229f0ed

Observation 3f643437-5a2b-43bd-8a36-28e141c0720b · outbound

This paper cites For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.336253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.718492Z digest=sha256:e960e4ea7b9762737430a0010c101cad6046e2d7d7e5919addf2cc68e37cf4d3

Observation d9417170-972f-4942-b930-3069e9e49353 · outbound

This paper cites • The format of the outputmust be identicalto the .txt file.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • The format of the outputmust be identicalto the .txt file

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:41:15.161476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.993083Z digest=sha256:193966d96d336e8d2790872608f15e5cf1969ed696c7f1577255f4063facc73a

Observation 950e8afd-d779-430d-98f7-bdd64b67600c · outbound

This paper cites in the least amount of time steps.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games in the least amount of time steps

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.974683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:12.040776Z digest=sha256:2703ac86c3305cbc4a2dca5d9bc74c66fcff2c725cc53969e507636556c6425e

Observation 9e02ff12-8e3c-4c01-a2f2-edbacc489886 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:14.763880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:13.661770Z digest=sha256:d581e1ccd283761d4ba82996515636d8ade45b48c3165345ebcea150a6921e24

Observation 59b48f51-89e1-4354-b052-6931d15109b6 · outbound

This paper cites Try to be creative with the solution, ie.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Try to be creative with the solution, ie

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.577985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:13.861797Z digest=sha256:f26b8053224ac0bbf290e2f75deb1541743a02e6042f453da94aa310198168b2

Observation 2f60979a-0dca-4fcd-9b60-c098dc3bde44 · outbound

This paper cites Your output should be the updated reward function file only.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Your output should be the updated reward function file only

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:13.953628Z digest=sha256:c216f680005920615fc2720f606f86cb7568e03c74582c4edc124e3d3d0ef5f7

Observation c3808800-ed00-446a-8938-fb0675132a38 · outbound

This paper cites Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan

Reference 2023

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T21:41:14.245278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:41:10.113675Z digest=sha256:ba327b349e8f22a940d17ac347645bb89186cadc40c258bd73889308cfd2763c

Observation 1bd7b80f-7613-4173-ba96-9553b70eddc3 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:10.068318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:10.068318Z digest=sha256:df36e8b29a0bed663d1ab31d0c9e033aa812caf47bf78974a948ff756515c69f

Pith citing papers

No inbound Pith citation observations are available.