Pith. sign in

Paper Citation Record · LEDGER

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2506.23626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23626 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:13.953628Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f13b1593-9979-4bf9-af18-bb1a1b7a84c6 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.571140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.179049Z digest=sha256:70d36079dc48a601f4b79cac3fe9144ebbe6a45b8f13a4887a6dc947f4cbce49

Observation b3449306-4e61-4056-9538-1e58989b0715 · outbound

This paper cites Instructions.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:16.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.299088Z digest=sha256:3f03767469ab98cf62dfea8d650e5248c584a15c968d166836b54fa1308ae33b

Observation 3a6e5395-af65-4535-bcb8-925621b60483 · outbound

This paper cites • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.765168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.519280Z digest=sha256:e16352c581299f734b4adf09ad589251fccb8f785ce08e57b6214ca58fb96f6b

Observation 02429b1e-248e-4587-8aeb-b749060398cd · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:15.533487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.606616Z digest=sha256:94b49e4e4c2ddd0dc6ce3bb83fadb78871fad3e9498e193c23841542a2d582d1

Observation 8cd012bf-5014-47db-a29b-7e198e3d1ee5 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.179890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.394406Z digest=sha256:104d26ec35983154e9630ec2bce6c31ccdd2e80a98cd347830fbc84be6305872

Observation ea1367b6-e3a6-4383-8564-c12a9ee351f8 · outbound

This paper cites • These parameters control the agent’s learning and behavior.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • These parameters control the agent’s learning and behavior

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.991458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.462055Z digest=sha256:73790bdc673ac2bdfdce5f26bcf702b811c977a3724d5cd3d4d38507db888083

Observation 3f643437-5a2b-43bd-8a36-28e141c0720b · outbound

This paper cites For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.336253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.718492Z digest=sha256:a14cd211ec0e17da5b9b799accafa54869adaf6dfaf1fe82e8962bb51c62e1a6

Observation d9417170-972f-4942-b930-3069e9e49353 · outbound

This paper cites • The format of the outputmust be identicalto the .txt file.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • The format of the outputmust be identicalto the .txt file

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:41:15.161476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.993083Z digest=sha256:6921a9bb1c647e7acadc7e3c5fbcb4156f5ab5ecce244c4d30989d39b7eaffaf

Observation 950e8afd-d779-430d-98f7-bdd64b67600c · outbound

This paper cites in the least amount of time steps.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games in the least amount of time steps

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.974683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:12.040776Z digest=sha256:069db13bf3d9b168bc45b51cdddb945a2388206a0140bf41989f5885d6cb9129

Observation 9e02ff12-8e3c-4c01-a2f2-edbacc489886 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:14.763880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:13.661770Z digest=sha256:cedb84f26a24aa6d3fe0b4b2a6522837304f7e9aab58db3c6d8f6df4c745ef54

Observation 59b48f51-89e1-4354-b052-6931d15109b6 · outbound

This paper cites Try to be creative with the solution, ie.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Try to be creative with the solution, ie

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.577985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:13.861797Z digest=sha256:8a07532b520df9ab840fec498e2f4d805217dd52b968553c96f3901e7ea984f5

Observation 2f60979a-0dca-4fcd-9b60-c098dc3bde44 · outbound

This paper cites Your output should be the updated reward function file only.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Your output should be the updated reward function file only

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:13.953628Z digest=sha256:fb07344a5d317021df1fa80881ea95d3c05c9a4d59d02b58078d95773037b918

Observation c3808800-ed00-446a-8938-fb0675132a38 · outbound

This paper cites Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan

Reference 2023

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T21:41:14.245278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:41:10.113675Z digest=sha256:f7f7684ade9357950dbcd17b8ab2b4bcbdee9cac3907fe5a2c51740fbcec2f9a

Observation 1bd7b80f-7613-4173-ba96-9553b70eddc3 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:10.068318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:10.068318Z digest=sha256:0d77b45141b756514fcfe5ee81145d49f620283e5ef9577dc345e51e611c7359

Pith citing papers

No inbound Pith citation observations are available.