Pith. sign in

Paper Citation Record · LEDGER

Cheap Reward Hacking Detection

As of 16 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2606.08893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.08893 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:21:10.240399Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b98d42d3-42e0-4dc5-b116-2cbe51c41ca6 · outbound

This paper cites Concrete Problems in AI Safety.

Cheap Reward Hacking Detection Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.187103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:2b0b8706344f63ebb2d30076644372ebf2039c91372fbcbc898e65bcccec5fe4

Observation f6c44add-2e91-4fe2-bed6-c673b284f936 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Cheap Reward Hacking Detection The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.183546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:b66e786a20ecbe16e7aba1b958116267bd80411d9ae425d95775938adc616241

Observation 75304bcd-3933-46db-a0bb-907d47025a69 · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.

Cheap Reward Hacking Detection Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:1ec523595597341590e8b3f612030ff47b6782693170fda7d44ebfa07010b122

Observation f4c636e4-01cc-494a-bd20-a951750d9980 · outbound

This paper cites Bisimulation metrics for continuous Markov decision processes.SIAM Journal on Computing, 40(6):1662–1714, 2011.

Cheap Reward Hacking Detection Bisimulation metrics for continuous Markov decision processes.SIAM Journal on Computing, 40(6):1662–1714, 2011

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:0daca90c5727bcf0fb16576bdad47af4cc184a236698bc06a2b36e9d708934ca

Observation 0cd2434e-3a9e-4f52-bce5-0ab395094a7d · outbound

This paper cites Density-based clustering based on hierarchical density estimates.

Cheap Reward Hacking Detection Density-based clustering based on hierarchical density estimates

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:cc1d963523912041f1279a0383804147e622c18369b5a936f8b37f47a5d72411

Observation df62f231-477d-40c4-8ffe-f23000abefd9 · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 30, 2017.

Cheap Reward Hacking Detection Attention is all you need.Advances in Neural Information Processing Systems, 30, 2017

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:ce086b72135435bc4e4a0d5c4d030a84e8ae7030aa36605d64989fb736a06f59

Observation 5a308ce9-e430-4ec5-9104-14a2540bed16 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Cheap Reward Hacking Detection UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.178304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:0e9feb5faf23ca195f3ae3a7419bc2e2e9280c0879aa31da2e54a1fef5206f72

Observation 311741cf-968a-4359-826b-21266dbd64a1 · outbound

This paper cites SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization.

Cheap Reward Hacking Detection SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:4e250461fa9cacb33023f5ba28f82658ebc4ecd169bc65e159303f4fc375a545

Observation 6548c4b7-cd22-4814-bec6-1667cfdc27e0 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

Cheap Reward Hacking Detection FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.181042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:a84df826575d27ffb31dccb57192f0070fa247382c6f34a65a8d68e3e71b98dd

Observation 7e5d8476-de37-433c-a6cd-cfa907876a1f · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Cheap Reward Hacking Detection A simple framework for contrastive learning of visual representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:62d1bed3276855ceecc66e8f933206ba40cba2a8597f717eaa175718744119c7

Observation 2a2af24e-cfda-4ac2-b784-81dff15a4a67 · outbound

This paper cites Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35: 9460–9471, 2022.

Cheap Reward Hacking Detection Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35: 9460–9471, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:4e8b29d373a50c073fcd1284c5bf610796835f6bc72ddd2e4b038df7cdbc85ba

Observation 038e4a77-d0d3-4382-910e-df59295d96a7 · outbound

This paper cites Specification gaming: The flip side of AI ingenuity.DeepMind Blog, 2020.

Cheap Reward Hacking Detection Specification gaming: The flip side of AI ingenuity.DeepMind Blog, 2020

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:d31431a38fe3713297745a5c33a7f78971ac614b90d7ed27386f7d41b02bd839

Observation 4b5f902d-6ae9-49ea-9ed0-625d913bfc38 · outbound

This paper cites Programming as theory building.Microprocessing and Microprogramming, 15(5): 253–261, 1985.

Cheap Reward Hacking Detection Programming as theory building.Microprocessing and Microprogramming, 15(5): 253–261, 1985

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:bd61352f297c33e0c843bc1dcc224e4aa38223f6230fa2453989c8ebdd00dd52

Observation b89209f8-608c-4713-859f-45358207db56 · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Cheap Reward Hacking Detection Categorizing Variants of Goodhart's Law

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.175446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:14b6b8a92bc2387201ae8088be48078df6a199532160fd5236e970f0d084e02d

Observation 443cbe0c-506b-4760-8ced-f122b6570d5e · outbound

This paper cites Towards understanding sycophancy in language models.

Cheap Reward Hacking Detection Towards understanding sycophancy in language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:667b1167e008a502bfdcbc50146262d7d4ec1262d176d361f0fca80a667fc415

Observation dc3b7ddc-d012-4c3e-89f0-82089c723d66 · outbound

This paper cites Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories.

Cheap Reward Hacking Detection Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.170778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:7092e15958553996ff57f120b9ea8b3812d00f86c186c4705f984543ddd98608

Observation 639a4d60-33d8-451a-a12a-3787297e5964 · outbound

This paper cites MICo: Improved representations via sampling-based state similarity for Markov decision processes.

Cheap Reward Hacking Detection MICo: Improved representations via sampling-based state similarity for Markov decision processes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:a82316758806ed3058b270394bc0716d38ee4a277c9c4460f10165855aa7fbe0

Observation 1a837e36-267a-4c60-b80e-1eb9aa0e2e2f · outbound

This paper cites Sliced and Radon Wasserstein barycenters of measures.Journal of Mathematical Imaging and Vision, 51(1): 22–45, 2015.

Cheap Reward Hacking Detection Sliced and Radon Wasserstein barycenters of measures.Journal of Mathematical Imaging and Vision, 51(1): 22–45, 2015

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T17:21:10.240399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T17:21:10.240399Z digest=sha256:2dabe9c5935c91e1c58d5a385a8a4edf2cffb02ff68ca4898de2b95cae646e73

Pith citing papers

No inbound Pith citation observations are available.