Pith. sign in

Paper Citation Record · LEDGER

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10209 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound

This paper cites Concrete Problems in AI Safety.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.686636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.686636Z digest=sha256:a17b7bf213ae781103fb1e9362ae7dff4e9a058c3ffe5ed09e23f1be5af1f372

Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound

This paper cites Measuring political bias in Claude.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.210935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.693517Z digest=sha256:6727ab03ebe4513103c551ba842bf303cc01b2850b9b2b0a203e2c39c9c6c7bd

Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.698106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.698106Z digest=sha256:c881e9b018e00e2f8f838c171d4d3cb30f2fd949dfd7f18d7a63b7aa86a7ed1b

Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.702774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.702774Z digest=sha256:fb1f4e0be50c0e270b24ae553f299131afa6751d20de725b004e75a814e1decd

Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound

This paper cites Boerner, Stephen Deems, Thomas R.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.708085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.708085Z digest=sha256:2b111b9f05948089a7bcbe9bdcee585250e60410840c4d8d455815ea558b7a8d

Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.712832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.712832Z digest=sha256:8dd18bb31be7244432d0b07082cd2901612557182bbde2f6207a9b5dda359edd

Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.721800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.721800Z digest=sha256:412866583406bfb94b9a3329c20f541b4fc9a97a100e226fb0f77f1d3b778dc0

Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.186919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.729878Z digest=sha256:c640c758ac3565f0fa37cf489916f931217238ad314fac813d028627250f4a83

Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound

This paper cites Deep reinforcement learning from human preferences.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.734216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.734216Z digest=sha256:efb4be81c9386c4b8de22063a9bf6990cbc78496f713f2e291de7136ca395dc8

Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.739636Z digest=sha256:377e2513cb0a40202159ada2f9c20e0d9f64009ffae7b51d54a7b471665b3ba1

Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.746751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.746751Z digest=sha256:aa04b53183090a98d320483cfa6b5cd1b746ff638c6c667ef412048fc7c32bff

Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound

This paper cites On the Relationship between Truth and Political Bias in Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.751481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.751481Z digest=sha256:5562cec44560dfcd5edb2cbe486994cfa77ce090ce1f38bbedc0328a5be320b1

Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.756004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.756004Z digest=sha256:82fd6b24ea877a055bef426bfd1e3089ef71a78649f428c76fb4516031e8aa66

Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.760327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.760327Z digest=sha256:860c0a566d54ea3f3293120365df507eac1bbf647220cf9de2c0e272dda626e1

Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.767334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.767334Z digest=sha256:05eb0b0b48b1aedee872c590cdac681de3f4cb1fdc6c3d5df2980eca218b5c50

Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound

This paper cites SGD on Neural Networks Learns Functions of Increasing Complexity.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.772617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.772617Z digest=sha256:22a7bd531b624071fc461f8302a0d455f4800abf212021a38305fcb0508cd8fb

Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL , 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.777339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.777339Z digest=sha256:99b4eb7d6c45450192819ed05f04c52d75725147a7faabfe7c8066dbc61cdf23

Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.782461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.782461Z digest=sha256:9d2f01d665a4af0c5ca0eee94415a82d1be447ea98ab091f068fa226acdff1e7

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:c0cf2ceca8ed5f37bc6fbe9e22d3f09d667fd71d38e58d0b9e9f9394f26c57de

Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.791545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.791545Z digest=sha256:ce08fc7c64e7fac2b6059207ac702b575be55a198479f21b0326fc451ee7c4c4

Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.796035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.796035Z digest=sha256:eb91c907cd7577bcb16b6269b6fe847500836bd8f653a2e1f483f664b210cbb8

Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.800654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.800654Z digest=sha256:c227588ee385639103e6cfe18890989d302043d2db3949fade21ac5943a660d0

Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.805283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.805283Z digest=sha256:b5e644388c3ac9f49d1d8f7fd9055684584a372d5e7f0b4955d9774f55e4ae39

Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound

This paper cites On the Spectral Bias of Neural Networks.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.809491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.809491Z digest=sha256:406866d1bea3b4bd2bd588b4783cd7d37136655be1e8b839b785cdcc37e93fd9

Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound

This paper cites an unresolved cited work.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:01.163686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.814145Z digest=sha256:bf3b5d46a0f57139d9c911bb2a778734c577ed156c258f9a10c2af4a4945e389

Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.822017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.822017Z digest=sha256:8791c14d8ff0ddff71b097859c5908d8678e8de757d0d94426cbdeec7ca8256f

Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound

This paper cites Defining and Characterizing Reward Hacking.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.828256Z digest=sha256:67c0110c55dca122ad113eacbac4e59cbdaddf36adfea49b7ee3d5c6fcfbeee1

Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound

This paper cites Learning to summarize from human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.845655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.845655Z digest=sha256:4c77e5fcb62831ead0aa7da77a890073982fb0fcc648c892e53239e23f86b22e

Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound

This paper cites Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.850544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.850544Z digest=sha256:5d01cd467bddbc1067b930fdfb6b81beb294288da22c0ae23293be6b6dff4d0b

Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound

This paper cites Deep learning generalizes because the parameter-function map is biased towards simple functions.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.855765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.855765Z digest=sha256:337c4cc3cfff4272bd9f3c5be69c22a50d77364c5e5adeab1b3662711ec26d5e

Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound

This paper cites Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.862922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.862922Z digest=sha256:262c9e094b9385a9af7ed08f665e5ca35269499adf7b93e558ea83972a6a1b65

Pith citing papers

No inbound Pith citation observations are available.