Pith. sign in

Paper Citation Record · LEDGER

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10209 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound

This paper cites Concrete Problems in AI Safety.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.686636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.686636Z digest=sha256:97ad6d9827351e36ef5c6d8017b4c782b20c3c68c7ea904fa5ad4c28026b7c3c

Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound

This paper cites Measuring political bias in Claude.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.210935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.693517Z digest=sha256:78384a6e042cdf4e9bb898be92d88d40b236d3259521002b4df9e9722387dfc0

Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.698106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.698106Z digest=sha256:21ff5e8478d8c715e5e8a8974e818e4257b91eea73b63b9afcd82008151eb2a2

Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.702774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.702774Z digest=sha256:904daaa435a0289c3a9242ff42ce0972112d90b564d10887bdb27a2dad729f5b

Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound

This paper cites Boerner, Stephen Deems, Thomas R.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.708085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.708085Z digest=sha256:54283b2f24fae445eb87e3e43ebd3ed9e19b0af95644801fb07c4da14be9cc91

Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.712832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.712832Z digest=sha256:9472ce45ac66d6b0c364932fbd574639b290e2b6258128810e2a81993ef1076d

Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.721800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.721800Z digest=sha256:20098c44946bc1de582dfab80ed320b55d1703502002df9f8dd1f6ac7ca7647b

Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.186919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.729878Z digest=sha256:c67d6627c6ee9a3337a9fc4b76ad7079d6b55f285fe7400bc65e7fc193e6e06c

Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound

This paper cites Deep reinforcement learning from human preferences.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.734216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.734216Z digest=sha256:7bf6816e13f3b46a9770d03be05a041f453283ca84668f54eb815488360b3d75

Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.739636Z digest=sha256:ebd927800d64875bbf0c348dee7349d28fdef293b3c0bd583e0ea1b2aa278241

Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.746751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.746751Z digest=sha256:91837e4290e18fa939457f15e66196af9cd2bd4b71d1735461b9db9e113bf10b

Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound

This paper cites On the Relationship between Truth and Political Bias in Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.751481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.751481Z digest=sha256:a061e9ac0b4c594cc50f49b3954bef5cad31f2cf28326df6388a5f6f4bd154af

Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.756004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.756004Z digest=sha256:833deebb17e966d8e78b8de9b93333dd3b357021d184bc4c1d91193465e7320b

Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.760327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.760327Z digest=sha256:e118983b81fb6a310f5281ba13125a090d7b0381a9a056fea85fc1bf5d8a714b

Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.767334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.767334Z digest=sha256:952450648d71be02045afcc43cadcac8b84a5b0b7e730f8b6f415d8a933a3f64

Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound

This paper cites SGD on Neural Networks Learns Functions of Increasing Complexity.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.772617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.772617Z digest=sha256:cdffc6f3ec30c7177f779e70f0095c0163c48f5380324e11336ddf1e473258bc

Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL , 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.777339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.777339Z digest=sha256:5cc7a12270dcd663e8818fa7ec48dd6dc2e9abb0910c81c6b987384321e889b5

Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.782461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.782461Z digest=sha256:339915b55cb8e4bf69a778108d0045611e99629abd6bd2a9c004501daa384cf4

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:9c6389851091690a099f936b60e9ea688c0fafa796a3d649f9d4f61566af7b44

Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.791545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.791545Z digest=sha256:d0a615c02a7723e3692b4dd3a4ce614587ad48dc584bb3a09dbd049fcbb6a183

Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.796035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.796035Z digest=sha256:6979eb4e6740deca0999c7acf76f3d4db963276b0494e5db8d173ada450e5711

Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.800654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.800654Z digest=sha256:bef66a860069e4874e664ef28e2c74a599a4adc5a45e4fb7144f801b33d6700a

Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.805283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.805283Z digest=sha256:aa1d873ef88925b459a2e49c9c2d0751f05da00da1fe4d1ee336baf29bc65683

Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound

This paper cites On the Spectral Bias of Neural Networks.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.809491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.809491Z digest=sha256:083d5c2a145942b530b86bc90e2e1c180127cdfa0f8d808ca4cec092eca63898

Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound

This paper cites an unresolved cited work.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:01.163686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.814145Z digest=sha256:d9a9c3dcf8b9ca2bc4f30b63933b1a00974c209b6a5ddeb6181f5fc96657705f

Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.822017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.822017Z digest=sha256:347202a92c300fa0e130a0b57b7be0d6fdc427bcd575685af72d0c2d4a5e63bf

Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound

This paper cites Defining and Characterizing Reward Hacking.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.828256Z digest=sha256:fa415bfb5628b8d787be76403a2ae4e97197ae322d280ea246396bee6e754138

Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound

This paper cites Learning to summarize from human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.845655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.845655Z digest=sha256:a10ef815c8d53aac51db3014a2d075c2d8f7e76a5e5b46ffd4dba57f268fbe28

Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound

This paper cites Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.850544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.850544Z digest=sha256:d9ac352bd7510b0af2636b88da511d47ccbede170895ca72959da6ea8a27b15f

Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound

This paper cites Deep learning generalizes because the parameter-function map is biased towards simple functions.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.855765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.855765Z digest=sha256:491623c70f2a098c7bc159bf6c89b4a257e906bdd8e56817d107516d28bed367

Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound

This paper cites Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.862922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.862922Z digest=sha256:b04156f90a00f2a2008a56e4fdf2d35ab2ff8abc507c9e64b352496ddf5f9da7

Pith citing papers

No inbound Pith citation observations are available.