Pith. sign in

Paper Citation Record · LEDGER

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling

As of 18 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2604.22981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.22981 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:05:00.013591Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:31.955883Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation be5091b0-a62b-49bc-ac7d-599812a5804e · outbound

This paper cites If r(x, y0..k) = E[r(x, y)|x, y0..k], then the average value ofr(x, y0..k) −r (x, y)should be close to 0.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling If r(x, y0..k) = E[r(x, y)|x, y0..k], then the average value ofr(x, y0..k) −r (x, y)should be close to 0

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.557246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:1c70423e4ceceea19492b68044390490e457d1bfd232ca8e21076b5812e4a54f

Observation 5e4a2f64-1e80-49fd-901a-8fdb925d2287 · outbound

This paper cites Sincey0..k is more informative thany0..k−1, we should expect to see the mean squared prediction error(r(x, y0..k)−r(x, y)) 2 decrease askincreases.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Sincey0..k is more informative thany0..k−1, we should expect to see the mean squared prediction error(r(x, y0..k)−r(x, y)) 2 decrease askincreases

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-05-26T13:42:24.593496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:258bf22f36243c17a7909f817547b0c4d9ff75c6dc58bb562180e4d12bea91be

Observation a9bce3cd-d509-47eb-93f0-8cdd50bd9241 · outbound

This paper cites What is the capital of Italy?.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling What is the capital of Italy?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.572252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:9da5b4de218652c98c5b7f26890dfb335e78af3db096505936c19772fda0866d

Observation cd944571-b660-4f56-9e35-3781291e1025 · outbound

This paper cites It is Rome.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling It is Rome

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.554097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:49bfcd14efd8248f49f8252a1ad5a2e5e493115e71ace20556f6c175e7ef09a8

Observation a7ecf277-a44b-4135-a292-35a059b43718 · outbound

This paper cites Rome, also known as The Eternal City.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, also known as The Eternal City

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.551205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:752de82c750d5042bbb673a7057e265d495e83376f80c57deab5f894f103237f

Observation 7a241ce9-a992-4319-9a90-e3e02e632abd · outbound

This paper cites Rome, GA.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, GA

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.563365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:c9967fff722841e6160cf32ab659a66fc23534f9a789d0d9dbb0cf29c316ca2a

Observation 5aaf6c01-d19c-41fd-8727-2f390cf29941 · outbound

This paper cites Rome, also known as The City of Love.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, also known as The City of Love

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.596565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:8f25fa9bae9e1d39ae4b8586ef6d1ae55b24134bf068e5078d3ea15ae002bb54

Observation 67466b84-f137-4545-a1ba-ee6e10a9c17f · outbound

This paper cites \n\n" separator) minus at the end of the previous step was 18 2 1 0 1 2 3 Reward Model Output.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling \n\n" separator) minus at the end of the previous step was 18 2 1 0 1 2 3 Reward Model Output

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.587513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:ae069bc7f895bb8865405d4efa1c47e5badbf8d13c12ae9a2649fe0acc09c683

Observation 8046382c-4f46-4d8c-95f2-8400448eebb5 · outbound

This paper cites Similar to the previous method, but sigmoid transformation was applied to the reward scores before subtraction.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Similar to the previous method, but sigmoid transformation was applied to the reward scores before subtraction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.566790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:36da0f9f62cb9f4cdb2dde3fa7376bafce826c4c55511c569b80ae71ba0534d4

Observation 7c574f0e-b9e3-45fa-8bdb-9036d509dcef · outbound

This paper cites Normalized Difference.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Normalized Difference

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.581651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:4529e79a71e0cf90db483e89f233b70a7ac62e908a945f651da81e9b12e43397

Observation 623d9013-7318-490e-8bff-f6d0e0811832 · outbound

This paper cites This was used to evaluate the checkpoints every 10 training steps to understand training progress.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling This was used to evaluate the checkpoints every 10 training steps to understand training progress

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.544749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:59d5886b131ff1b51b6c69902b9b2ffcca1ced8bf6d515bdca093e170cba6746

Observation d91248a3-0142-4d9c-874f-adc6005bf530 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.590360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:7c55916c913bba9290327b43cdca179e6580dc127e1b7238e7b75dd0392942c2

Observation 395e19b0-b562-4292-9e50-9b4647b677f3 · outbound

This paper cites The prompts for evaluation are also sourced fromDolci-Instruct-RL, but we only used the 10%validation subset with no overlap with the PPO training subset.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling The prompts for evaluation are also sourced fromDolci-Instruct-RL, but we only used the 10%validation subset with no overlap with the PPO training subset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.560193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:64f9a3830cc418b0d5fb82f8b8f8c2455e7122382ccb7d20ce2e6052520fa3c8

Observation 2d8b95e0-d29e-4f26-b95f-e4fd9a0ddac9 · outbound

This paper cites warm start.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling warm start

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.548230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:97a469dfc7072d9c64ca29270e47552e411b8ec8f6670a5f319d75bd4b7533da

Observation 69b1cbeb-0e57-4997-9d1d-c9a0b6609106 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.584526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:746cb9f7c966244a8a82b188c0b11e856959596769345fdc373148e54b589d62

Observation ad260881-b957-49cf-80ae-47b819cdeab2 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.578416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:1082866450cdce95b8ec6946813a5466f5674076a4f5336d34098f64957e98c1

Observation c56e4abe-6ae8-45ee-a89f-9ec30623b440 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.575416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:1f713675b1e488496892bbb0e77e5519716f07642eb501e00da057e2a401926e

Observation 807ae89b-8af9-40a1-82cb-664db535cdaf · outbound

This paper cites A" (Response A is better).

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling A" (Response A is better)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.569442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:946d29e5f6ed380496157a563b4fba9bceed118d8f9e73283a3b5a57ccf373ff

Pith citing papers

Observation 164682d0-2cef-4822-b803-1943d511c074 · inbound

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training cites this paper.

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:30:35.757355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:30:31.955883Z digest=sha256:b1e4aea8651aaa6f0e7afec8a432773ccb5f99e4a1b1e232cab220dd2c95dcc6