Pith. sign in

Paper Citation Record · LEDGER

Intra-Trajectory Consistency for Reward Modeling

As of 21 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.09096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09096 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:38.960214Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4132a0e4-49f0-476a-91db-a8c37c4f63d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Intra-Trajectory Consistency for Reward Modeling Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.831652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.802950Z digest=sha256:f1be3d2217580bd39363ad5d0e4300614ac5660a0baae39e82f921e19f421c11

Observation bec86bc4-369b-4086-b50c-f16b257b9303 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Safe RLHF : Safe reinforcement learning from human feedback

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.778388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.901900Z digest=sha256:2d3dbaffe746813e654920724eb59bcee2e7c34abfc6ae8233733cab8530e219

Observation 2fcd8772-3748-4012-983b-cf3ccc1b3c56 · outbound

This paper cites Model alignment as prospect theoretic optimization.

Intra-Trajectory Consistency for Reward Modeling Model alignment as prospect theoretic optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.715947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.972354Z digest=sha256:adb46b4242e65b7181220cf749c2eab6ef58e7a46878f41a3bfca0d7ab2dd0d3

Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · outbound

This paper cites A Survey of Direct Preference Optimization.

Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:34.053645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:34.053645Z digest=sha256:d6126c32d1d80a1569de662f2a5a46577991d9027c804ddbbadcfd15ca42576f

Observation 6fde3e4a-7d27-4dd2-8899-cc5c84a54d8f · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Intra-Trajectory Consistency for Reward Modeling Generative verifiers: Reward modeling as next-token prediction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.658480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.125960Z digest=sha256:6e249f1e01420f9c722da3031b0e0840cc026ecabb3a1c03b8307aa53f425950

Observation b9fe2fbb-cf39-4bd0-ab3a-daf565f93a65 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for LLM reasoning.

Intra-Trajectory Consistency for Reward Modeling Rewarding progress: Scaling automated process verifiers for LLM reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.586236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.198454Z digest=sha256:65bdaf53060998ad8d4358ae6eb290d454584f60899e27adfa9bd91aa8a53ad9

Observation 6661a49b-0f37-47d9-b026-caa023ae1740 · outbound

This paper cites Scaling laws for reward model overoptimization.

Intra-Trajectory Consistency for Reward Modeling Scaling laws for reward model overoptimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.504214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.320396Z digest=sha256:b0826599a0f91b9b301f415bdc5120aa05aec162ce90e8c9b751af514111ca48

Observation 726cb910-5925-4f6c-b4a9-48241819474f · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for LLM s.

Intra-Trajectory Consistency for Reward Modeling Regularizing hidden states enables learning generalizable reward model for LLM s

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.412624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.412805Z digest=sha256:325e562e39e6e79e02d0f26d588d637663c3868a3f1901e47eb57bdfd9d3b8cf

Observation 53f1914c-cd13-4cac-aaaa-4e94bc21976d · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.303562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.526511Z digest=sha256:d3c7dbd31aef920e387ea55c17b9e27e1b7d55b216a1a1dabf9ea2ac69c51cba

Observation edff0ebb-c86c-41a1-a74a-25c15ce9b1c3 · outbound

This paper cites Warm: On the benefits of weight averaged reward models.

Intra-Trajectory Consistency for Reward Modeling Warm: On the benefits of weight averaged reward models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.175665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.624566Z digest=sha256:4a59cf96988471bc881724b5aee0d2f945f91a96d76fca4238c3a8ef6cb3b351

Observation 32e99435-33a6-4fd5-a7d6-30f7ff031b90 · outbound

This paper cites The trickle-down impact of reward inconsistency on RLHF.

Intra-Trajectory Consistency for Reward Modeling The trickle-down impact of reward inconsistency on RLHF

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.062844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.718080Z digest=sha256:064ee3d1a5f7313500677f461334314da0016efb00499c2e3507190502c050a7

Observation 51ecde41-b6f3-43e0-b887-94d27783c4df · outbound

This paper cites Rrm: Robust reward model training mitigates reward hacking.

Intra-Trajectory Consistency for Reward Modeling Rrm: Robust reward model training mitigates reward hacking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.952080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.818224Z digest=sha256:895c8bba41616d28473fc8b46f0a48b1d9d9c1d942ef81df9be4a091d196f69e

Observation bbc1ae30-c7e8-4a0c-860c-cac1e1c9f90a · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Intra-Trajectory Consistency for Reward Modeling Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.827940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.911868Z digest=sha256:2af2c4fa91ad4822f0377d517d716479c80235bef609e15d52c9b62657bd944e

Observation 2b9af422-9ecf-4211-89ba-30c3ec0a3010 · outbound

This paper cites Odin: Disentangled reward mitigates hacking in rlhf.

Intra-Trajectory Consistency for Reward Modeling Odin: Disentangled reward mitigates hacking in rlhf

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.701760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.020744Z digest=sha256:067a29bf11180050ff4e9e437e1ceb719826389a35f970edf247d9da9c955b05

Observation fdc4bee4-9ae1-446c-aa16-55969f355381 · outbound

This paper cites Improving discriminative capability of reward models in rlhf using contrastive learning.

Intra-Trajectory Consistency for Reward Modeling Improving discriminative capability of reward models in rlhf using contrastive learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.590224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.128852Z digest=sha256:7d1266f5f2a3e132882a6fa7b4851cd880930a2ad20583b30388531d4bac49b0

Observation 49f91bbc-1be1-4b00-9db4-52a9d972270c · outbound

This paper cites Rethinking reward modeling in preference-based large language model alignment.

Intra-Trajectory Consistency for Reward Modeling Rethinking reward modeling in preference-based large language model alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.454000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.197601Z digest=sha256:ee171b8d8253490e13ce9b6bc3d71dee5a8aa2d6510013e21ba318a67b0da241

Observation 9049e081-f6f4-4c1f-a5fa-e908588fa405 · outbound

This paper cites Let's verify step by step.

Intra-Trajectory Consistency for Reward Modeling Let's verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:35.330672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:35.330672Z digest=sha256:834ea7a73cb1301af98b4bc097d61a95519a31a565f0f5c938aa2bd64d455779

Observation 5e927d83-6394-4273-ac20-844e64b50103 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Intra-Trajectory Consistency for Reward Modeling Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.340960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.442708Z digest=sha256:1f1fcadd66c482768faf23bfde5089b6393473298a0a37372761312231badd3b

Observation 7a7ca000-670d-417c-8a6a-7a8db72ec35f · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.

Intra-Trajectory Consistency for Reward Modeling Rest-mcts*: Llm self-training via process reward guided tree search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.200619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.554133Z digest=sha256:55a4ffabc8703991792dd874502bba80f60ae9352ef5ff5ca0292045881486b4

Observation 0e4c4813-eb08-4f65-bbb2-cbf80ae1a160 · outbound

This paper cites Gemma: Open models based on gemini research and technology.

Intra-Trajectory Consistency for Reward Modeling Gemma: Open models based on gemini research and technology

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.040192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.617671Z digest=sha256:521a285ac776a73c0b380bb3aaefb6634713d4e75fe7adf32bf06d620f3006f1

Observation 0b1e1919-0da9-432f-98da-33781cc30cd4 · outbound

This paper cites The llama 3 herd of models.

Intra-Trajectory Consistency for Reward Modeling The llama 3 herd of models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.927236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.758591Z digest=sha256:b2a67666635eb38e0d13573e09e347c519a6ef281290eee94dbe40958215a95b

Observation 379d687f-f425-49c7-bf25-29e493da0680 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.851006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.867529Z digest=sha256:f1fa08974caf4d68ae307a26d4644cf41d25437c276961d545358e288b4d4f19

Observation e15458fe-de22-4138-89a3-13ee36986d9f · outbound

This paper cites an unresolved cited work.

Intra-Trajectory Consistency for Reward Modeling Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:09:43.590225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.924570Z digest=sha256:28bd48965c2ae599c074c961404337310e6812fc73783351ce5cc7b22a88abee

Observation 2c87082c-53e2-4bcf-8941-90a39fabb102 · outbound

This paper cites Solving math word problems with process-and outcome-based feedback.

Intra-Trajectory Consistency for Reward Modeling Solving math word problems with process-and outcome-based feedback

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.247491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.016024Z digest=sha256:fdb6acefc950ab169c6c6846cfeedc088c925709752277597ef8b814e704ae75

Observation b6be155d-90e6-4dfb-8466-2f519790f09e · outbound

This paper cites Token-level direct preference optimization.

Intra-Trajectory Consistency for Reward Modeling Token-level direct preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.898366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.115444Z digest=sha256:e88e7a3bf39c44bc9077824e12d62f021791be80057e99798b19a8c3dfd8eeac

Observation 835f56c4-ee2f-4d5e-b96b-004c8aae0fb5 · outbound

This paper cites Process reinforcement through implicit rewards.

Intra-Trajectory Consistency for Reward Modeling Process reinforcement through implicit rewards

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.614422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.195715Z digest=sha256:7f7bea18c6862d45d5cde287a5065965b1f32ae223c69037a3a4d6761b67618f

Observation 521f59ea-afed-41d1-be65-117b7ccc2e9c · outbound

This paper cites Gen ARM : Reward guided generation with autoregressive reward model for test-time alignment.

Intra-Trajectory Consistency for Reward Modeling Gen ARM : Reward guided generation with autoregressive reward model for test-time alignment

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.503125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.318045Z digest=sha256:6b5a741cda6f3444668f75c09dae1879ca967384729791192f1e212633979cff

Observation 1be149c5-8847-42b9-a1aa-f1115b8585de · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.

Intra-Trajectory Consistency for Reward Modeling Fine-grained human feedback gives better rewards for language model training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.368824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.407602Z digest=sha256:e1e4b5ec2019048a6c8fd42e8a96ccebbd67cfe49620cc669baf3a613a8e5249

Observation 196c5c63-3519-4ae4-95de-a01b34f6c74a · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Intra-Trajectory Consistency for Reward Modeling Safety alignment should be made more than just a few tokens deep

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:36.488861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:36.488861Z digest=sha256:c19d39922b557c58a7541666f816610076c3b1bf46250c29569e100db59b8ee4

Observation 6237dda3-1d2e-4d98-b747-9e20ee8a911f · outbound

This paper cites Process reward model with q-value rankings.

Intra-Trajectory Consistency for Reward Modeling Process reward model with q-value rankings

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.234540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.596630Z digest=sha256:3c3981078c1bfa7862301147e8cf46e3bb2832f5e5edbebe82b76da7703645a9

Observation ce204ed8-5904-45fb-9405-ced160567cf9 · outbound

This paper cites Coda: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding.

Intra-Trajectory Consistency for Reward Modeling Coda: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.119546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.685789Z digest=sha256:f3338abd6cbee11b6d4e17dc00fff572f88bc91d47c5744039bcc4569da351f2

Observation 98bdd41b-ec82-485a-943d-8e407496663e · outbound

This paper cites A survey on data augmentation for text classification.

Intra-Trajectory Consistency for Reward Modeling A survey on data augmentation for text classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.971159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.777800Z digest=sha256:839df056dddab13a2e231826916fbfa949cc250c7f1c8fa7bf2cfbcfa4aeb0aa

Observation 44606918-c76c-4a6c-bd26-703a45d2b9f7 · outbound

This paper cites Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel.

Intra-Trajectory Consistency for Reward Modeling Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.850902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.873930Z digest=sha256:0e7011909cb1fccd8a1da6c4e2a44a843fe7fc51a133aa7e858410e8bb92b199

Observation 8e096c90-a31e-446b-8466-1a2f3dbd2c6d · outbound

This paper cites Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling.

Intra-Trajectory Consistency for Reward Modeling Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.732648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.958032Z digest=sha256:42791f8d9545fb04ee2d2ba23172ebf62594001dc862d4f39236710967ffbfd4

Observation 62cc8540-6c22-42c8-bf54-bd30eaa62a12 · outbound

This paper cites Qwen2 technical report.

Intra-Trajectory Consistency for Reward Modeling Qwen2 technical report

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.612274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.074903Z digest=sha256:beaf0d3740aedc844dc2c6913be8d92f745f37a75527de9639e00febaa71d87a

Observation 5b5989df-c9f7-487c-abac-0ae2dad00254 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Intra-Trajectory Consistency for Reward Modeling Measuring mathematical problem solving with the math dataset

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.503487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.183892Z digest=sha256:e1c02877ca7710a21d3785b71f3485c4bb69e145340a96d004fa70cd269191ad

Observation 73391db7-3f72-47da-8f0c-1a3638a177ca · outbound

This paper cites Free process rewards without process labels.

Intra-Trajectory Consistency for Reward Modeling Free process rewards without process labels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.385834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.281339Z digest=sha256:6b2b3167aa99e076103b8e41c9f1e894568a98eb2b62dc7da0c84ae1579eb39c

Observation 3bc09aff-6ed7-4505-8345-735472625b6b · outbound

This paper cites Mistral 7B.

Intra-Trajectory Consistency for Reward Modeling Mistral 7B

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:37.400183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:37.400183Z digest=sha256:3cf0b986a1b03a2ba822f1c1656237393fcb15c96d8c839d11536c69d610aa76

Observation f2697b11-fb0a-4340-8c89-b747571cf43d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Intra-Trajectory Consistency for Reward Modeling Direct preference optimization: Your language model is secretly a reward model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.255820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.508705Z digest=sha256:3c2ea51dd684796d24ee4b93c5536e543c443aebd01221cb950fa9489d897fa9

Observation 10235350-5bec-4bbc-a2e7-56b114fc2575 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

Intra-Trajectory Consistency for Reward Modeling Rlhf workflow: From reward modeling to online rlhf

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.128081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.639203Z digest=sha256:6751021458da4df3ac798320cb11b02b615659dbf64509c866e45675854825e4

Observation 6e5a6512-f8d7-4337-b74d-7e6e175ebe34 · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling.

Intra-Trajectory Consistency for Reward Modeling Rewardbench: Evaluating reward models for language modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.978309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.747142Z digest=sha256:a99e0a5be92a03bc3da66017995b0ab19714c71a55ce4eb42a64e3a1a55621c0

Observation e05e6164-2d74-45cd-a9a0-df3fa4aefba7 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Intra-Trajectory Consistency for Reward Modeling Llama 2: Open foundation and fine-tuned chat models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.799148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.856360Z digest=sha256:982827c77853caeae8fe226a966f45ec7fbefb72846efa9bafce850f02af62fa

Observation 43d67609-7c95-4db5-b8a2-0c528715a0b9 · outbound

This paper cites Secrets of rlhf in large language models part ii: Reward modeling.

Intra-Trajectory Consistency for Reward Modeling Secrets of rlhf in large language models part ii: Reward modeling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.614967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.001579Z digest=sha256:9d00cf735bd7494221a2783c3374a304ac8794c3d80a2cde4ef5272a48acc18f

Observation f81486a0-a0ec-4cec-83a8-fc12d2ec4624 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.426846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.111936Z digest=sha256:3ade769ed8a77e19d2bf14f08a1be97eb6ae2b229c338bec7e24f97a001517e8

Observation 3c2aebab-b8cb-4e42-b102-a5ccbcf6015e · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Intra-Trajectory Consistency for Reward Modeling Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.259515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.239973Z digest=sha256:4e3de89d08b8c34a9f3abfd01324a07502c2e1283ed5c95f201e7d9ce26db652

Observation 30b3becf-60e8-4b2f-b91a-08d6e6de7e84 · outbound

This paper cites Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.082421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.357506Z digest=sha256:8c4d6bad7a6d3805989b841ab3bc173ca377de4d63b0744b3e5e09340e06a854

Observation d8e979ef-ed7f-4cfb-ac10-6836fa9f4ea1 · outbound

This paper cites Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant.

Intra-Trajectory Consistency for Reward Modeling Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.901062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.466365Z digest=sha256:6f508fb792d7859809b7f99c6a4d4732a75bd6aa49a968b2f7003aa8f4ebdadb

Observation 7d9606f7-8284-4b33-8833-7550696c0906 · outbound

This paper cites Proximal policy optimization algorithms.

Intra-Trajectory Consistency for Reward Modeling Proximal policy optimization algorithms

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.751532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.545848Z digest=sha256:1d82940ae3e16a909f0efb79b84b4911f2579debb6b765453f5b2d126f0ad48d

Observation 49fbd15f-3dca-4368-b712-494501cc534c · outbound

This paper cites Understanding the learning dynamics of alignment with human feedback.

Intra-Trajectory Consistency for Reward Modeling Understanding the learning dynamics of alignment with human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.555383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.671903Z digest=sha256:ddb919f8c150c99c027bfd48faeacf8490eace026ebe94108e0971ba85fb118e

Observation f245d885-99b7-4bee-be2f-03ae603fcd91 · outbound

This paper cites Improve mathematical reasoning in language models by automated process supervision.

Intra-Trajectory Consistency for Reward Modeling Improve mathematical reasoning in language models by automated process supervision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.388219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.760953Z digest=sha256:f6a1ad110850a622c00f751cd00fd7b84290da8f0675ad728d4bf18f8e956262

Observation dab70296-bd7d-493d-b28f-2bcddad4ce2f · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Intra-Trajectory Consistency for Reward Modeling Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.187419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.854123Z digest=sha256:5b8779674b89d6311d6e900a5a683dd57776ffd0b6b461db6c3a301cb3bf8356

Observation 1327d8a5-e129-4bc2-a7e8-75d29996d2b4 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Intra-Trajectory Consistency for Reward Modeling Lora: Low-rank adaptation of large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:38.960214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:38.960214Z digest=sha256:149d33231829b9722828224132a95d7c94b587cb9f6dcf7e20d575112c6e9428

Pith citing papers

No inbound Pith citation observations are available.