Pith. sign in

Paper Citation Record · LEDGER

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2502.08922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08922 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:18:40.394477Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:46.031769Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T17:08:01.250732Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb139baf-7580-4c82-8155-cc86f5db0d28 · outbound

This paper cites Meta-rewarding language models: Self-improving alignment with LLM -as-a-meta-judge.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Meta-rewarding language models: Self-improving alignment with LLM -as-a-meta-judge

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.599511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.508691Z digest=sha256:335f07bfccafa588fa9df23d11aeb88aa48a13f81abdd7ff892b4d4f91481858

Observation e724814b-ff4b-49ce-bd8d-18587919ea51 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.557512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.557512Z digest=sha256:7e5985718c6e251bde68f2aa2bc9a83dd1a27f5d341d017498c0cb5fc6e3477c

Observation 9e892fca-77c4-4940-9743-4061a4d54fa9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.615070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.615070Z digest=sha256:f36c5aa40d272fb3c230ea276459bbeab88c23cf16fbf5b7ef21424c36481304

Observation 88a07b6c-c796-4124-84bb-16bfdcd50f94 · outbound

This paper cites E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.585312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.664489Z digest=sha256:b5c5f785ac892c157d70fd2c6a43f7d31267bb0333706cafcfc955ba19438635

Observation 7b5a1aec-238c-4517-8551-cdacf8658b67 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.720611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.720611Z digest=sha256:73c15c4e05c7bbd5278417aa3883e76812c508b37ab51a4601aa988425c4ffcb

Observation 8421daa7-8b6a-424e-8e6b-1658d8da8220 · outbound

This paper cites an unresolved cited work.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.782053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.782053Z digest=sha256:e6b298791bcceb292649c56c1018c84a07002e560e9bd34159b64851099b54b3

Observation f0305702-5474-4972-9d5e-3d4203c67e02 · outbound

This paper cites an unresolved cited work.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.785911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.785911Z digest=sha256:8cdfed5f06557b8c6233cb39a2dab6f9b0c7a8746ae2fe149a86441ea6947364

Observation 6e509c3c-332b-4979-9a0d-763c8681ae98 · outbound

This paper cites H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.564457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.789152Z digest=sha256:fa9b89cf61ca427043c95ac6553707c6cf828beef0531f89d8e12a3cd0206441

Observation 2c615e10-6313-4219-b04a-95e09e6e31af · outbound

This paper cites Discovering latent knowledge in language models without supervision.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Discovering latent knowledge in language models without supervision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.554612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.792319Z digest=sha256:3a885e0a507a3309882a12dc7aa63d9f3810cfe29e4f08847464b7ec411b4b57

Observation d46c78dd-3243-433c-9bfb-621697db17cf · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.795597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.795597Z digest=sha256:33aa4228a302b4f0d3f1b4e77f6df5db096ce354be9e52c226d29e9a4af61d69

Observation 1163456b-9200-41ef-9695-08a2fd6e6d56 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.803268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.803268Z digest=sha256:11d9dab2973f8338c39b7dcf0e9cc435e8894c42c953395e7c43667a8fbee3ce

Observation c1249323-4353-49ba-bdd0-b7cbd6e58b82 · outbound

This paper cites an unresolved cited work.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:18:41.432377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.806277Z digest=sha256:bfbcaa87d160a03c763f224e8518ae7ab26b87038bfb041aa692a28abdf60ab0

Observation c7fafe0c-56c6-451f-b9cd-cf2034c80ae6 · outbound

This paper cites Training verifiers to solve math word problems, 2021.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training verifiers to solve math word problems, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.809418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.809418Z digest=sha256:9d6b5ad8b520c00fbcce958566c62d500af503bc3910ab25d1987511ed489386

Observation b6f2cf57-9420-4492-867a-7d51ea2b01a1 · outbound

This paper cites MetaRM: Shifted Distributions Alignment via Meta-Learning.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models MetaRM: Shifted Distributions Alignment via Meta-Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:18:40.888353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.812498Z digest=sha256:e2f6734c2b2be1792d6f8e83b40169f30de437177443966a65cec80a0b5cb29e

Observation a435c9e9-de7c-40ac-ac1c-dbf3d4c2c9f1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.815907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.815907Z digest=sha256:ec0948fd5e2c9a10c2d84beeee5dd4c60b73eb409cc809d8224afeb14122b563

Observation 0faf2016-1197-42bf-9496-d41fc3b97446 · outbound

This paper cites and Bengio, Y.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models and Bengio, Y

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.350955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.819576Z digest=sha256:fba60a44e0788c8243e5e08afb046207684dcff2519e96e65b4361a6714db01a

Observation 81736b11-9a44-490a-b19a-5d4fe0d9ae56 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.822951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.822951Z digest=sha256:ae79cd9d824a73f0287eb970f472873f17d02d67ea2dbf6b60879b1069058c81

Observation d4e7d5dd-168e-4055-9f3a-aca257409d27 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Survey on LLM-as-a-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.831125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.831125Z digest=sha256:d2aa1409d4fbf85a1715254e591f48819329092a40ef87f2e6ab1309b194d1f8

Observation a82e82e3-79b1-4e5c-8908-5d982ded6360 · outbound

This paper cites Measuring massive multitask language understanding.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Measuring massive multitask language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.340024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:39.833827Z digest=sha256:78133e125e50e2a541516c80ea135f1f3e784705c5c71841aa67bfff05c7c933

Observation c61ef4f5-603d-4696-bfcf-d7a32ee4e1da · outbound

This paper cites Large Language Models Can Self-Improve.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Large Language Models Can Self-Improve

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.836718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.836718Z digest=sha256:ea046c0e381ec1b918c426d341504e82a393372f4b4c14b6293420617708f665

Observation 89abcfe5-4980-4830-a70f-a54d6375978a · outbound

This paper cites Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.880512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.880512Z digest=sha256:e0a732eb5c9176ac2c3677b8093f1a17da353a384ca94584ad4df7f149074e1e

Observation fa350cd6-20a9-4804-aa81-23c4452eae54 · outbound

This paper cites Mistral 7B.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:39.944964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:39.944964Z digest=sha256:89c26ed1c00ac911297181de4093ca24c97c494cff91e49dc983b9d49d1e43b8

Observation 4af70f34-2ea6-4aab-8e0f-5eda5533175c · outbound

This paper cites A survey of reinforcement learning from human feedback, 2024.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A survey of reinforcement learning from human feedback, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.001004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.001004Z digest=sha256:133108b51b8d5bdd7442df161082bd9b42b7292814df1e18aa39bb048e075140

Observation 967174d5-25f9-43bf-b316-f897652d2b9c · outbound

This paper cites o pf, A., Kilcher, Y., von R \.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models o pf, A., Kilcher, Y., von R \

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.328414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.054020Z digest=sha256:dea1dbec6af38a8a7a30a41acea97dbd0c1977ed016862b9fe351bfab0aa8b3d

Observation c9581455-2476-41a8-b899-d9ccae36168b · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.095737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.095737Z digest=sha256:c1306259176d59c85f2c30f6b9b0c457be6ddc38dc280c2507c498626310c9cb

Observation 0cfd3253-0bdc-4c0a-accf-e50797e06f7a · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.157736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.157736Z digest=sha256:86b7bbec89f2253ec2f9efbd487cf4c6e2e84a16c274cfbaaa0bba8274e8d264

Observation 295294a0-e936-4b54-972d-5442e5f58595 · outbound

This paper cites and Hutter, F.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models and Hutter, F

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.177163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.177163Z digest=sha256:f398b3f35fb4337156d6c92a33695989c429afd8e27410051c5d8e13d2ab8d78

Observation e09f68dd-25f7-442b-9268-0b9981e5ea09 · outbound

This paper cites BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.180688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.180688Z digest=sha256:46040327599b286b649bf36a8a79976edde34652dfb03ad8ca065525a8f83639

Observation c5a02edd-f8ea-4c74-9aca-1e8d39ffe7e1 · outbound

This paper cites Introducing ChatGPT.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Introducing ChatGPT

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.312585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.184253Z digest=sha256:499ad73698bd552f5f392bb3233a865a18097bdc54f7442ebef5414789bf8afc

Observation 20395fd6-42f6-4f65-821c-fdbb441104fd · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.187371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.187371Z digest=sha256:227b5621bd94ff78cd518b7c31ade6ffd0caa63b30af40a9e3edcc704350d3e7

Observation 7de1d6b8-31cd-43dd-8c9e-2f83d7250299 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.193451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.193451Z digest=sha256:e5264418978896f4fc8b3c2fb6fb6368b5155dd3819616479f1881d5f371b3a8

Observation 70fe3b7b-f63f-46f8-ab25-0070b9bc4792 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Disentangling length from quality in direct preference optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.196466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.196466Z digest=sha256:e3ea19f139e590b7f84f55bf9a09034f8a4e8bc4b917a0f5c1f7a3bf8b05bc05

Observation 2cc8016c-02b9-41da-a29a-b5e2a5431e7a · outbound

This paper cites D., Ermon, S., and Finn, C.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models D., Ermon, S., and Finn, C

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.200063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.200063Z digest=sha256:9fc5c2803c6a1ddbfd4df6ddc7646af1ac21f81559afb4b8c040e7638d784ebf

Observation 903318f1-9c08-4bdf-8239-36e367371e71 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.204091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.204091Z digest=sha256:467c5ce4719646de52944da8b1a43ef0e035a45f007b2877cffe6e50ba97a228

Observation 17deff01-442c-461d-9e67-b0f3dc9307cc · outbound

This paper cites Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.288557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.207167Z digest=sha256:598936f90c124649119be0e8a3bb0927f3bc92f1d991f0e6b39d80dab2f2d4f6

Observation 7f7c52d3-b3d6-435b-9a49-90fde801f122 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.210360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.210360Z digest=sha256:440e6a40a89040fa414fa854bc2ec2717615f79d7c07f940f2db9bf42834615d

Observation 25f07d7e-6a95-4b3b-ac2c-c14c9fdce18b · outbound

This paper cites an unresolved cited work.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.214834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.214834Z digest=sha256:74121d04d3ba84650be72309fce5f3e0c98d21e7a518b60522fbf4fb36ece1cd

Observation 530e25d9-8427-48e0-b08e-4145d69cc392 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Llama: Open and efficient foundation language models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.218076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.218076Z digest=sha256:a296893955e7fa7904ec0a034990bece36f5d4fce52c3fcb49791695880042cf

Observation 88a9ef1a-fdf6-43df-9a9a-48aa06193a4e · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Aligning Large Language Models with Human: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.221309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.221309Z digest=sha256:81c1283ae0da15d481ce6f1b95b18270aef5831fa3133d920484676099098ee4

Observation 7164f128-cdf4-44ef-91c6-d7204eaa7eed · outbound

This paper cites CREAM: Consistency Regularized Self-Rewarding Language Models.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models CREAM: Consistency Regularized Self-Rewarding Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.224406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.224406Z digest=sha256:befb4e56a1b9e1cc1a3cb43b8cefb4415948125d76d6a225d6d1c2d593f31fd4

Observation 45ba3edd-7f76-45ac-9694-d4a798e03727 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.227467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.227467Z digest=sha256:a13ef6ea3899e80fe97f0ccb152d7f407a94436855f35b47fe2e94fab857f093

Observation 64166ff1-2b6d-4123-b172-19f7c6def6a6 · outbound

This paper cites Unsupervised data augmentation for consistency training.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Unsupervised data augmentation for consistency training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.266269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.266269Z digest=sha256:bb77d2c38908a8167592449f52b463aa32bc68de5ab3d2912d145a1bbe5251b1

Observation 598702dc-9791-4f54-a428-300b670adccc · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.316040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.316040Z digest=sha256:ab2679990df1978a91c7619db87e9e8ac657c82ac1079ae39f0a71d7e7be8335

Observation b18fa00a-7ce2-4f66-b358-f64b253c7e37 · outbound

This paper cites Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.257658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.357141Z digest=sha256:5597a4ae1d0228d64374f76401ef5697f5770517681e08ab6585687574ef658f

Observation 99baab46-c283-4db2-bea8-20b2703cc36e · outbound

This paper cites Consistency regularization for cross-lingual fine-tuning.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Consistency regularization for cross-lingual fine-tuning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.210130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.378764Z digest=sha256:ae4bbe23af24ce21c4bfd17572521d7b51204337c5547025a316c203e3f61af6

Observation e33d1246-26e6-4c54-b606-ade210b624df · outbound

This paper cites E., and Stoica, I.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models E., and Stoica, I

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:41.083420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.382175Z digest=sha256:7be8f66d1316cbcf72847f3119d32af7b33d63b0194c2b4a80a340b6f5a09c5f

Observation 176818cb-b9d0-4c68-8670-6542311d105e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.385881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.385881Z digest=sha256:1f1f8d930bf9030d9d81959c0be58a93f8c193981a629e8fa96fdd045151818f

Observation 4cb36be4-555d-4d63-b195-7f4f991998b5 · outbound

This paper cites Lima: Less is more for alignment, 2023.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models Lima: Less is more for alignment, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:18:40.949771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T23:18:40.390304Z digest=sha256:8edbdb86c0c17d7506ebc2aaa053b4b5d5319c49acb0e995c42d46a3e909fb2b

Observation c8fedeb3-353b-4ef0-a942-1cd7d86361a2 · outbound

This paper cites write newline.

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:18:40.394477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:18:40.394477Z digest=sha256:2ac13b862d7204dcd6791f11fce673389a88a3d5dacb2a2d91d5f37b6e281531

Pith citing papers

Observation 7a523748-4ad1-4f59-b1a4-e6c6a8b30bc4 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.205165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:0937907aa2751cb6d37cdf62dc319918e97aa3a7edd1c33c98ee3643078caa0f

Observation 97b1f075-2d3f-4a02-bc76-f2266f53f3a7 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:46.031769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:46.031769Z digest=sha256:13beddcd2e56e612189ace508df3cf152ccec218e1f18efc804cae52b624b104

Observation 18db4afd-80b6-4da9-b0f4-c8e5b09750d9 · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.746071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.746071Z digest=sha256:76738c99ff6923eeb9b9d94a7956a436a123b86b1eda3b3302960553611b2b22

Observation 9b43603d-a844-4c6f-9943-96b546c31597 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.253192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:dc116b48776213a8a82788e1c0742029393b8a7527137c53db66f77187219d78