Pith. sign in

Paper Citation Record · LEDGER

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2407.19594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.19594 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.003907Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca078fb9-faad-404f-8986-e3a52e6b7170 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.709019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:4796d12334540fa836ce03479216eaec5b2d50a06985115853f11dda2cf5562c

Observation 52402e35-8588-418a-9cd3-f34ca807c882 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.211488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:244f34b6893fce6ecf3347b34605d0a2773f9cdef19c5c82e574b73acab85f5b

Observation 84ed587f-2345-4098-8b28-592c52a1dbb8 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.003907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.003907Z digest=sha256:76b42c1e0ffc9ac9e444e18bf601b51cd3d1b594e7cf4584ce661322e3ff6164

Observation f60556ac-df75-4b36-bca0-634d49709598 · inbound

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers cites this paper.

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:20.983397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:20.983397Z digest=sha256:cbd3d7037816b1eaf85ed4c7b98706c0ca5dcc97b4e496844d45c1eb6b35717c

Observation a01fedc5-c6b4-4bc9-9208-1279ac3466e3 · inbound

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks cites this paper.

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:51.799718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:51.799718Z digest=sha256:b7bab5f13c2dcccf516e49f23cf1ae9b79aa31a23e8fb7f44d0df0a69cf9d043

Observation 5c91b95a-4b68-441d-9198-df6418d1d8a5 · inbound

LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification cites this paper.

LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:48:04.284134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:48:04.284134Z digest=sha256:e94e3f33e235f7c8fe17d54e8aa5a2aa4664ef2cd5626b05cd162ec3a6f78037

Observation e66676d5-37b4-4900-9a0c-5c695b23be66 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.652545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.652545Z digest=sha256:1d69ff186f4ed783f4b5a4c5f4a66ebbdaa26f66be4e32a3076b590fc9b3dac4

Observation 922c3bed-baff-4497-b9b8-297001d79be1 · inbound

Unlocking Recursive Thinking of LLMs: Alignment via Refinement cites this paper.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.109745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.109745Z digest=sha256:874a4c1361918e6af8e4db474adb5ff990af3224174cb2250ffc9f7228f84573

Observation f95238b0-b2d5-466b-9e29-bdafb150f4d8 · inbound

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models cites this paper.

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:49.730548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:49.730548Z digest=sha256:4745f5d69cc96ec5dbceb2072df7f76b4068468b0ac6f502052c8ae1e05457cc

Observation 1fef196a-87d2-4433-a6a7-37af59368ef7 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.585513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.585513Z digest=sha256:d32d4c78da20e65b597e6510a83737ca530fa0f4c483517bd626fc9806fea7fc

Observation 420db723-0bf7-4c3b-83a8-1b5f407f5672 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:37.395043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:37.395043Z digest=sha256:8bcb693c3c28b9f72d3a3a29116f147979228c20fa86a072e5d4b535231da677

Observation 95be38fa-aece-4e88-8109-e4c96cdb5f47 · inbound

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation cites this paper.

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:47.916012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:47.916012Z digest=sha256:644bbef1b08884ecc94ec97b18817fa7861d5d528e9b49d4b2fd2a098dae4d1f

Observation c95bc744-5fe4-4afd-8009-e645732914b9 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:33.549019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:33.549019Z digest=sha256:bbf5845daf28d863a0a3d88eb2e27e48c29aae36c336af1266a6f47e24365b4d

Observation baeaa4fb-5bd9-4b9f-84da-61c5d5ffab84 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.310079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.310079Z digest=sha256:21d5472bb7cb3ccff2f5e402a4b2dc5a54d1aab8cf8a30f2982f8f20fc608e5e

Observation bfd9e083-b3fe-4b12-94e5-8e64088cc42f · inbound

Revisiting Active Learning under (Human) Label Variation cites this paper.

Revisiting Active Learning under (Human) Label Variation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.062777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:54.062777Z digest=sha256:3abfe00a7b8604377390df598c4b0b4f1b5bdf13e4b3a0d32c9501d6bdf216b7

Observation 5cf96810-b462-4a32-874d-ae31e45beba5 · inbound

SGPO: Self-Generated Preference Optimization based on Self-Improver cites this paper.

SGPO: Self-Generated Preference Optimization based on Self-Improver Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:49:12.374832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:49:12.374832Z digest=sha256:485ee3a3f08491c7e12a92cf411b20b6a984eeb45a622bbcdb620d85177ed30c

Observation c7ec67a3-c7ed-4663-9d70-ceb9de936158 · inbound

Multilingual Self-Taught Faithfulness Evaluators cites this paper.

Multilingual Self-Taught Faithfulness Evaluators Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:25.167291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:22:25.167291Z digest=sha256:ffdd25fc58e1f13357dd83340b442e907d9d708ef2c9285077f1a2ab7779be84

Observation 71b2bf67-f471-482f-ba00-5427082fa2ee · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.698152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.698152Z digest=sha256:515bf8a97ee7e77000e813f0c942ceb64677eeaf92dcb5a56f8eb0382d4a1ec1

Observation c4ab6dcc-ad00-4bb4-8283-79b3df2fef12 · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:25.666027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:25.666027Z digest=sha256:d0a07c10e764513cfa260b578e40d8c35c52d17512af371b91205e0d1c51a23b

Observation 388d9b6e-b63c-45f9-90f4-387e4da831d4 · inbound

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process cites this paper.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.687864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:021d63c4c8f512a7d3ec51d43674b917596c244839eaa28e0f6cda2d3f821318

Observation 03b422e5-c77a-4319-b3bd-13d999609f2d · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:08:01.288616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:b65473ca5e488f99997f1559504786ffb3a88063a19caadf5b8d18d087714128

Observation fcf5096a-b107-46f8-ab63-0b466d348d59 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.407696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:68acfc5db84feae09a268091c3344ba076dcf44b9b2afb415c3785629f274f4b

Observation 8da6dbac-c112-4210-853b-3ca4df8d9d87 · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.972571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:734bc7376e2e1781ed4cd94da29e1654770283fe61a32cf409c2b748bd483401

Observation 3efc4b0a-cf2d-4b52-9170-587accd10d67 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.295277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:b8c8f43a737352a5437040659f710f88cb30fa5c103c01b19c9fafdc78e3e617

Observation 2f98913a-87b1-4de3-8e53-0b135e21af01 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 264

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.551756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:a6e2700d1b7f7a168eddf3b118d2de4420c555d21dc02761794835e24f20e20f

Observation 103da097-1211-4c59-961f-d31d121ebf1d · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 218

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.789963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:072ebb4ea6409c7ed792a0b72da3c209da558c85cc24f6b335c6c2ef260402ce

Observation 99a6daa6-1d24-418e-b7ff-d2b1e72ebef6 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 217

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.731218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:0df879bf6deac458be57cedeea399f361856b3087958734937bac006c6bd9c11