Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2407.19594.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.003907Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation ca078fb9-faad-404f-8986-e3a52e6b7170 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 254
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52402e35-8588-418a-9cd3-f34ca807c882 · inbound
LIMO: Less is More for Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84ed587f-2345-4098-8b28-592c52a1dbb8 · inbound
Reward Reasoning Model Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60556ac-df75-4b36-bca0-634d49709598 · inbound
Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01fedc5-c6b4-4bc9-9208-1279ac3466e3 · inbound
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c91b95a-4b68-441d-9198-df6418d1d8a5 · inbound
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66676d5-37b4-4900-9a0c-5c695b23be66 · inbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922c3bed-baff-4497-b9b8-297001d79be1 · inbound
Unlocking Recursive Thinking of LLMs: Alignment via Refinement Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f95238b0-b2d5-466b-9e29-bdafb150f4d8 · inbound
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fef196a-87d2-4433-a6a7-37af59368ef7 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420db723-0bf7-4c3b-83a8-1b5f407f5672 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95be38fa-aece-4e88-8109-e4c96cdb5f47 · inbound
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c95bc744-5fe4-4afd-8009-e645732914b9 · inbound
Enterprise Large Language Model Evaluation Benchmark Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baeaa4fb-5bd9-4b9f-84da-61c5d5ffab84 · inbound
Bridging Offline and Online Reinforcement Learning for LLMs Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd9e083-b3fe-4b12-94e5-8e64088cc42f · inbound
Revisiting Active Learning under (Human) Label Variation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf96810-b462-4a32-874d-ae31e45beba5 · inbound
SGPO: Self-Generated Preference Optimization based on Self-Improver Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ec67a3-c7ed-4663-9d70-ceb9de936158 · inbound
Multilingual Self-Taught Faithfulness Evaluators Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71b2bf67-f471-482f-ba00-5427082fa2ee · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ab6dcc-ad00-4bb4-8283-79b3df2fef12 · inbound
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388d9b6e-b63c-45f9-90f4-387e4da831d4 · inbound
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03b422e5-c77a-4319-b3bd-13d999609f2d · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcf5096a-b107-46f8-ab63-0b466d348d59 · inbound
Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8da6dbac-c112-4210-853b-3ca4df8d9d87 · inbound
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3efc4b0a-cf2d-4b52-9170-587accd10d67 · inbound
Trust Region On-Policy Distillation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f98913a-87b1-4de3-8e53-0b135e21af01 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 264
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 103da097-1211-4c59-961f-d31d121ebf1d · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 218
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99a6daa6-1d24-418e-b7ff-d2b1e72ebef6 · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 217
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.