Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2409.20370.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.831378Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T12:26:31.581809Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 92f1ac4a-af66-4441-a1e6-c82f7e85207d · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d58cf3-19e4-4a9c-b0dd-b33b719e9368 · inbound
Multi-objective Deep Learning: Taxonomy and Survey of the State of the Art The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 199
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f7f0db-581a-470f-b3a7-d376c1862577 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75b57ebb-2c3e-4acd-9fe1-0523b5f07d79 · inbound
Reinforcement Learning from User Feedback The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d12446-f629-4000-8a3e-681432b58a87 · inbound
Boosting LLM Reasoning via Spontaneous Self-Correction The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90fd3dc-1854-450a-b2cc-e1efdc81e7d4 · inbound
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baadd476-aa13-4853-a946-7bf01cdf7f50 · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f80efe-77d1-4ad1-a195-0bc56ddfee7f · inbound
Debate2Create: Robot Co-design via Multi-Agent LLM Debate The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2ef0f4-bce6-4a99-9403-4dad04cc9755 · inbound
Token-Level LLM Collaboration via FusionRoute The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0daa22a3-e4fe-4bc8-8bd5-3fd64f227851 · inbound
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 204
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1707d8-0451-4278-be2d-ea5f5f65f148 · inbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.