Pith. sign in

Paper Citation Record · LEDGER

Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2410.03742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03742 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:40.569414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T12:56:24.594609Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3545d381-49d0-4bca-a7a7-5e5819485c6f · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.846903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:e52668b60bca4391fe4f0ca0c7dcc5d6c0f5545f9d9c2e4aa634d36bddf4ac1c

Observation 65ed0f96-370a-4e77-85e7-397bfd35c7b2 · inbound

Reinforcement Learning Enhanced LLMs: A Survey cites this paper.

Reinforcement Learning Enhanced LLMs: A Survey Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.569414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.569414Z digest=sha256:831cca1f23aa3a09a71bf9d72614d65d17000c65791ba9cb0bc83c3a7de4b7c3

Observation fe590e6b-b6e5-4f90-a71c-5e6eb9fd070e · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.230084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.230084Z digest=sha256:f87074948faeaa8547a1394ef6b2820633b468b8c8a81c5d359189de53248cbb

Observation 0346ca06-8dc5-42b6-8bac-5b01371268fd · inbound

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO cites this paper.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.531765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.531765Z digest=sha256:d486c1ae70b07a341402b9731be42e1c32234c6e4c07d9f86bd045e43ffc17f1

Observation 18917754-a9e0-4cd7-9bb8-634b69e74c0a · inbound

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards cites this paper.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.607519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.607519Z digest=sha256:4994f59875061e4ea98bf3c394a034ac0bfdb49bee2edf86886a1eadf99e5370

Observation aa761de0-13a0-428c-a2a0-faf5ddd51be0 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.597062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:f0e86848b1bc8710b6f7fbdf54e95b41cc7f5a504878016fb54fea4608102c89

Observation 06b06874-1450-4894-9e15-1e324d0cc986 · inbound

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation cites this paper.

EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.159528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:59:40.497067Z digest=sha256:13f2e314d247e3d3c7ca56fa4fd5427b987854153330c21b63ea8efe76867916