Pith. sign in

Paper Citation Record · LEDGER

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.08965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08965 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:59.969588Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03fe4b88-0c49-479c-a137-bec123b343a1 · outbound

This paper cites an unresolved cited work.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:04:00.872967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.629212Z digest=sha256:895600e00131ad77ce6740e810bc3e3e041a10697762d12e35c74911201a772d

Observation 76d43501-a5be-4718-a772-01626ecea597 · outbound

This paper cites strong signal.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO strong signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.712230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.701979Z digest=sha256:a95344ef4aa05c10cdd122c43508739ca7c2b8b57d7019fcd4b1f2ca7c154522

Observation 45c4b94b-5f19-4f9b-8836-fea49deded49 · outbound

This paper cites Weak Accept or Strong Reject vs.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Weak Accept or Strong Reject vs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.271775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.969588Z digest=sha256:5f8250a05182dc33497d558ea09ebb81a67230a4d9ac6eed14172fa4774efdf6

Observation e7bb3a19-2414-47b3-bf2d-6a8c61a5c1b6 · outbound

This paper cites Generative Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Generative Reward Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.316747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.316747Z digest=sha256:009e36e5b463c2638b547b596c9725f19b26ada744feceefd1563fb0d24984e7

Observation 0346ca06-8dc5-42b6-8bac-5b01371268fd · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.531765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.531765Z digest=sha256:1d836e6790f19eab1962358f9d810ad7794e1c338e9dd47e33344e96e0443e23

Observation 6921bfe3-fb3a-4a1c-9b50-f04875cc2178 · outbound

This paper cites The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling).

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.559616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.800755Z digest=sha256:bd4406df4d6fff6ee4b5019d3c3dd084b34a88233370c261059b1a69bcc4327d

Observation 0d90e109-0ec1-4a4d-a75e-8c517cb91d5b · outbound

This paper cites This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.426144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.861614Z digest=sha256:0ed5062a44bad5a90c833edfff5725f5a9e9624c5da15ea7122574637671b736

Observation 7e971c68-9613-45c3-98c3-6249acb37590 · outbound

This paper cites Proximal Policy Optimization Algorithms.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.448741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.448741Z digest=sha256:639a67ab15bc9843a3fa85720362cf355fd0ad8a0f061ef1b41a72e8590925ee

Observation 93599e44-d838-4697-8c2d-62358cf345b5 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training Deep Nets with Sublinear Memory Cost

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.265968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.265968Z digest=sha256:6c3008dffd3e42b23fb425b660bf4ad9752fdc322b09d6ef527fc867c61a4292

Observation 0a49f408-87ef-4be4-8efc-4928a420e7ba · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.109692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.109692Z digest=sha256:f9ff8256032e292641ef9cd26066d8e474b7eeeb9448001592e64850e0f5b4c6

Observation c48772e5-ded9-4861-8e38-87baf9bd560c · outbound

This paper cites Training language models to follow instructions with human feedback.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.185692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.185692Z digest=sha256:8c20e7920d28e6d2bea1d3804eab0fb44828bad986ecad5fddb708e7c812d388

Observation 98fe3752-c817-4a50-880f-dea6985cabb7 · outbound

This paper cites Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:01.019646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:59.380040Z digest=sha256:ea849ab6409978e9ed77bc45994b5179861cafcdfa15c2745606731bdba6f981

Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · outbound

This paper cites Critique-out-Loud Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.046421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.046421Z digest=sha256:f91620e82eb16c6720447bf9a1dd697b03f0c187bbb62aa1f170711a9d6098e8

Pith citing papers

No inbound Pith citation observations are available.