Pith. sign in

Paper Citation Record · LEDGER

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO

As of 17 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.08965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08965 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:59.969588Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03fe4b88-0c49-479c-a137-bec123b343a1 · outbound

This paper cites an unresolved cited work.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:04:00.872967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.629212Z digest=sha256:fc88c9d5d9837f1f09989cfb0b97d84c483dfaa0cfdce0b07578dbd1f1f25486

Observation 76d43501-a5be-4718-a772-01626ecea597 · outbound

This paper cites strong signal.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO strong signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.712230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.701979Z digest=sha256:207a98814dbd7e391d8f389006e21b8ff1c85702de08967e7035e5f7c9f369b0

Observation 45c4b94b-5f19-4f9b-8836-fea49deded49 · outbound

This paper cites Weak Accept or Strong Reject vs.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Weak Accept or Strong Reject vs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.271775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.969588Z digest=sha256:2474068cae429712c22200583cca5927ac33cc68e6b244300ed992742827a4a3

Observation e7bb3a19-2414-47b3-bf2d-6a8c61a5c1b6 · outbound

This paper cites Generative Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Generative Reward Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.316747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.316747Z digest=sha256:cf58797256d056495689c5e1dfbf1b611950f97fa4d7f5e7012f31a2451d0cd4

Observation 0346ca06-8dc5-42b6-8bac-5b01371268fd · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.531765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.531765Z digest=sha256:9be095e89569d7e6c4b4f1af9b94ffee1b1335de6d53d20ff77a2c1b4d5c84c0

Observation 6921bfe3-fb3a-4a1c-9b50-f04875cc2178 · outbound

This paper cites The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling).

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.559616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.800755Z digest=sha256:acd4bc41f372a247c27786b7d3cc70aca0ef23a1b8ecb0bb359b25054bef3708

Observation 0d90e109-0ec1-4a4d-a75e-8c517cb91d5b · outbound

This paper cites This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.426144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.861614Z digest=sha256:05dbc955818add5772b77473fefd78bf56034ab104af40fa99ce0f4a81469d69

Observation 7e971c68-9613-45c3-98c3-6249acb37590 · outbound

This paper cites Proximal Policy Optimization Algorithms.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.448741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.448741Z digest=sha256:640d91cfa8c3a69e8a33ee5a5ce10c97e1fee5fc1e46d073b97611557b1188ed

Observation 93599e44-d838-4697-8c2d-62358cf345b5 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training Deep Nets with Sublinear Memory Cost

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.265968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.265968Z digest=sha256:b350cca8542ea7fa39dcde8969d938f7b5983c2ebc02e46f3532910d3b13a34a

Observation 0a49f408-87ef-4be4-8efc-4928a420e7ba · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.109692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.109692Z digest=sha256:b0cbc0b83787a1a4aba2b58cdb8c1f0c296a07320cadfa9d62cecd9043ffc0a5

Observation c48772e5-ded9-4861-8e38-87baf9bd560c · outbound

This paper cites Training language models to follow instructions with human feedback.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.185692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.185692Z digest=sha256:4cfcd23620962822f75a9b9c8f15c3f089d0ada032d05a9a9811c6dae49a3ccf

Observation 98fe3752-c817-4a50-880f-dea6985cabb7 · outbound

This paper cites Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:01.019646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:03:59.380040Z digest=sha256:f9ecf79b999c5d7ecf3a31a767723ef72eedb545a153fc83ae3b62af404e52b0

Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · outbound

This paper cites Critique-out-Loud Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.046421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.046421Z digest=sha256:a7eaa4e9fd4514203e458aedf8a08cdc982fb75b5f1e885b0c3f57b5ce7ef00b

Pith citing papers

No inbound Pith citation observations are available.