Pith. sign in

Paper Citation Record · LEDGER

RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2307.12950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.12950 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.380705Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T06:38:37.182755Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f87a3c6c-6d14-4a1b-86ba-537b10bb1783 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 298

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:38:37.185884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:f3257327d6febe90adae63a2c155079f1ed8fd15ae54d7b65c129f0fa890d201

Observation 37060bab-bcc1-490a-83ca-f6405add8166 · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.380705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.380705Z digest=sha256:e9e9f32807fc7dd3af92c1b1230a8b4d33854d92441256327c995138ea9bf5e0

Observation 962730ec-8fea-40a4-8f7e-dda1b4b8d5fd · inbound

Aligning VLM Assistants with Personalized Situated Cognition cites this paper.

Aligning VLM Assistants with Personalized Situated Cognition RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:40.224132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:40.224132Z digest=sha256:76fd21f7c6ebd06a301b37ee9a3b0657be9cf7e21a5f5173ac7c641efa46f72d

Observation 4159dc25-b6e1-4d9a-8ace-106e05b18e14 · inbound

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning cites this paper.

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:47.265500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:30:47.265500Z digest=sha256:d87ac7db7dd7c2cf0c100cd0ef5f10e385fde2637f69dcc21921acbefe07cab4

Observation b97bb204-c99d-4d5a-84a2-b4dd046ed8e1 · inbound

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization cites this paper.

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:26:45.501716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:26:45.501716Z digest=sha256:8e3bbd1d0cc30b394488866a529ec3226bf91e6714cc41d6632cb8431111a397