Pith. sign in

Paper Citation Record · LEDGER

$\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.08639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08639 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T23:02:23.432721Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T07:12:28.794227Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ef134bbb-3b20-4cc5-98f7-5003f38a2c54 · inbound

R.I.P.: Better Models by Survival of the Fittest Prompts cites this paper.

R.I.P.: Better Models by Survival of the Fittest Prompts $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T23:02:23.432721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T23:02:23.432721Z digest=sha256:99ae1f57e24eccc5d780523dd2ffd9969b22d9ee15107eafe11caf9fa684daca

Observation 8d6b33bf-3fec-43d2-9bf1-692c3fee2721 · inbound

Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking cites this paper.

Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T18:55:14.737534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:55:14.737534Z digest=sha256:668ed0f276d03c67210ba852b4545896b2c37860420da61fb32fde80840b1f5c

Observation 1a3e7e6a-baa7-4362-8964-7cf7b57cd679 · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.097604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.097604Z digest=sha256:fa4c4d7524d56f5cc7a867d9c593e7456a63e19489e9e9a4da53ba4906c8b4dd

Observation e8882991-f408-49fa-a646-9d5517477e26 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.415535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.415535Z digest=sha256:694f07c2779cb39e96c15a3047edc67b15917fc21e6ef106fb70058b86d675d0

Observation 34d1da5e-9660-4de1-8f71-242207ba596a · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.688099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.688099Z digest=sha256:008d56c004124d9ac1ef077c7809649381d54336353c289a7c457780fdc679ee

Observation 43b0d0fb-1411-4b10-97f6-3786b9ae1168 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:05.455019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:57:49.396570Z digest=sha256:1512399064dc2ee33c0c05045f8d66db63943916e7b83fd1a2db0623234c04a0

Observation 4f9bf3a9-0d75-4e9b-a1fc-cd3af649db25 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:26.495774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:26:54.426050Z digest=sha256:0a1ecbed0d079f0db5e8962be4b1f89f01f9f000981eb356241a7481561f08b9

Observation 7ab647a7-099c-4336-85f5-f71521799a71 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.796709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:08:39.328446Z digest=sha256:131dcbd45280af0f86642b5144cf0f724776d2f6c39e520a8547b448d9e4c574

Observation d0c9cb75-268b-48a9-baaf-601c0a28d473 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:53:22.786099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:53:22.786099Z digest=sha256:ee7ccf0c61233b5fdf67a7ccc47c2721b3997060fc8aba16ab4e347a8d556643