Pith. sign in

Paper Citation Record · LEDGER

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm

As of 18 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2505.01706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01706 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:18:40.431498Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da9d4998-f419-4811-846d-12cbe65310da · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.907673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.907673Z digest=sha256:8f09a5975f6803979d5b83feef036325ac0a2bbbba2de2654432e2600976ae1d

Observation 17a81456-0a56-4f8f-8587-92106ba508e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.947889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.947889Z digest=sha256:84b80a461fb3b5be4b845704496f3677f82e8689bc343c6deee5a362d30cc729

Observation 5b48d90a-eaa2-442c-beb9-bcc31db1c918 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Rank analysis of incomplete block designs: I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.952368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.952368Z digest=sha256:7ecf96bdc87eaeabdd0cbc8b2234e44e59beeed7a32255dbab84b0985dd1c50d

Observation 213a1309-5cca-4770-b8a1-14dfdeba5ee6 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:39.956959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:39.956959Z digest=sha256:0a7f95b3f626835e1f39a12b547826a0abd4b3424b3de1ca3b07e160bb56f101

Observation b9675d9a-123f-4c1c-b829-6b7d8892609d · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.034871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.034871Z digest=sha256:8b00d2ab4e7d40871448dbf3599fe8aaf0c95710e28d87242b2b0a94d001d913

Observation 4b214a57-c63a-44cc-a86f-fc50abdf03cc · outbound

This paper cites Deep reinforcement learning from human preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.092021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.092021Z digest=sha256:715d745d5a82d2e38da7c9793b17e336dcdc29b07ff70e20191ff5b239dd71ed

Observation 99edc449-12fa-4e1d-9084-e2a3c0e2577d · outbound

This paper cites Diverse Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Diverse Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.097323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.097323Z digest=sha256:1bf5af77d1aef370dee72d76423b65d2d2ccf75f8b7cd6774c30c61b4cb59b6d

Observation 46ec5300-70a3-4c60-a8f9-a7f7f07df7b7 · outbound

This paper cites 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm 2D-DPO: Scaling direct preference optimization with 2-dimensional supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.795844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:18:40.132613Z digest=sha256:92da8012dbe06cc0610803d33cd2fe179303a16789852555dc27ec3d79dc6032

Observation b99ca145-d560-4b7b-9ab4-46d5e33770e3 · outbound

This paper cites A Survey of Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm A Survey of Direct Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.171743Z digest=sha256:82882dd0ad6674d4860d08d4d8ea45942198d3dea6be21de451778e0d3747960

Observation 5d16bf50-25f6-4d67-9b7f-da58c07bff7c · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.205387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.205387Z digest=sha256:89b40cd5340338885a0d8f8ffab9ba14671a5bacad835852a180cd707df8ef6c

Observation 1341ab5e-4e74-41a3-b22a-a327ab782103 · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.210266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.210266Z digest=sha256:5d3d915f27faa3e8a1795344bcadac415df3b597d87e7011731e8656c8e83547

Observation 3fd6d03a-72fb-401b-ac24-a6ebaf2a3947 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Direct preference optimization: Your language model is secretly a reward model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.215067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.215067Z digest=sha256:d241e9d4a0c591ee7c6b7379e865e279d5d6b26c971ef79d576a9e6be22607e3

Observation 57710674-310b-406e-a080-2ada650490c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.224347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.224347Z digest=sha256:01b29d681eaed7c7de10348c4c632efb17a7dbbeab9b909a89330610206afaec

Observation 9953d08a-c9b7-499c-bc11-e479704e9215 · outbound

This paper cites Large Language Model Alignment: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Large Language Model Alignment: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.269907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.269907Z digest=sha256:2c025586fb70f9cf5da2eba3d50b7f7dd006d0fb2c48ea8b3f945d23968ffa92

Observation 857476ab-f5b4-4fba-aeb0-c03eae9761fe · outbound

This paper cites Things we like: human preferences among similar organisms and implications for conservation.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Things we like: human preferences among similar organisms and implications for conservation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.775002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:18:40.307558Z digest=sha256:d54426e16d97124d96a431dbf861ad96633e1bbd3603b6287e03da668b073a0d

Observation 7f60104e-3e48-426c-8b1a-31cc8ce16fdd · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Aligning Large Language Models with Human: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.312439Z digest=sha256:0de8ea43215f8460e90fb5c9154c54e48ea8be5bdb66aba43dd032e0c10e5da0

Observation 793f0b0b-2aae-441e-a55c-c64bb7e2653a · outbound

This paper cites Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.317605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.317605Z digest=sha256:ff72217c2d6d857b07b08664c3a22d916fd67fb04a07b12c429f43e844a39a47

Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · outbound

This paper cites Token-level Direct Preference Optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.424127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.424127Z digest=sha256:d0d0319270fe9bd63125dfad3ae5ce201bfa1b7e55fabd1980dccd673f545ea0

Observation af46cbf7-a3e4-4e51-8ea6-658b6ce5e590 · outbound

This paper cites Beyond one-preference-for-all: Multi-objective direct preference optimization.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Beyond one-preference-for-all: Multi-objective direct preference optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:18:40.678798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:18:40.427730Z digest=sha256:3efa7725c21fcb9a024185e0629fc6727dcc269c19cceac0f2b587c7e124ccbf

Observation 05846fd4-7eb7-416d-b9ab-dea33d3e4f0f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Fine-Tuning Language Models from Human Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:40.431498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:40.431498Z digest=sha256:2acd703026dd7312f3e4c6a997d12b4db601327a6668b3214c87c11ecb7a9019

Pith citing papers

No inbound Pith citation observations are available.