Pith. sign in

Paper Citation Record · LEDGER

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment

As of 17 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2505.21395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21395 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:05.082364Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22f83a12-a945-4a69-908e-4ab81863aa5f · outbound

This paper cites write newline.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.632371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.632371Z digest=sha256:1649baf6241318f8804d4f2665df1572b94c9efb40d69aa5bc7dec5685e8fec8

Observation a060b2f6-9f79-45b6-ad81-05cac294c2fc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.693724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.693724Z digest=sha256:ff0ab40362d12ab99e6c4ba464461e1d71d7b94e1c02bb62f7628d8730423d34

Observation 45695cbb-b815-4b7c-91bd-efdc0d7692df · outbound

This paper cites M., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., and Sun, W

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.808597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.808597Z digest=sha256:310c5b466055d04e67edea43614eb129115a5ac98b545a3d5c440596083c9825

Observation f2fea3a5-9512-4a06-9378-3726f05d2273 · outbound

This paper cites Harnessing Density Ratios for Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Harnessing Density Ratios for Online Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.889918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.889918Z digest=sha256:d849cb3e44b6d2dc510f4ddd8a892018ee2b4aa404dbb6d5dbb36b847053aa89

Observation 0e61b7ea-ceac-44e3-a682-6d065032228d · outbound

This paper cites Scalable Online Exploration via Coverability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Scalable Online Exploration via Coverability

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:45:06.394885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:54.982699Z digest=sha256:91d85749ea52a024a2532ce6a919c58586333a51ea1cc6af7d6989a86e2cccf4

Observation 59df60b0-282d-45f0-881a-64d8913f321f · outbound

This paper cites G., Guo, Z.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment G., Guo, Z

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.055754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.055754Z digest=sha256:704e05e08a9aaee165330af27e7048a9a471b6fa48dc722c453e548734db0a05

Observation 84b05808-b330-4592-ab14-99f21e382e4e · outbound

This paper cites M., Schneider, J., and Ng, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., Schneider, J., and Ng, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.169617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.169617Z digest=sha256:4be1081113c783737f2d615db103854bd2d7bff942a5c7ba950035f23a68a78b

Observation 63de237d-f678-4794-b06d-4da79c494ac9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.238757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.238757Z digest=sha256:940901f33dcebbbc5367dc476dfa5f923ef4bfbc2aee4ac5a799b11b20dc51b4

Observation cf6783ee-6407-423e-99ae-962abfa1c415 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.377088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.377088Z digest=sha256:1cc9de3021293ac7a90f13d172a8a7d67e12b9cae5858b0a237b6f02d047a264

Observation 14710f35-8db2-49c9-93e0-44bb0f7bd651 · outbound

This paper cites Contextual bandit algorithms with supervised learning guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Contextual bandit algorithms with supervised learning guarantees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.517503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.517503Z digest=sha256:ce92a05ad645543c54d979761fc75373e45459e7bf7f28e9f4dd6bcabee9bfdd

Observation 22a88ce9-3be9-4e59-bf16-9bcc8471aaa5 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.605096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.605096Z digest=sha256:6736df7a9a00d986f2afaf2fded7b6ac3be4b43bbaf200a32664bbf32d21e610

Observation 240e18bc-3414-4dcc-84a3-ded8531ff22f · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.718587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.718587Z digest=sha256:95777318afb5f9eff3e3b0cd8e3acbdf15549c0b1cb38c5af85c0e4ec3de0d0d

Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.826681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.826681Z digest=sha256:228129ba76087f278b6d310c4b6871b1e05ff5b1d136ac9733323cb3841e0b77

Observation 294d0f4b-5166-4dcd-8e91-7a8a6a731517 · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Dataset Reset Policy Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.917428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.917428Z digest=sha256:80dcd81f708567c666a008da1133971a4b534e259eb5efcfde5cdafbce5edaa8

Observation c63e23e4-6167-4dce-9523-9a09316c548a · outbound

This paper cites Robust and private stochastic linear bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Robust and private stochastic linear bandits

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.010458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.010458Z digest=sha256:74569add44ea3cc7e6ac48a1ad114d2082be78e7d5abbfbd1ab8e468b7baef9d

Observation 51b7fda3-0a4e-485d-9cfa-2a913759b7d1 · outbound

This paper cites and Hsu, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hsu, D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.102692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.102692Z digest=sha256:015337a18448049cf47c9c49f74c94070212b10dee4682cd1408e8044b7482a5

Observation 404f5f1f-107c-481a-9b74-dff23f57bfc9 · outbound

This paper cites and Jiang, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Jiang, N

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.239470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.239470Z digest=sha256:2538e72088e1099909b24e36d297d1c0f196aaca11c8ae1be89235ceb83cc2bc

Observation 3fe4d398-1f28-4680-b100-e0cb89748d74 · outbound

This paper cites and Sentenac, F.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sentenac, F

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.337210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.337210Z digest=sha256:0f999e339288bda5c421696fdedda0606bb17d8d3c7f88d1a78e1c1cd6e8fce3

Observation d938f58b-be16-4a0d-843a-81517f30b9b4 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.414470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.414470Z digest=sha256:34ad0438407c2fd5bdb4e0b3e50dd40a0661ff05b2217ae1003d56b438f33f0f

Observation e50edb24-1c97-49c3-b789-41f0bbd070c8 · outbound

This paper cites Distributed Differential Privacy in Multi-Armed Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Distributed Differential Privacy in Multi-Armed Bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.541880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.541880Z digest=sha256:664558a77381adbbbb057b3f97a086487d983809e3f52fc7f22d11be964883ba

Observation 264da0f3-a7cb-455b-94c4-2314d2555433 · outbound

This paper cites Shuffle Private Linear Contextual Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Shuffle Private Linear Contextual Bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.606840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.606840Z digest=sha256:251d79cd511a49e7f40afa8d2eb4e3f840f38ece8f5ce8a0431b381ca42c92b0

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:ed78d2ee84dfc671fb1020f4fbd835a7feecd55c41c73899e58f35ebdfe115c3

Observation 05e2ec5f-65c4-4eb4-b309-760c60063db9 · outbound

This paper cites R., Zhou, X., and Natarajan, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment R., Zhou, X., and Natarajan, N

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.734155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.734155Z digest=sha256:b3cad8af38a967945c6daf04dfe343578eda38d8330d5cc1d42ba88417ba3f61

Observation a20de056-40a9-4f70-a0a4-d6ffa2a43c34 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.793524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.793524Z digest=sha256:830b53183bae65f3199999b4a251ad9889d6c9abcf867c340e44675df119aa85

Observation d34bfd27-8df7-4de2-aeec-e32b1e7132e0 · outbound

This paper cites and Du, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Du, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.852370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.852370Z digest=sha256:f74b3771797540b8fa2db0e618288f89f963f62959458a3045f279bd78844e73

Observation 25e84e02-d205-473a-ac09-e48864c73dde · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Minimax-optimal off-policy evaluation with linear function approximation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.912450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.912450Z digest=sha256:1da4261392f995c3ec67487ab743206b96ab0eb569e97bd883375da972965cc6

Observation 2fe68c96-036c-4c61-ba86-8735a69f9b79 · outbound

This paper cites Calibrating noise to sensitivity in private data analysis.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Calibrating noise to sensitivity in private data analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.002442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.002442Z digest=sha256:4e058bb661d8a0fcf8147b1c1dc889ea1cc986a155f9ae2e5f045bb44acc46bb

Observation 3280134e-d360-4f8c-a9b0-dbcfe4d6476e · outbound

This paper cites Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.064925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.064925Z digest=sha256:a7f698ba337c9d0f92abf7b616105c3210ccd929b96b74c95e3c7001ce30d3b5

Observation b8409717-f79f-4245-b9a6-2ea6febaafa3 · outbound

This paper cites Importance-weighted offline learning done right.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Importance-weighted offline learning done right

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.137384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.137384Z digest=sha256:88eb367bfd18ac46e03b07c9831b65e0a286602487228810c6b6b0a6e276898d

Observation 47c6f8f3-e850-427d-a039-f2b44a779a37 · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.195037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.195037Z digest=sha256:eca950abb8dfabffed5860b1d3f5f109f253fe319254832c5b067282f1f3ed59

Observation eb4675ef-62f4-4075-aedf-98bf12cc7089 · outbound

This paper cites Local differential privacy for regret minimization in reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Local differential privacy for regret minimization in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.243840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.243840Z digest=sha256:c462c6a5b2dea98a5144d570ec47a9f62cafc590975c5ee84beb1746f716d7ce

Observation bc140ad7-f20b-4a3c-a01a-da6da7cd4447 · outbound

This paper cites and Hopkins, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hopkins, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.323106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.323106Z digest=sha256:b028378f2137f5d2d56cc8ec67f57ed0ae217c7145f4bcb1ef5c3e05bbe5203b

Observation f12048df-d6ce-4d49-a913-cb2bba9b9466 · outbound

This paper cites B., Kamath, G., Majid, M., and Narayanan, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., Kamath, G., Majid, M., and Narayanan, S

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.392906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.392906Z digest=sha256:cbf6a9dc1646b1e9ef676018b49c332922c807c1c2a79646442be0f24eb09e30

Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.450115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.450115Z digest=sha256:44e143e6b7bdb1dfe4dd902d95afb5f13f948d58d2bd2884faeb5540204ccb33

Observation 8b846a68-5427-4060-bbec-7f6dc4f75ef7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.534624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.534624Z digest=sha256:3ac0c27ca879e32a8195874386fd9ec7cffaef469c3bc119b00b680861f8255b

Observation bcdcba6a-9b70-4803-abe7-3893d49db018 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.616181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.616181Z digest=sha256:5dd0cc6e52456d275d77d6088da1e94501fefe54bcc989d9361c3efa7d0982bc

Observation 8020fbf3-3226-454f-882d-00ff24c036a2 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.676519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.676519Z digest=sha256:9aa9f49470aa6ed85b8e15f5ffeca1b5c46d6a15b717cd7edffc57c774688502

Observation 0ca989b1-0cf9-41af-b94d-c4c55d00eb5d · outbound

This paper cites and Langford, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Langford, J

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.751781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.751781Z digest=sha256:7d99c51f421d9850fc7c059d5c10f30e09a52ced7231b3964b95d2dc9f716fff

Observation 91464f4c-943b-4e25-b532-9f3ef1473a6d · outbound

This paper cites The Broader Landscape of Robustness in Algorithmic Statistics.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Broader Landscape of Robustness in Algorithmic Statistics

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:45:06.102730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.815194Z digest=sha256:b0e8300ae9d6d5c094d738c5fee51aee284124ebea6f338a5fc9218e0a736c9e

Observation 14b5b11e-9475-4f5b-afa7-042b1927bccb · outbound

This paper cites P., Lee, H.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment P., Lee, H

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.831953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.831953Z digest=sha256:da85823d67fbf2a45a4efb1397b12bc4d552012d237291fa116e690acaef98ad

Observation 071cfa26-ec20-4869-b252-6081c14f389a · outbound

This paper cites and Brown-Cohen, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Brown-Cohen, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.444969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.836692Z digest=sha256:5415af0b8bed24afa371540d370cb865671b33e03b59ca6fed01a1681bb7c673

Observation ad056ec6-d3dd-492f-8217-2ca786e0b09a · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.208728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.840951Z digest=sha256:83267fd9954e3e61fb732692b6002575ad001758cb7076a8e87aa80c2e933e3a

Observation d9b1c00a-e13c-462e-b624-9bb8d9f8675d · outbound

This paper cites Differentially private linear bandits with partial distributed feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private linear bandits with partial distributed feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.045189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.846124Z digest=sha256:7c597f467a622c4ccadb790e73b2258a39f56e7f3ba51480b351f129ae094b5c

Observation e2e13a3b-36b1-44a4-b8ba-79f9a6ba2927 · outbound

This paper cites B., and Yu, Y.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., and Yu, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.887776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.862941Z digest=sha256:057f7f2daec3e30392f943145206fc915e8c34673498d0d87ffced0fa1cc2b4d

Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.881471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.881471Z digest=sha256:8a9cec91e71a901484d423ade83696cfa0d3d30d62006a42e6295fd4f632645c

Observation 206d38f8-6001-4688-a32f-53524120eb29 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.905368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.905368Z digest=sha256:dd2f29ff2d0768f88328887ecd691656be5a61bfc863b3bc06bcaad3b9aa43e3

Observation 52856b81-da6d-4fc1-8ae8-d8eb8190a6cf · outbound

This paper cites Y., Yan, J., Jayaraman, D., and Bastani, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Y., Yan, J., Jayaraman, D., and Bastani, O

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.640870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.963295Z digest=sha256:81323396e3426f7255063c195bcb49479a0421272a9a5c9ecf1055d23b88e82e

Observation bbd543d0-84ec-4ba1-afb1-dea967749296 · outbound

This paper cites Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.989120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.989120Z digest=sha256:6e04a203a49ea64bfdb5422f67720e2b4f6ad7caa778e86f519437ccdcf1ab1d

Observation b7b5f535-eca7-4929-b022-dd2a8c11a880 · outbound

This paper cites Corruption Robust Offline Reinforcement Learning with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption Robust Offline Reinforcement Learning with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.018188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.018188Z digest=sha256:662355ba08deb671b9cf02e51a7b43273aa581cfca955be7896f4884f30263d8

Observation b1eb1efe-0e88-4f7e-b1e1-30674758ec0b · outbound

This paper cites and Talwar, K.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Talwar, K

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.067713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.067713Z digest=sha256:6b8b87c9b7c34d42706f342e622ede0b53da3365274d8873a6ed2229db335738

Observation 57dbc156-7351-41c5-a40d-874bdce57ac9 · outbound

This paper cites and Thakurta, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Thakurta, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.445447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.166200Z digest=sha256:2c5a73cac6258234b383654d82723de830e2a07d130e1cb31242ce902ef27958

Observation d20f5511-e422-4b99-9a44-aa19f99bb3fd · outbound

This paper cites and Szepesv \'a ri, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Szepesv \'a ri, C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.207803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.256381Z digest=sha256:d4e932600bd9e2d612e2592d1238fc84406aea89b138eb15f9ddaeb65b8efd2c

Observation 4825c11a-6490-491e-817b-18b2bfbc143e · outbound

This paper cites Nash Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Nash Learning from Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.339973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.339973Z digest=sha256:87b1aa04ca2b13f339ff16d973eb2b4361808b1141a409c1bee35e917c20f1cd

Observation 07f6a628-19e0-4ccb-8873-35e77c39a2fc · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:45:12.005072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.441762Z digest=sha256:f867dc159229e301f9f3136fd1a77a5a7427a2df71c1e4d7be9cd88ef3ed8f74

Observation 1b7c65ef-b82f-4a21-8d04-b3f32a471f12 · outbound

This paper cites ChatGPT : Optimizing language models for dialogue.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment ChatGPT : Optimizing language models for dialogue

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.754842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.569884Z digest=sha256:34c27c25569094efcf02acd295236d3273600db50f2414300674eb7aee5933ad

Observation 6f5f7b64-9e18-40d0-b252-5700b8401841 · outbound

This paper cites Training language models to follow instructions with human feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training language models to follow instructions with human feedback

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.599409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.720589Z digest=sha256:5b992a6dc271c4cf21d5c17334596e4e4121cedb600eb4c76b2e6fb9e0ba303f

Observation 14c19cfa-deb2-487a-959f-8f20c231999f · outbound

This paper cites and Wang, Y.-X.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Wang, Y.-X

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.361226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.865665Z digest=sha256:b2c63b455700496ec192599f75d8acd0e84281ce119fdb12df61314e096086fe

Observation 9288cd28-495f-4282-b1f4-204b56685608 · outbound

This paper cites D., Ermon, S., and Finn, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., Ermon, S., and Finn, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.164787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.994773Z digest=sha256:e600d84a78ccae200cf7c6689a9fc05a3f8e18dea5c2d771cbfb6c5e57f83f91

Observation 3d0142a0-ccc1-4823-9912-5ef0c8bdf33e · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.913677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.170713Z digest=sha256:baafc3c8dc477b38a3d94fca4d1759cff349e4ac08f12f2bb92f520975b8aace

Observation 1bf8d535-f667-4f41-9793-1c298983eef5 · outbound

This paper cites Multi-Armed Bandits with Local Differential Privacy.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Multi-Armed Bandits with Local Differential Privacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.345604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.345604Z digest=sha256:0bd8323c3fb8a965a2d5ef8b2adc3baa1a1fcb53801c58df3ac2db635302d88f

Observation 4518b6ef-e776-4b72-8342-f7fcb33ae2f4 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Agnostic System Identification for Model-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.521221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.521221Z digest=sha256:740fa758147a9aace0759e7080fadd0a392eedf3773264b30828f9d871cb557a

Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.736916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.736916Z digest=sha256:60b8e984f66de29bb7d3dcfdd6f5ff90c3031cf0ec7213f23f00a5518b01dea6

Observation cc91668b-e6fc-4d75-9fa3-64bededec017 · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.649260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.899253Z digest=sha256:751ae36f8986c4be0b3e7e346a8bd15ae3fec1e1aabedd53cb56b540563b0054

Observation 08a18c92-2d24-44ba-aec1-fa6dbc4d9703 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Proximal Policy Optimization Algorithms

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.057955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.057955Z digest=sha256:79756727132a9afac3eca95476421626b53287c999ba63528eedbc3778bca70a

Observation 39eca17c-cd02-4985-94a5-84b9be50371e · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.438346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.151879Z digest=sha256:f68e2bd07bfc71c5bb102354eb5b12f1153bc348708a68c3aaf44c6f0661973c

Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.273913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.273913Z digest=sha256:569e716433008524394a4364a6524fe6ba9b3c72237e17eaed6df5b0aed3067e

Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.439564Z digest=sha256:fbdf90140a4c0b11d53aedfe0f5170b9352b77fd56eb1cbcbff4f80d4e77f92e

Observation 00b963f3-8f72-48a2-a7b2-787bae6547ce · outbound

This paper cites The importance of online data: Understanding preference fine-tuning via coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The importance of online data: Understanding preference fine-tuning via coverage

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.250643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.549181Z digest=sha256:cf0f615995f9833d006e538f72b17f9b7b343d56e1d1708132641e2e7e7b1627

Observation a871c859-9728-4a2d-a6a5-96ddc928938a · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.714395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.714395Z digest=sha256:f9376af6dcb7265d45f14678f247c956e168deb317939fdc298f38ff468635f5

Observation 349350fa-2915-4818-9aba-b70be6b32109 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TrustLLM: Trustworthiness in Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.876180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.876180Z digest=sha256:02078de1944b2bdb9f98a856c8f03f5859f0c3d61897171a889bec120a80195c

Observation cee8bc4a-7444-4747-8fa6-c8e6159e49cf · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Principle-driven self-alignment of language models from scratch with minimal human supervision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.086380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.019288Z digest=sha256:a2247c7fb6c99f30dc89fe93c7113f77e26fca3f8b6c5bc3909bcf64b557862e

Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.173114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.173114Z digest=sha256:b194ba6f78eeab6964cba11b3207c6f940d67b50ac8e02d0f95e91aee344e2d1

Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.361604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.361604Z digest=sha256:c47d9893a1e92f48d6550cb304a1c5e9bc45c5d16a2ce9db367924149dc2f21b

Observation f84149dd-0a3d-42b8-8d5f-c638c3e517b2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.530685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.530685Z digest=sha256:f71c73a38ce28238b6fe473b9a145c7533999f3acaee6be3c629f871f82d5de5

Observation 32fd6239-8a15-4a5e-8b43-1d4eb839d10e · outbound

This paper cites Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.676167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.676167Z digest=sha256:294d777a6e3c02d8fdc6592c80a459e3abb5be065ea067cc4e83e47c565efa6f

Observation 13631173-84fa-4065-84da-7aa8bb414245 · outbound

This paper cites Private reinforcement learning with pac and regret guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Private reinforcement learning with pac and regret guarantees

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.913242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.801236Z digest=sha256:adc4dc28fc76f7ef076e58dc537b3a7b378f3a9466ee974e22f950a9068b8a2d

Observation eb59928c-97bb-4b8e-80ad-1443f37636a1 · outbound

This paper cites TRL : T ransformer R einforcement L earning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TRL : T ransformer R einforcement L earning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.749209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.936084Z digest=sha256:bb677fee970035234da047692d953b02f3115719e76aa2a2e5efbf1ddb22afda

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:d8abbf1e0ba422c87e6a11c8a820edd03bdc6b0e25aedbd8f76a4c81afb927f8

Observation 6dcc33b2-c25b-4b26-9f24-290ac9944091 · outbound

This paper cites The Central Role of the Loss Function in Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Central Role of the Loss Function in Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.217073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.217073Z digest=sha256:0d79aa75f625d6e642288c9c834cac350ea3ac91412df47acad58bded73807e3

Observation 408a1957-edf2-4e6e-a5c0-60f311fc7d14 · outbound

This paper cites Oracle-efficient pessimism: Offline policy optimization in contextual bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Oracle-efficient pessimism: Offline policy optimization in contextual bandits

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.514295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.369874Z digest=sha256:0aca82ef12263e86a3eb49fbdda9afce3494ba70cadb5f13e8d712e5533db845

Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.561160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.561160Z digest=sha256:b9563b6e3ef5ffb3218921408178d5d7ee232bd0defbd6147cd86c9c214fb7f7

Observation ef0913f4-fe91-46fc-8f3c-89eb5bc4d9c7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.755041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.755041Z digest=sha256:02044f400457b8f97d7f94b76d208e599bd616e95428e93a25626bbb177e62cc

Observation 38d55713-5938-490b-a9f4-e55aa35bebe4 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.353478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.884175Z digest=sha256:ad88bed01ca8083a3e2ecd26a60b3700a6e96b8382cf3f441f5e4d3e16fb72ef

Observation d4a93fb1-1819-4d37-9e85-7a6155634ec4 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Self-Play Preference Optimization for Language Model Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.046755Z digest=sha256:c1aea27eac7ee3e1127fe6f7d92c309e379428c737c77b032c1c393e166575a3

Observation 24c71555-cb6c-429e-81e9-b096df6f3489 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.153167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.207476Z digest=sha256:06dad4d5248467969d1cff3bcde9809e7d265ac59e082c0564637c0dc1e3bd30

Observation 2d05d8e8-2c3e-4b5a-81ae-0e73ee18c877 · outbound

This paper cites On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.350482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.350482Z digest=sha256:05986bcc5c8d549f3d7b66e0c7ebd8b301e0fcf571d7edf5647bdde2728a6b1d

Observation c70c23d6-df86-41e0-8da4-2b2b588b97e4 · outbound

This paper cites Foundations of Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Foundations of Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.495496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.495496Z digest=sha256:9ab59032f39306b2c25be72a5eabc2758e37e9b376fee770d5b74d4b87a85a13

Observation 5f943956-9991-46a4-94aa-57ee8a6fb445 · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman-consistent pessimism for offline reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.964046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.635728Z digest=sha256:f18d2454708c8eb95779312e524111ceea89c0f0113e5a1d5e71d60817bba586

Observation 47693abf-7fe6-4f98-9725-ed925a50e51a · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.783637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.708960Z digest=sha256:11098b59bda802817f2bab7aac7006e22870377ee034bb521036cb725bc9d0f3

Observation 2c102604-556e-4131-b27d-df8e3dfe5770 · outbound

This paper cites The Role of Coverage in Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Role of Coverage in Online Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.851174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.851174Z digest=sha256:48dec80eefdc99d685219848daffb2568758007b0c78a2472105637bef72798e

Observation f24e11ad-5bc0-4e7a-a0df-1b77114804f8 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.971759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.971759Z digest=sha256:0bb8378866d81f5a8ad5cd22a6586ee334297dfd253de89268db5e936dbf3d35

Observation 1e623a67-03b9-4747-8d5a-b1b369fbcd82 · outbound

This paper cites Differentially Private Fine-tuning of Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially Private Fine-tuning of Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.056564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.056564Z digest=sha256:a76d229846720f35005724f169fde4c7795c69e36d17efc985374d258c28ef38

Observation 99c7d7ac-d4bb-42ab-90b3-8e8cca818102 · outbound

This paper cites Offline reinforcement learning with realizability and single-policy concentrability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Offline reinforcement learning with realizability and single-policy concentrability

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.527175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.163239Z digest=sha256:44e43181ec9e8efb4ec3589eb6a7c6fda157c850424c88c3a1b23f8bc94b7ee2

Observation 2a0bc272-c655-4f37-a4d1-9a8ef6274c35 · outbound

This paper cites D., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., and Sun, W

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.315920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.365119Z digest=sha256:2e6920dfa4bfab1a0d3a490f99ebd1af2244cba9c8cb18245ca0104cc85d3448

Observation 13b583d9-506f-4e74-bd85-4e04f870217f · outbound

This paper cites Corruption-robust offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption-robust offline reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.096754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.456100Z digest=sha256:9858c3fbb6315e4529bdc8ae8bcb257dc583d91008f258ebd92d2de1592bab2a

Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.554366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.554366Z digest=sha256:4583dd666d0a0d64e0fea933f0bed90d9a79591d41fe316b3efe870ce3156a33

Observation b58fc734-8508-4caa-bbe5-520dddb3eb8d · outbound

This paper cites Locally differentially private (contextual) bandits learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Locally differentially private (contextual) bandits learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.834248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.710387Z digest=sha256:c1efca32d8ae46622264a8eaeb88b25a4c43eed58c195ad71b2054923b9046d3

Observation 4b5a4910-4606-41c6-a9f4-f1cab496e3d7 · outbound

This paper cites Differentially private reinforcement learning with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private reinforcement learning with linear function approximation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.623841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.842225Z digest=sha256:2b84a0e3e4f8db7ca77239a772c441fede87a797cf7476ce4cda8e2b7f29f72f

Observation 8708e2fc-7628-4e8b-ad26-4173b24d4d05 · outbound

This paper cites and Tan, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Tan, J

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.448453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.955804Z digest=sha256:7337126b59b9dd18b40d67744980693e2a7250bc775b903ec998d1bab15fc6e6

Observation a6de8d3e-efe2-4eab-93c7-daf7f2c38802 · outbound

This paper cites and Zhang, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Zhang, W

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.202397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:45:05.082364Z digest=sha256:af222b67b0c30080269e9bc61698ba309e5ab86c8cd5a4d286a60e3761af3eb4

Pith citing papers

No inbound Pith citation observations are available.