Pith. sign in

Paper Citation Record · LEDGER

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

As of 9 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2506.03066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03066 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:15.714867Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:48:15.394329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.307998Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9703832-2747-4b73-8c47-4467da086c5e · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning: An introduction, volume 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.114854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.114854Z digest=sha256:1cb07115b97499a98a2be2e28b0e97308acc93e432f35352aae15aed2e514a5b

Observation 70fe8ede-870e-4e44-ae03-67b1798c8153 · outbound

This paper cites Controlled experiments on the web: survey and practical guide.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Controlled experiments on the web: survey and practical guide

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.203695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.203695Z digest=sha256:07c56fd26958c106ecd6a954b94c499a18e46185be18ca14c1c2d8d693bf4f4a

Observation 2de9bc8e-27eb-4749-95be-5ba5d2ac8ec5 · outbound

This paper cites Deep reinforcement learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Deep reinforcement learning from human preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.296587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.296587Z digest=sha256:8f33e93f5397250fe16c4975dd4af8d72e03a49f05151391516fd41b86f447eb

Observation 2916cde9-d8b5-465c-97bc-8968f3927110 · outbound

This paper cites Al Sallab, Senthil Yogamani, and Patrick Pérez.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Al Sallab, Senthil Yogamani, and Patrick Pérez

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.405514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.405514Z digest=sha256:61808870b64a85d6ec1ed0286166825aa23155b11647d210c8c91c9dab2d6f32

Observation 739bf0c4-c7c9-4cd1-bc45-3be231cf17ea · outbound

This paper cites Training language models to follow instructions with human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.525911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.525911Z digest=sha256:7ff78368d10b724fa60769957f5acb7333f313630d696cd1bb25d458cf04b6c4

Observation ca01d7dd-fa85-480d-9975-8a0b310cb5f4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.608779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.608779Z digest=sha256:0d9c54f1aa8a05afe3efe155c49b55ec48baf1fca8d9d86230a3c4ceeb3e0d0e

Observation 99c7fe35-9628-4b78-b4ab-abb6f85799f3 · outbound

This paper cites Dynamic programming and stochastic control processes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dynamic programming and stochastic control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.716180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.716180Z digest=sha256:3da3f3fe8a6c9976e98f9de76ed27a645f79f882864ec8e5d55b0a744bffaccd

Observation 192aba53-7d83-4907-ad1b-7861438a4a85 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Markov decision processes: discrete stochastic dynamic programming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.827799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.827799Z digest=sha256:73193db92a922b8694bd9f6c2fa2a648778b9a482951d8fe4d0317dd0db39e87

Observation 5d6406f9-6fc8-452a-874a-8d9f581f79e9 · outbound

This paper cites Inverse reward design.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Inverse reward design

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.891470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.891470Z digest=sha256:b551cf8cc08fc89ee43da3057bc58d4a7f6a0449e2b1463c71d944a332ef57dd

Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · outbound

This paper cites Reward Design with Language Models.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.973403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.973403Z digest=sha256:7609d28eff5c59478c562e49ad825112c884b00dffc4ad63bb29edbf7e38f47a

Observation 6f2ccb55-f03d-4c47-9196-aa7b47ef13c6 · outbound

This paper cites Defining and characterizing reward gaming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Defining and characterizing reward gaming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.056676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.056676Z digest=sha256:0081d700c87f71f964e7c78c021bbac2ff8b4c21927ccd39a18c9b8583dcbe23

Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.242520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.242520Z digest=sha256:dbfbf40a61f4e00a8816103cfca96516a421759112de35bcbf726be9c6feb877

Observation 31d800c5-7bad-4ed5-8efc-6e2d80ecf082 · outbound

This paper cites GPT-4o System Card.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.353435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.353435Z digest=sha256:dd3badb7cc757ed0e541ddaba719c7637ee84057b5c173e570bb9c3ab2efb8a7

Observation 1534fe5b-74b5-4f5f-a2bb-9c1b5495e071 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.439190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.439190Z digest=sha256:4cf4ef99209725c165968d697e968f87516bf08b5ae7c825102327f858977ef0

Observation e5c0ff52-d75e-47a2-bb45-e07577dc1094 · outbound

This paper cites Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.434223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.554482Z digest=sha256:74628dad22363e0a9efa2fd01b6f9a31bf8d2b044e5357dbb32e990af77353f1

Observation 5def0db4-fbcf-475d-b4d3-3d100ea593d7 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Proximal policy optimization algorithms, 2017

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.665749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.665749Z digest=sha256:7e1474e0556c54e61f2a5c2357b95dfda49ee1c891f60f6ef324b1f64550afc5

Observation a80d8297-66c6-4289-8a23-988584021b24 · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Open problems and fundamental limitations of reinforcement learning from human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.322900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.755552Z digest=sha256:b719e63648542f4964677e4c7fb96cb0cb83ddca7206bd68100aed19f3504919

Observation 2b6d06d7-162f-4354-abc4-664102270b38 · outbound

This paper cites Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.164237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.793743Z digest=sha256:dbba215d8dfb894d88aacfc04ac6b6f0ab96bfc49835ea08b19a921196bec954

Observation 4656b496-a7a3-473d-9590-2e947e847154 · outbound

This paper cites From r to q^ * : Your language model is secretly a q-function.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function From r to q^ * : Your language model is secretly a q-function

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.963362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.799620Z digest=sha256:b319a36ed9f97b043156a190b34422f95339e7947a136a896ffa14641f4557b5

Observation b76e2247-cc93-4d0b-841a-15c87dfc1c31 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:23.787862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.839237Z digest=sha256:2bb2bffaf268606120fe03075e45c57c48b0533a482e5e93cd401bd5329cd049

Observation 2105833b-782a-4fa4-b5d9-68dbaa0f8459 · outbound

This paper cites Random utility theory for social choice.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random utility theory for social choice

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.906112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.906112Z digest=sha256:26a1aea46af80388cfab1be3991c7bc212cae1aa090ab2877f4b095e22e7dfc1

Observation 7519c8fe-e391-4f82-b614-f4e0aa2e2cd9 · outbound

This paper cites Modeling ordered choices: A primer, 2010.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Modeling ordered choices: A primer, 2010

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.658611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.971612Z digest=sha256:b98f3707c26a8d53a10c47ba4d7265ad4dc8fc2170cd854f2969bdbb173c880d

Observation a1d649d2-8446-4572-87bf-486af6d3ef41 · outbound

This paper cites Econometric analysis 4th edition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Econometric analysis 4th edition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.504855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.030921Z digest=sha256:cbba38448ec9dc086cc9ea1c1d022743ffed0d54fe0e91766679df78bbfaad29

Observation 571facf2-89e0-497e-8702-1fd5b464a369 · outbound

This paper cites Sensory evaluation of food: principles and practices.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation of food: principles and practices

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.282474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.101142Z digest=sha256:7b66ae76651654e97d6c18d8b9173e859464a75822c58db872e9e0e7bdf99095

Observation 097214b3-cd0f-4089-a3e8-c302ce0d7ae6 · outbound

This paper cites Sensory evaluation techniques.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation techniques

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.069453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.166328Z digest=sha256:6f326b442c37009ff848b3a28ef95227d57c1634203e39b2f42f2c1ca1a8765a

Observation 0b3b9f97-ef92-479c-bb26-deb82ab8b5de · outbound

This paper cites Nash learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Nash learning from human feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.880574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.215572Z digest=sha256:47117b409d1457499ae7fcb7c5d04e40405fb01bfd369fbed7a310727e064846

Observation 9cca57ca-66e0-4d12-bbd7-2c97d1726096 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A general theoretical paradigm to understand learning from human preferences

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.715947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.318773Z digest=sha256:7410801516ae0737171e6590a873c2855abcfaa7a8abd643c20899ffc11b526b

Observation f44cfb7f-350f-4cc4-ae86-c53b6055920e · outbound

This paper cites Models of human preference for learning reward functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Models of human preference for learning reward functions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.557457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.395294Z digest=sha256:ae7fce79737c4a88d9fd0f8591bc561e5e3627ebcc5e6edbf38069b4e1b54dd4

Observation 1120e8b4-c18f-4faa-8d3b-7be8f969d57a · outbound

This paper cites Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.399902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.569486Z digest=sha256:84e4290567144adb4639961418b9d38e691ea8737e8eec6d581120903cbb6bb6

Observation 16fc529c-6624-49cd-89d7-04165b8a6dc4 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based online learning with dueling bandits: A survey

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.244253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.737974Z digest=sha256:2b5bad4d5ce02e0d8062e030755cdca349a4187252fc4218345f8d1551f4bdf4

Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.854603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.854603Z digest=sha256:2422373b9fd1dd4f22fad7471126c40269c9957975f19cf0b767a56ab8377cc2

Observation 155b393a-2365-4814-8dc9-7b9c21daf1bc · outbound

This paper cites A survey of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A survey of reinforcement learning from human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.936227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.936227Z digest=sha256:9d228061ce21ec64521aa5966798c9edc22b71299e26808078c7a008a2eadfef

Observation ce96e202-23b6-4bcb-aa22-30374648414c · outbound

This paper cites Scaling laws for reward model overoptimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling laws for reward model overoptimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.106083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.957795Z digest=sha256:7cf6447812cb69cda7c397ceb5026789ec3f04f6adca30facfe14729745099e5

Observation eb2419d6-f622-4eaa-b050-50fb1789f50c · outbound

This paper cites Model-free preference-based reinforcement learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Model-free preference-based reinforcement learning

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T11:20:15.864764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.050288Z digest=sha256:552a906eba661dc89bfbfa08bea34de3694aae5a1c781e986ffa76a1c8736a1b

Observation 852239dc-55ea-4a60-bd60-6d5db5a9e88b · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.111892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.111892Z digest=sha256:eee63c54b6ced18309adacf16f0477afb9b32703f032734179e415b53e9c19b1

Observation 47499b1b-0fdc-41be-b7ed-be6849a3b848 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.164246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.164246Z digest=sha256:713eaa4a36cc81d567e293fefab724b617d117da73590eaaec47204f34b823c6

Observation eb008823-5efe-47a6-9bb5-5b36aaab89fe · outbound

This paper cites Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.959204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.204264Z digest=sha256:43382bdcfb0740a312298ec73596efb369630943dba746ff9505ddcebec36194

Observation c53d914b-6a9f-429c-97e9-e5b27852a45f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function RLHF Workflow: From Reward Modeling to Online RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.243909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.243909Z digest=sha256:fe6b174a84b93fe32561ba1d63e0f073429610ecca3a1608cc37bc38da4dee5a

Observation 72383b3d-0aa4-47c7-952e-3f082282bcbf · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.818904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.281397Z digest=sha256:ba4da6514b7383f137e650133c928fea51354c593391bdf3fc8c5d4f17d61b8b

Observation faf8bccb-9cd4-44ef-92c4-992109603475 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.316647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.316647Z digest=sha256:4dfc6f2f215d3dc46f82a528a41dc31b8dc8de88ff1b0b0a505431f775c2c508

Observation 4e900513-ea97-4973-99b3-e7ed056381b2 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.642253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.349331Z digest=sha256:022c908bd5589d8c1c7e73a8f3233c5d07900a42c52a673b5effc4b93ad0a8ce

Observation fde34b94-7148-45ca-9177-fabbbd68ac2d · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dueling rl: Reinforcement learning with trajectory preferences

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.447430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.386029Z digest=sha256:29c0bf69a5b82047068c9a932c907dda3a9555f0b4f8be14e85413c7b7312639

Observation d3376d72-f218-4f16-979d-2b5ac0b7f007 · outbound

This paper cites Lee, and Wen Sun.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, and Wen Sun

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.420179Z digest=sha256:eaea291e54ed6b10063ac899eceedaea231ac208c46a24195c6a1ce60b03c87b

Observation c27ff57d-d834-4f70-9e04-45b415b4032f · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.078126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.440411Z digest=sha256:4db1d132640fdacf46b38175ec65e964beb092362bad1b33f04f39905cfe39c6

Observation cb357bfe-a7f0-42b9-8cf9-09477257b24f · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.936211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.473666Z digest=sha256:eb6460708a8d57abe3948fdf2221ed8f6bb12b6009a8a91cf4578b6157df5cf6

Observation c998c1d4-4d99-46c4-bd9a-40bf5731e69c · outbound

This paper cites Provably feedback-efficient reinforcement learning via active reward learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably feedback-efficient reinforcement learning via active reward learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.811129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.509008Z digest=sha256:597a274bf47ebccbd7914b9037118fd24e538f4e002134bf6ab2e3262711e2d0

Observation 0233317e-db04-4053-aae1-f918664243a0 · outbound

This paper cites Making RL with preference-based feedback efficient via randomization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Making RL with preference-based feedback efficient via randomization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.514997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.514997Z digest=sha256:839f9ddfb962a44524e2aa4d58e805143d4dff5d796cc69a69ba2f6e461d76bf

Observation c477f12d-8642-45e5-9f7c-6d8ee20d0c30 · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.572546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.572546Z digest=sha256:b3b9227dfeb7ec985b9952b830416c0830fae8b154ebce52c81bcdf05c929bf2

Observation dd8c175d-60e3-4c22-a873-2d522bf430d3 · outbound

This paper cites PARL : A unified framework for policy alignment in reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function PARL : A unified framework for policy alignment in reinforcement learning from human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.698009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.622186Z digest=sha256:56dc9afd63b2e4d7cb370ff3b0365724e6f16ac0db79a35e8ed0d0681d14c371

Observation 5e79b9c9-f0f6-461b-8c8b-9946d45997bc · outbound

This paper cites A Theoretical Framework for Partially Observed Reward-States in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A Theoretical Framework for Partially Observed Reward-States in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.676261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.676261Z digest=sha256:88c5b0f200b3bf427b7bfb28e0bc4aa32b5cda2f942d00ac1c26700487873ce8

Observation ed9d7dcf-c55c-4cba-b980-2322526f4b1e · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.714921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.714921Z digest=sha256:00ccc32f8539afa7e8df469ccf24a7482a02602fed7961b95eb8eecf2453e641

Observation 6690ffa4-8c86-4f9f-92e0-78b78d683504 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based reinforcement learning with finite-time guarantees

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.768418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.768418Z digest=sha256:0253fa0c1bac0e73f788289915c8fbae4a505bfeed8e3a7716a3943ecdbb6754

Observation a625c456-97f1-4a52-8f94-1a6d6771ea27 · outbound

This paper cites Zeroth-order optimization meets human feedback: Provable learning via ranking oracles.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order optimization meets human feedback: Provable learning via ranking oracles

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.474693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.828580Z digest=sha256:e679064fd4890ccc8978bf8d15b2e925d33ca4aa463736e3fc8cd027db96f0df

Observation 7340b7ca-af8f-47b1-826f-1103e648a73a · outbound

This paper cites Interactively optimizing information retrieval systems as a dueling bandits problem.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Interactively optimizing information retrieval systems as a dueling bandits problem

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.869262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.869262Z digest=sha256:6a3e8981ef8768adf440073965366822dd566a539e293b57c29070c592f6ae99

Observation 37158a94-e721-48dc-9a56-9d4451dae6fc · outbound

This paper cites Beat the mean bandit.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Beat the mean bandit

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.339023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.903719Z digest=sha256:707c008095bf0c8219774b2a85cf2d9c460581f779b55fd035f34e6293be2e8a

Observation 36498c12-90fc-466f-a433-0d8d4b0b5411 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Generalized preference optimization: A unified approach to offline alignment

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.260964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.939206Z digest=sha256:68e6ebf1fb932c34e44b3d4607a5eb83e93a83134c70e7d28ed088a4916edb55

Observation e48c9491-b216-4f2c-8dce-4cfe88ce4f50 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.972232Z digest=sha256:3676cb558e9c38a629cb16549807fc67dd3830aad4ff53d60f97d914bf167cc7

Observation 5f9fc27b-a6fa-45d6-aceb-bc102ab9b3c0 · outbound

This paper cites A theoretical analysis of nash learning from human feedback under general kl-regularized preference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A theoretical analysis of nash learning from human feedback under general kl-regularized preference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.005951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.994701Z digest=sha256:bf137ff9c8c92bbbe7be5cdb3ece345f2dfd8e19bc7785d6fc74990d89a3deec

Observation 5cd7e6c0-792b-4999-8f7b-1ef62dfc01ff · outbound

This paper cites Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.045330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.045330Z digest=sha256:57d041341cf2f266bfdfc4c0b977ba76166f6032396887291957e986011932d6

Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.109193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.109193Z digest=sha256:729859cbe728d14deee32f8998844af7a442973d4f224ad309678d8995f20c1e

Observation 7d609393-a485-48e9-831d-0d2108ca7be1 · outbound

This paper cites Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.166568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.166568Z digest=sha256:fb8ed7902409733eae06e5155e06d72caf1e502f9af57ebfa245a0893bcc0b4d

Observation d1155787-5804-4de0-aab3-ff4b133c206a · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Stochastic first-and zeroth-order methods for nonconvex stochastic programming

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.251100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.251100Z digest=sha256:8cc03cea19c1a411ee94ea20b02a02fb4291abc9b02baf0f2bb81db5a1626075

Observation b2cbb1d1-4a3a-4b1a-98d2-495f1ea29c8a · outbound

This paper cites Random gradient-free minimization of convex functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random gradient-free minimization of convex functions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.327923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.327923Z digest=sha256:d49ea926c543cafe70b4168332b1cd626482800defac8bcd8aa69678f3ac5ef3

Observation 4b2c1b49-9dbc-4dff-9042-fe52767810be · outbound

This paper cites Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.876255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.392308Z digest=sha256:53a0f1c8bf6a03bcb0ffc745de12a4d6beeba90b9051b74cecd60e412a3d2ba2

Observation 883be79c-b080-49d3-b44d-2ed6fc91247e · outbound

This paper cites Zeroth-order stochastic variance reduction for nonconvex optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order stochastic variance reduction for nonconvex optimization

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.744327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.448455Z digest=sha256:e0b01cd2b32b0bdaf60098ed8659f2948c0d33ab9a5cedce3e586635c6191f6c

Observation 971ae2dd-6d2d-4f65-ab50-7b305d4b7b0f · outbound

This paper cites A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.522274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.553872Z digest=sha256:b07c669e382ce4546bcfddee3e52134e66697290ee6e33aac31c75bfb1be6c9b

Observation 6aa4c8eb-31e8-4123-8fd8-86f3fa3e4b86 · outbound

This paper cites On the information-adaptive variants of the admm: an iteration complexity perspective.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the information-adaptive variants of the admm: an iteration complexity perspective

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.280656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.628579Z digest=sha256:ce7133e1f54a4b14c44263117349b9866c0e23770116bd160e01458d3a6662b8

Observation 32a34c98-7b60-4904-a31b-ac0bc3ab8a24 · outbound

This paper cites Fine-tuning language models with just forward passes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Fine-tuning language models with just forward passes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.029182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.679574Z digest=sha256:6c90fe9efb2c724ec4ed0e6f9ff1f53d98b98c6d018b19f1ededfc577fb51028

Observation 07876c0c-5d52-47f5-8453-efbe4022bbcc · outbound

This paper cites Evolutionsstrategie.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolutionsstrategie

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.777335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.753674Z digest=sha256:3d8f027affdc22f4e3da315bb975e2510f2b60124311a17f876282f6bb231ff1

Observation f1de5ea6-2bb9-4c81-a9dd-ca48cfc2ac06 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.770317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.770317Z digest=sha256:eb1557c35a8102104b12c75a8c503b3fe17a8c78af68b6e3acd063ab979babda

Observation 284ff06e-fe75-4924-ad8d-f9a1aeeae412 · outbound

This paper cites Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.519056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.834484Z digest=sha256:78a8d85d7f66a15a2baa8ef6fec2ee7e8f5b17714f6ac7fb6b8473d7b8e1e6c7

Observation b1c5defc-e852-4e0f-9cd1-8710f494396f · outbound

This paper cites o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.225652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.893600Z digest=sha256:d90801ce96e476bfe096f7c5b65921d9668fd8b26dfb79a0c0eb140722a68d99

Observation 4514c777-95fb-45ff-8643-5349d481cee0 · outbound

This paper cites Preference-based policy learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based policy learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.013767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.926887Z digest=sha256:f3a8566be545432721eb03d0b15a4dcdcbca1e8ae067ce607d3b538104c5fbb0

Observation c1086b22-e42f-4c98-8177-ce41ecf25690 · outbound

This paper cites sign SGD : Compressed optimisation for non-convex problems.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD : Compressed optimisation for non-convex problems

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.771367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.993052Z digest=sha256:e5afb7d8d1b77019b4e3eefbb783e7f3cfda95cde4731f6681fcdeaf9796f283

Observation 4624754b-48f7-4449-9214-3f6e92dee3e3 · outbound

This paper cites sign SGD via zeroth-order oracle.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD via zeroth-order oracle

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.561389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.061239Z digest=sha256:b03325c9dacfa3fc712dd679e979f14a48e0518514f29876e212a90c28bd1817

Observation 1d23d772-18ad-4127-855e-212208825ea0 · outbound

This paper cites Thurstone.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Thurstone

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.425354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.109276Z digest=sha256:2aa6da4d4029e6b08a66a8278e8265c473f694f95aefb9a583acdd671ac936f4

Observation 6f98a202-ef5e-4072-ac1b-49f50c2593b3 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:17.289639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.174397Z digest=sha256:233ee85f07de440f909e56f1faae6b66bfaee2204543e8e5d1ce1cd928fde74e

Observation eea2b3cf-b74c-4cc7-b4b3-701a9358fa0b · outbound

This paper cites Reddi, Satyen Kale, and Sanjiv Kumar.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reddi, Satyen Kale, and Sanjiv Kumar

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.228756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.228756Z digest=sha256:51c7f3d24d0077aa14c71cd5857d8d78e0ac89eac5c5ffcc8e45ece5112d2935

Observation a778a116-10e9-461d-acc1-169b01716d4a · outbound

This paper cites Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.145625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.262179Z digest=sha256:eb896c3b0499c0fdb4edff9becad1663027d1ac4e8a2d4e3216e7fc8f570202d

Observation 35a2cadb-2800-410f-b344-516972a410a6 · outbound

This paper cites Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.927021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.331068Z digest=sha256:7fbc790aebfefeda4547d0dffb39af6c05144470fc2da2555d404dc4b9c163b0

Observation ebf7a969-1673-4a56-a423-dfe76fb3dcca · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably efficient reinforcement learning with linear function approximation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.706697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.383213Z digest=sha256:d3b1d735c0b9054dd9c9f581c91e43971767b889f233610d1560b5149bb8248b

Observation c2d9b706-27a8-433f-a24b-df58a46d0be7 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function High-dimensional probability: An introduction with applications in data science, volume 47

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.443901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.443901Z digest=sha256:3e3af46bac9a09baf1285e546e08799c702d17e07782197654c418b426d4e7cd

Observation a1e85407-5e5e-4403-9f38-41b9ca43909b · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:16.605630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.515429Z digest=sha256:7f905134be80673bd749f374787b5855b9f91f4f0d7dc11e6723ea6772b2f7e2

Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.560054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.560054Z digest=sha256:9a21c53c293cc68eb1ae588ca21fa3f5e2fc3f04a7ea657ae9a22e4e55c76b4a

Observation 7e5ac7ae-8755-4084-8948-5a91bd1a59c4 · outbound

This paper cites Convex Optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Convex Optimization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.471296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.630602Z digest=sha256:4ba55ca834fc4485dd593c7dac9506465a5b7688b92be3bb0b44594889d559e0

Observation ca26ef71-734b-4f3c-ad42-300c2dcf7080 · outbound

This paper cites Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.361093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.682673Z digest=sha256:9c443e75ca8244b4d117bdc07d87fbca74ed6e884f60b36659d475add754481e

Observation 9d2fc4c2-5b85-45fa-8a99-90eafaa3e14a · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the global convergence rates of softmax policy gradient methods

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.252160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.714867Z digest=sha256:c1522d8fddd239090b3d992366f5d42be829dce6ec1c7a8cc1533ce1a4421a22

Pith citing papers

Observation c6d37a12-b035-4679-b920-5dea3a8d1499 · inbound

Efficient Federated RLHF via Zeroth-Order Policy Optimization cites this paper.

Efficient Federated RLHF via Zeroth-Order Policy Optimization Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-21T00:19:49.870688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:48:15.394329Z digest=sha256:4757a1e87b1bb8079ac469b3c7caa8226ff5b231112102a3d675c5db634b6e98