Pith. sign in

Paper Citation Record · LEDGER

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

As of 10 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2506.03066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03066 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:15.714867Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:48:15.394329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:35:19.307998Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9703832-2747-4b73-8c47-4467da086c5e · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning: An introduction, volume 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.114854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.114854Z digest=sha256:5629df9df928ad4622163051e34d8221b72a3de24324e33147f66c03968e30f5

Observation 70fe8ede-870e-4e44-ae03-67b1798c8153 · outbound

This paper cites Controlled experiments on the web: survey and practical guide.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Controlled experiments on the web: survey and practical guide

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.203695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.203695Z digest=sha256:f4776dfcbecaec12b6ee49f1219aa4dc98b6ce48433af2aff51b59629c97dda9

Observation 2de9bc8e-27eb-4749-95be-5ba5d2ac8ec5 · outbound

This paper cites Deep reinforcement learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Deep reinforcement learning from human preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.296587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.296587Z digest=sha256:ad24401d29cbf48f5f38993b514f1afdce99e448b821a76e463743212ad3eeec

Observation 2916cde9-d8b5-465c-97bc-8968f3927110 · outbound

This paper cites Al Sallab, Senthil Yogamani, and Patrick Pérez.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Al Sallab, Senthil Yogamani, and Patrick Pérez

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.405514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.405514Z digest=sha256:c081aaa7c2fca200e491681c5cca89bf145728bdd1f94d8ad635089f1962e0c7

Observation 739bf0c4-c7c9-4cd1-bc45-3be231cf17ea · outbound

This paper cites Training language models to follow instructions with human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.525911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.525911Z digest=sha256:9af7ed9ba2592e27e19a350bde8f688391eb1f695b55ca4633f27f975a8439c9

Observation ca01d7dd-fa85-480d-9975-8a0b310cb5f4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.608779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.608779Z digest=sha256:eae29c54ae69e199e4039e8e6dd6e6d56245faa9de030c3f39b44807acb504ce

Observation 99c7fe35-9628-4b78-b4ab-abb6f85799f3 · outbound

This paper cites Dynamic programming and stochastic control processes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dynamic programming and stochastic control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.716180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.716180Z digest=sha256:4162460dee379df5aeb482ab37ccb276002583249ba461114da99f9d9e750b16

Observation 192aba53-7d83-4907-ad1b-7861438a4a85 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Markov decision processes: discrete stochastic dynamic programming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.827799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.827799Z digest=sha256:fe2e85d71f1cd840f1253daf8924ab3985d9388875f3d900e58658d2ff4e4f61

Observation 5d6406f9-6fc8-452a-874a-8d9f581f79e9 · outbound

This paper cites Inverse reward design.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Inverse reward design

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.891470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.891470Z digest=sha256:17f03f10fed48fbd39c5ad51ae9cf587c3b9ca7e3712e162a4cfbdd27c1fe9c5

Observation de54f59e-6712-4905-8a83-7fe6714d6b9d · outbound

This paper cites Reward Design with Language Models.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reward Design with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:10.973403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:10.973403Z digest=sha256:5ea7559462d5da75f99d191a2346620c377468befdc5118ab83ac4607c593fb6

Observation 6f2ccb55-f03d-4c47-9196-aa7b47ef13c6 · outbound

This paper cites Defining and characterizing reward gaming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Defining and characterizing reward gaming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.056676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.056676Z digest=sha256:45bf061e6101db215e72ccd9eb6056a4d2edfa86232b60685d54d4ea40f70b59

Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.242520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.242520Z digest=sha256:75fc4aec27c9c65ace817c37df4f54cf8b119030156a76d06faec24f858abc2d

Observation 31d800c5-7bad-4ed5-8efc-6e2d80ecf082 · outbound

This paper cites GPT-4o System Card.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.353435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.353435Z digest=sha256:5ba1ca97856d324211ec0e6a4a1c6fc2fc76ba54404bbe301686c2e0ec8dcc84

Observation 1534fe5b-74b5-4f5f-a2bb-9c1b5495e071 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.439190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.439190Z digest=sha256:ec3c0ffa0d65670be245ca7e38f68d4d24ccb2fdd5bbb3b40e50df9f8f416310

Observation e5c0ff52-d75e-47a2-bb45-e07577dc1094 · outbound

This paper cites Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.434223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.554482Z digest=sha256:88f47e743f94958857003845f014375844ef2da25c83a1ded1d3b6a622464158

Observation 5def0db4-fbcf-475d-b4d3-3d100ea593d7 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Proximal policy optimization algorithms, 2017

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.665749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.665749Z digest=sha256:521eba14878dbf357d60b0c0bfb34f2adc5658865ff8c82ca707ddb1aee354c4

Observation a80d8297-66c6-4289-8a23-988584021b24 · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Open problems and fundamental limitations of reinforcement learning from human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.322900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.755552Z digest=sha256:81e3876de61564879aec0cc985d359a6f66ade14a5ed16964df75112b9d11f1a

Observation 2b6d06d7-162f-4354-abc4-664102270b38 · outbound

This paper cites Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, and Tianlong Chen

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:24.164237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.793743Z digest=sha256:4560f8c1812d0acc484b86b9b73c238c514a9b2262be9e6f5002f3808d79fb14

Observation 4656b496-a7a3-473d-9590-2e947e847154 · outbound

This paper cites From r to q^ * : Your language model is secretly a q-function.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function From r to q^ * : Your language model is secretly a q-function

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.963362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.799620Z digest=sha256:065b88251925f79e8125b1407fde83d93235dd44c152424dbc63891e0f11b359

Observation b76e2247-cc93-4d0b-841a-15c87dfc1c31 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:23.787862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.839237Z digest=sha256:aa72e7f6e50110b9b2f69f8fecb3de355d4f1d91d9e74ffac89295674156ec93

Observation 2105833b-782a-4fa4-b5d9-68dbaa0f8459 · outbound

This paper cites Random utility theory for social choice.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random utility theory for social choice

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.906112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.906112Z digest=sha256:acda35d2a6c32e2dfc8b3ead8d6dc796f5b4b981f28437be8f2f8b67411827d6

Observation 7519c8fe-e391-4f82-b614-f4e0aa2e2cd9 · outbound

This paper cites Modeling ordered choices: A primer, 2010.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Modeling ordered choices: A primer, 2010

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.658611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:11.971612Z digest=sha256:f095b3eef4e0078dffc4631d2a76ba25e673a1f7806456866a9291b056df285b

Observation a1d649d2-8446-4572-87bf-486af6d3ef41 · outbound

This paper cites Econometric analysis 4th edition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Econometric analysis 4th edition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.504855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.030921Z digest=sha256:b6f4076b66aac3a70e224e3c91b6607d84e5f3dcc92c5fd58a679ddf01628c6c

Observation 571facf2-89e0-497e-8702-1fd5b464a369 · outbound

This paper cites Sensory evaluation of food: principles and practices.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation of food: principles and practices

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.282474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.101142Z digest=sha256:a89bdcd7510014423771a604e245bb4fe0057b40920c902706d1edd682ae7c6a

Observation 097214b3-cd0f-4089-a3e8-c302ce0d7ae6 · outbound

This paper cites Sensory evaluation techniques.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sensory evaluation techniques

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:23.069453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.166328Z digest=sha256:91cd0a5f88be918f001279cd997fca46b7dde73f7b465985f082cf5b49ea1406

Observation 0b3b9f97-ef92-479c-bb26-deb82ab8b5de · outbound

This paper cites Nash learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Nash learning from human feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.880574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.215572Z digest=sha256:894fad2a8d443206b6eb16a6b7eb88c083564aff6928b3a5c9abc8903df21a09

Observation 9cca57ca-66e0-4d12-bbd7-2c97d1726096 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A general theoretical paradigm to understand learning from human preferences

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.715947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.318773Z digest=sha256:9bb1a23f87c0a66ffc26d810b17c393b5d87ad229947bb9f33b97aed8c567343

Observation f44cfb7f-350f-4cc4-ae86-c53b6055920e · outbound

This paper cites Models of human preference for learning reward functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Models of human preference for learning reward functions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.557457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.395294Z digest=sha256:d53089f7feab9b8ff2123063a2e85554377825a2e36b6b4f91941445af4d2f23

Observation 1120e8b4-c18f-4faa-8d3b-7be8f969d57a · outbound

This paper cites Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement learning from human feedback without reward inference: Model-free algorithm and instance-dependent analysis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.399902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.569486Z digest=sha256:a2290bd44706a1b3b0ccec7076dfb1887bcbc57b2c2b35dec28f99aac7e075b8

Observation 16fc529c-6624-49cd-89d7-04165b8a6dc4 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based online learning with dueling bandits: A survey

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.244253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.737974Z digest=sha256:e06ed127aff9a8e7e3a5cb5c0c67d79dca18048bb72fd53e56bbd2029d8359c2

Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.854603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.854603Z digest=sha256:0955a210c500e2d24c4c98241c639e34dba12f86d1f07acc700eebbe7a5c84c3

Observation 155b393a-2365-4814-8dc9-7b9c21daf1bc · outbound

This paper cites A survey of reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A survey of reinforcement learning from human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.936227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.936227Z digest=sha256:53968762338d944177bd5c59b868ea167fd4aa2a9c390ab0dd62fe778ce3ea36

Observation ce96e202-23b6-4bcb-aa22-30374648414c · outbound

This paper cites Scaling laws for reward model overoptimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling laws for reward model overoptimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:22.106083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:12.957795Z digest=sha256:0f24983e9fefefe323b101b2f9da64c76b227093c0347a8ca89e571230b4af58

Observation eb2419d6-f622-4eaa-b050-50fb1789f50c · outbound

This paper cites Model-free preference-based reinforcement learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Model-free preference-based reinforcement learning

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T11:20:15.864764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.050288Z digest=sha256:92b461dc770ddb5ad68fb3cc489cd99009867a4a533de1bb0130f88aa5beee0a

Observation 852239dc-55ea-4a60-bd60-6d5db5a9e88b · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.111892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.111892Z digest=sha256:cfd419b301bb750c710422d88fc3a57cfefddb1d7756b654d4c89bc43a8080f6

Observation 47499b1b-0fdc-41be-b7ed-be6849a3b848 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.164246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.164246Z digest=sha256:8ca2e57603cbf79a608a797ed47400d4a3c14de2d788aa689aec71d5db67e54e

Observation eb008823-5efe-47a6-9bb5-5b36aaab89fe · outbound

This paper cites Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.959204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.204264Z digest=sha256:f94b91086285a4f880e87e91876a68970286f9be576e7211dfaefedd0afd11a9

Observation c53d914b-6a9f-429c-97e9-e5b27852a45f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function RLHF Workflow: From Reward Modeling to Online RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.243909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.243909Z digest=sha256:9f37f36aaa838c74b01ea81b4b9e97007025530ed9fdee91b78756a2a0464c98

Observation 72383b3d-0aa4-47c7-952e-3f082282bcbf · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.818904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.281397Z digest=sha256:36996de1a14952dc3a6bd043c8007bc61892aa14979a618b63f506e861fef8f8

Observation faf8bccb-9cd4-44ef-92c4-992109603475 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.316647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.316647Z digest=sha256:9492c770341551b35ed2a006e2449d665473005bd02309eb8b2ed61223e388be

Observation 4e900513-ea97-4973-99b3-e7ed056381b2 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.642253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.349331Z digest=sha256:d48aca671c7b4ae852c339b168da06ff3ff3a20e18edc8b74b22efab7144b617

Observation fde34b94-7148-45ca-9177-fabbbd68ac2d · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Dueling rl: Reinforcement learning with trajectory preferences

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.447430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.386029Z digest=sha256:108eb28a830dcbcfe709d70382fd3b9e86d592891c2a007490c65c2c384b8f0c

Observation d3376d72-f218-4f16-979d-2b5ac0b7f007 · outbound

This paper cites Lee, and Wen Sun.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Lee, and Wen Sun

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:21.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.420179Z digest=sha256:0d11f5cc621e73806f089ad78527136dec9eb203b361b03d0c7f90498ea9bdf9

Observation c27ff57d-d834-4f70-9e04-45b415b4032f · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:21.078126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.440411Z digest=sha256:5c7722486c5545a2b84ade152492a880a22b7fe6317a6685b3711d36df82adb8

Observation cb357bfe-a7f0-42b9-8cf9-09477257b24f · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.936211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.473666Z digest=sha256:7252b7aca11e7c9e01be6a73c1064ab4595f7671fef6ef3589063901e127356a

Observation c998c1d4-4d99-46c4-bd9a-40bf5731e69c · outbound

This paper cites Provably feedback-efficient reinforcement learning via active reward learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably feedback-efficient reinforcement learning via active reward learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.811129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.509008Z digest=sha256:dc038b4d60f3f9b7236a6dab35837ddb9a5a20b62302f925bdfbe08959f8f210

Observation 0233317e-db04-4053-aae1-f918664243a0 · outbound

This paper cites Making RL with preference-based feedback efficient via randomization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Making RL with preference-based feedback efficient via randomization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.514997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.514997Z digest=sha256:16926dda51b0a17658e03f52e02092441a7865532e4c8461cb8c4798c0831920

Observation c477f12d-8642-45e5-9f7c-6d8ee20d0c30 · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.572546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.572546Z digest=sha256:c46af3987b2b1201f0a128d40bb556e3b248d12313a5f2eb7031a1f7fb9ff951

Observation dd8c175d-60e3-4c22-a873-2d522bf430d3 · outbound

This paper cites PARL : A unified framework for policy alignment in reinforcement learning from human feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function PARL : A unified framework for policy alignment in reinforcement learning from human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.698009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.622186Z digest=sha256:1af93389239e32e7e02f909af355167ffd157d22826da9c706d440e0a46db47c

Observation 5e79b9c9-f0f6-461b-8c8b-9946d45997bc · outbound

This paper cites A Theoretical Framework for Partially Observed Reward-States in RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A Theoretical Framework for Partially Observed Reward-States in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.676261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.676261Z digest=sha256:997384751ecb9de61c4f30eb75c24ef14ee9909a3dae1c603bdf3cb5b92aa9fd

Observation ed9d7dcf-c55c-4cba-b980-2322526f4b1e · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.714921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.714921Z digest=sha256:5b4a9ba64318cdbc78393985d0b434c55e3f116e45993186a30a96dc20c1f773

Observation 6690ffa4-8c86-4f9f-92e0-78b78d683504 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based reinforcement learning with finite-time guarantees

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.768418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.768418Z digest=sha256:02325701fe0fd6831df7ef8a265f39fe556724873b3009016e688f64492d8042

Observation a625c456-97f1-4a52-8f94-1a6d6771ea27 · outbound

This paper cites Zeroth-order optimization meets human feedback: Provable learning via ranking oracles.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order optimization meets human feedback: Provable learning via ranking oracles

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.474693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.828580Z digest=sha256:f355bb8266d92075a7204541bb156725bfeafddaa975d2ea756ed82eb8133aee

Observation 7340b7ca-af8f-47b1-826f-1103e648a73a · outbound

This paper cites Interactively optimizing information retrieval systems as a dueling bandits problem.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Interactively optimizing information retrieval systems as a dueling bandits problem

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.869262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.869262Z digest=sha256:663054a5336fb92b790f678d60fbefc11d22a93685a1e16ca53a5dc3c243a287

Observation 37158a94-e721-48dc-9a56-9d4451dae6fc · outbound

This paper cites Beat the mean bandit.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Beat the mean bandit

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.339023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.903719Z digest=sha256:da376144033636176b8231b1b6a5eae442c1b2df17a342092749ccd3d3b04bbb

Observation 36498c12-90fc-466f-a433-0d8d4b0b5411 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Generalized preference optimization: A unified approach to offline alignment

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.260964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.939206Z digest=sha256:e630778fa676f0140a5e7ac2ab672e92d86b16f18e310c190283e9b8ef198928

Observation e48c9491-b216-4f2c-8dce-4cfe88ce4f50 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.972232Z digest=sha256:4f478b34760d18cefbbe65e6b2ec8b2a3f2fa4162cb0463960f0f9fc3eb112a8

Observation 5f9fc27b-a6fa-45d6-aceb-bc102ab9b3c0 · outbound

This paper cites A theoretical analysis of nash learning from human feedback under general kl-regularized preference.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A theoretical analysis of nash learning from human feedback under general kl-regularized preference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:20.005951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:13.994701Z digest=sha256:c5c9a6d2b59ba6fdda437efad54bf65ae684cd3decc621f7a1fb69bffdc1d806

Observation 5cd7e6c0-792b-4999-8f7b-1ef62dfc01ff · outbound

This paper cites Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.045330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.045330Z digest=sha256:bdeab0ee1a449ce72bb0dad360df7fa47fe21d3ef2ff19fa587f8cfe7665435c

Observation e7f15eb2-1f01-4a41-b31c-5f4fe31cdcbc · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.109193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.109193Z digest=sha256:8c3d94e76e89918b5b0a570b17efc216496ce37125041e15af16c9c0c811b427

Observation 7d609393-a485-48e9-831d-0d2108ca7be1 · outbound

This paper cites Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.166568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.166568Z digest=sha256:e7d60b2d8ef59ed796237cbcb491c2dfec90f49318476296626df674d7bcaa21

Observation d1155787-5804-4de0-aab3-ff4b133c206a · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Stochastic first-and zeroth-order methods for nonconvex stochastic programming

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.251100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.251100Z digest=sha256:ff431ce4c6c66c36bb409145e6df9b1379e8cf8a3215d58e0704a11c0e11d749

Observation b2cbb1d1-4a3a-4b1a-98d2-495f1ea29c8a · outbound

This paper cites Random gradient-free minimization of convex functions.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Random gradient-free minimization of convex functions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.327923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.327923Z digest=sha256:ce16b1f33b03ccfcfa24762533ee98d5b10ca1574101b9feb2f3c0e7f901d57e

Observation 4b2c1b49-9dbc-4dff-9042-fe52767810be · outbound

This paper cites Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.876255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.392308Z digest=sha256:28129f4602aed03640451eafdb818e7b4f912ba0bb310d517db95304309eb0b7

Observation 883be79c-b080-49d3-b44d-2ed6fc91247e · outbound

This paper cites Zeroth-order stochastic variance reduction for nonconvex optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Zeroth-order stochastic variance reduction for nonconvex optimization

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.744327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.448455Z digest=sha256:3f2509b312c6c8e1137219ad0f722e83329c7953b65a036d39f025d79886fe33

Observation 971ae2dd-6d2d-4f65-ab50-7b305d4b7b0f · outbound

This paper cites A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.522274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.553872Z digest=sha256:c315f583229f9dc76c246caa6e9b86d479d42b083314e2ea6db1da04a9edaa1f

Observation 6aa4c8eb-31e8-4123-8fd8-86f3fa3e4b86 · outbound

This paper cites On the information-adaptive variants of the admm: an iteration complexity perspective.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the information-adaptive variants of the admm: an iteration complexity perspective

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.280656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.628579Z digest=sha256:72f02642cc3ea2a23c4c1deae7fa73d7e0005b04818c053326dbb5388aa61dba

Observation 32a34c98-7b60-4904-a31b-ac0bc3ab8a24 · outbound

This paper cites Fine-tuning language models with just forward passes.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Fine-tuning language models with just forward passes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:19.029182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.679574Z digest=sha256:51c3d3f370a011db67f454f14258de810141851c50c938276ef28126977e889c

Observation 07876c0c-5d52-47f5-8453-efbe4022bbcc · outbound

This paper cites Evolutionsstrategie.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolutionsstrategie

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.777335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.753674Z digest=sha256:74f167432ce4a611b388d630a289f6fcd720f40c00826ac1ea46771f634498ff

Observation f1de5ea6-2bb9-4c81-a9dd-ca48cfc2ac06 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:14.770317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:14.770317Z digest=sha256:adb19fd36fbfb25849f004c901d6c16801edc518a200fa93f3e2137be7a9b8df

Observation 284ff06e-fe75-4924-ad8d-f9a1aeeae412 · outbound

This paper cites Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.519056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.834484Z digest=sha256:42f302aa8c4e483c01df8516e3d67d09cd511cb18d2b1bba6ea05cfd8c0c2ce5

Observation b1c5defc-e852-4e0f-9cd1-8710f494396f · outbound

This paper cites o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function o r \'e nyi, Paul Weng, Weiwei Cheng, and Eyke H \

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.225652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.893600Z digest=sha256:57c85dfb3dc7fffba3fd9ed0b83a2f14a7b6131b7cbf38a924ba0ff0463bb174

Observation 4514c777-95fb-45ff-8643-5349d481cee0 · outbound

This paper cites Preference-based policy learning.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Preference-based policy learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:18.013767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.926887Z digest=sha256:ea94c6c3ac5e91bb16555518b6375e8289f443ae64b7c9ba141596244e6f3b48

Observation c1086b22-e42f-4c98-8177-ce41ecf25690 · outbound

This paper cites sign SGD : Compressed optimisation for non-convex problems.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD : Compressed optimisation for non-convex problems

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.771367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:14.993052Z digest=sha256:aadda670861f7b3cd597f7cb04b5607a18e580917cb607f702a6bbd7466aa59f

Observation 4624754b-48f7-4449-9214-3f6e92dee3e3 · outbound

This paper cites sign SGD via zeroth-order oracle.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function sign SGD via zeroth-order oracle

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.561389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.061239Z digest=sha256:4715a8bf00a48c77cdb12fcaa9172b19cee2c32e4419b21655eb618514a1230b

Observation 1d23d772-18ad-4127-855e-212208825ea0 · outbound

This paper cites Thurstone.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Thurstone

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.425354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.109276Z digest=sha256:d94d0236c375aa3d5dad6b73c95cbc706ff08cbd0a6bdc2682d40e2112844a77

Observation 6f98a202-ef5e-4072-ac1b-49f50c2593b3 · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:17.289639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.174397Z digest=sha256:ea607df95bdd82b62a01f9c4a4a713935a0ee21f680e3e657a77c94eac97bdff

Observation eea2b3cf-b74c-4cc7-b4b3-701a9358fa0b · outbound

This paper cites Reddi, Satyen Kale, and Sanjiv Kumar.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reddi, Satyen Kale, and Sanjiv Kumar

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.228756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.228756Z digest=sha256:35026732af982d41f3324bb2f05a3005316035c526b89c76dac1fe16b383f656

Observation a778a116-10e9-461d-acc1-169b01716d4a · outbound

This paper cites Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Online rl in linearly q^ -realizable mdps is as easy as in linear mdps if you learn what to ignore

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:17.145625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.262179Z digest=sha256:698ece0355929c6e85d3e8af9ac3dea50bbfd1ea30be5293f870eb8c45d4b9dd

Observation 35a2cadb-2800-410f-b344-516972a410a6 · outbound

This paper cites Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.927021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.331068Z digest=sha256:c5ede88b4171b4e1526d1db5914e8da2e481eb55ca605f00c4e8342e177e2817

Observation ebf7a969-1673-4a56-a423-dfe76fb3dcca · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Provably efficient reinforcement learning with linear function approximation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.706697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.383213Z digest=sha256:023432698e6b58c3c714048321aa9109a7b2e2f1c9518348afd6f7882351de4e

Observation c2d9b706-27a8-433f-a24b-df58a46d0be7 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function High-dimensional probability: An introduction with applications in data science, volume 47

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.443901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.443901Z digest=sha256:8bf911fc4689b7db971a30c956cfad2daba7397dc5f676763a8aaa595d0fde70

Observation a1e85407-5e5e-4403-9f38-41b9ca43909b · outbound

This paper cites an unresolved cited work.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:20:16.605630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.515429Z digest=sha256:43d0fbd185b12bd3a9f64e9ef66748e98994e790f50b7dd7634995ccd37a7822

Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:15.560054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:15.560054Z digest=sha256:c865e28b8ded9a9c8f66180a7a9b19225a65e2fa63ef72a6e499e7ca32f4b2e0

Observation 7e5ac7ae-8755-4084-8948-5a91bd1a59c4 · outbound

This paper cites Convex Optimization.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Convex Optimization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.471296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.630602Z digest=sha256:b595657d9a54138d9721f6ffabccdd97e91ad99303ce79217261a1097a899bfa

Observation ca26ef71-734b-4f3c-ad42-300c2dcf7080 · outbound

This paper cites Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.361093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.682673Z digest=sha256:9023c5c7b22134833addfb6d69f8904947ba1bd816057b9ee3bbf1b9d0e72d6a

Observation 9d2fc4c2-5b85-45fa-8a99-90eafaa3e14a · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function On the global convergence rates of softmax policy gradient methods

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:16.252160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:20:15.714867Z digest=sha256:37e187bd4b89f876fa1aa1efd4d5b0d0dc7188f0b706484bf3757069503308fd

Pith citing papers

Observation c6d37a12-b035-4679-b920-5dea3a8d1499 · inbound

Efficient Federated RLHF via Zeroth-Order Policy Optimization cites this paper.

Efficient Federated RLHF via Zeroth-Order Policy Optimization Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-21T00:19:49.870688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:48:15.394329Z digest=sha256:a195a60f14543ef5f630661030882cd6d7b97cfb45b2d06e71a794cf668ebd87