Pith. sign in

Paper Citation Record · LEDGER

Online Preference Alignment for Language Models via Count-based Exploration

As of 11 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 6 inbound Pith citation observations for arXiv:2501.12735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12735 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:58:02.706784Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:59:14.448812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:37.298243Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef53fa48-6b45-4b67-8dd4-5df4ab33c1c1 · outbound

This paper cites write newline.

Online Preference Alignment for Language Models via Count-based Exploration write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.282596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.282596Z digest=sha256:b48609d703f82b76701c1fc007e487072dfc6a23613c04d5ea554ec3dcf238b1

Observation fa65da28-31df-4ceb-8871-a66f902761bb · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Online Preference Alignment for Language Models via Count-based Exploration Improved algorithms for linear stochastic bandits

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.289381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.289381Z digest=sha256:3b8544d7814223fc50a13b179219d1a4def5bafbc9293215c2c2ebd0254d6770

Observation 37ebaddc-f5eb-4bb8-b750-5cc9186a9199 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Online Preference Alignment for Language Models via Count-based Exploration Reinforcement learning: Theory and algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.295415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.295415Z digest=sha256:2bb9640218ee7299589c8c5cb167b2012f192caa6eca9ed4af9e481de178c030

Observation 335e907c-2771-49c7-9467-6a51634de33f · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Online Preference Alignment for Language Models via Count-based Exploration Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.300277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.300277Z digest=sha256:a609737be88582e5898cda532a0a4df629ee8981eae5ed33e43598753d03dbde

Observation a205591e-4273-4ad1-9286-0445c7a496d4 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Online Preference Alignment for Language Models via Count-based Exploration A general theoretical paradigm to understand learning from human preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.306276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.306276Z digest=sha256:db703dc805b97993af8dbbddeebf7bdbdfec660f167b6a20700072322af7c3db

Observation 3db2ce02-fa1f-4a4d-9889-dd930552cb92 · outbound

This paper cites Dynamic bottleneck for robust self-supervised exploration.

Online Preference Alignment for Language Models via Count-based Exploration Dynamic bottleneck for robust self-supervised exploration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.901418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.311797Z digest=sha256:edbef3d1c93e090b411251eb5512808a50f0e301cc3c613dae54e5d3be4f9f2d

Observation 61b7b2fc-bf1c-45b0-83c0-bb113d54659a · outbound

This paper cites Principled exploration via optimistic bootstrapping and backward induction.

Online Preference Alignment for Language Models via Count-based Exploration Principled exploration via optimistic bootstrapping and backward induction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.885692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.317754Z digest=sha256:7a760a482f8aa07cdef3aae03dd3fb2645228bd084576d2deef00d7151ea51e1

Observation 3f01d01d-c28a-46fb-acec-1ec4fedcad04 · outbound

This paper cites Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning.

Online Preference Alignment for Language Models via Count-based Exploration Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.869749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.323338Z digest=sha256:53d49de0bd348047f60d3c89d0f55558e82168b874ff88547ebe8067bfc8b4ba

Observation 5663f3e6-4208-4386-9523-0f3419e42898 · outbound

This paper cites Pessimistic value iteration for multi-task data sharing in offline reinforcement learning.

Online Preference Alignment for Language Models via Count-based Exploration Pessimistic value iteration for multi-task data sharing in offline reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.853497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.328525Z digest=sha256:39a67d73cf61969587b5ceb2f18d50301ead0304cc2632d1adabb8f617dd239d

Observation 1b9ff02c-0539-4d21-b69d-ec8d6295a0cf · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Online Preference Alignment for Language Models via Count-based Exploration Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.333468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.333468Z digest=sha256:28b0d1bffb5b309bb1b31c486fcbad133b5786318c6762033465bda4815cf538

Observation 22139940-ed4e-454f-88a4-306882a4fed1 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Online Preference Alignment for Language Models via Count-based Exploration Unifying count-based exploration and intrinsic motivation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.338317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.338317Z digest=sha256:9ad17db4e0879c01e1909d02a6784f1a1b1909cacc7ae5cb7ba75b84fcf62f84

Observation 0298d3f8-8a7d-45ee-889c-57af64d913f9 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Online Preference Alignment for Language Models via Count-based Exploration Rank analysis of incomplete block designs: I

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.343378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.343378Z digest=sha256:9881a18aed8d31019e4f9a237f6bb69094434bdaa13b34a6fd4f07dc922bffcd

Observation 91614892-6328-480f-83ad-5b5073a4cfc9 · outbound

This paper cites Exploration by Random Network Distillation.

Online Preference Alignment for Language Models via Count-based Exploration Exploration by Random Network Distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.349211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.349211Z digest=sha256:208ada01b63a81aa50c75c0114a5a19a5c54e48698af393b82adde2ae04b881c

Observation 5daa008d-20d7-4214-a794-a7d172dd9392 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Online Preference Alignment for Language Models via Count-based Exploration Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.354557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.354557Z digest=sha256:2da25d1a48454b3e6ecc9baf36fe3cd962942066cee0fe9327165912f4edc112

Observation ae3edc00-5f1d-441e-8640-e3d5dd19f24c · outbound

This paper cites Deep reinforcement learning from human preferences.

Online Preference Alignment for Language Models via Count-based Exploration Deep reinforcement learning from human preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.359974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.359974Z digest=sha256:0d8664630806c07e6cb3650ede3f6a77322e9eecef34302dcba0678c544a542a

Observation a83b5501-fa90-47cc-8437-fe75775a3151 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Online Preference Alignment for Language Models via Count-based Exploration Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.365180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.365180Z digest=sha256:cefe7c0e3391decbf10ce8c65d8477d106e168a471d41250a696daef409a4bbf

Observation 73039118-e167-4d2e-9a6a-31bfef77f552 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Online Preference Alignment for Language Models via Count-based Exploration Training Verifiers to Solve Math Word Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.370586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.370586Z digest=sha256:6037a7780204ba0e62a69c93638ab004da2de7795acf857ffc6d613198044481

Observation 96c70b83-bb3e-4d7d-880a-138e758e79d4 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Online Preference Alignment for Language Models via Count-based Exploration UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.376355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.376355Z digest=sha256:5c9c48ebd7b2f40eb9122265fc24db11b4d9aa937d39245f830814f8d1768322

Observation c77a1b4f-c896-445a-8633-bd811307e2fa · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Online Preference Alignment for Language Models via Count-based Exploration RLHF Workflow: From Reward Modeling to Online RLHF

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.381265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.381265Z digest=sha256:53dce39d518b1f337b666aea05f973011dd35bf29135be5e4e568134ec8001bc

Observation 0353558b-556c-4724-afc8-e0a726a9b717 · outbound

This paper cites The Llama 3 Herd of Models.

Online Preference Alignment for Language Models via Count-based Exploration The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.386536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.386536Z digest=sha256:c586e100e5c2b3d6357e3883e5ab67e2fa60a3ae61d085f6a7bc631c2362105d

Observation 5a1c866b-ff66-4914-a6a9-2bf484d45a3b · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Online Preference Alignment for Language Models via Count-based Exploration Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.393058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.393058Z digest=sha256:ce2aaa5446820fbfcc138c6394dbc25f644ced36d1a813fbe187e75c9be6be11

Observation cf3c38bb-a4c6-4643-9338-a732163ca67a · outbound

This paper cites Efficient Exploration for LLMs.

Online Preference Alignment for Language Models via Count-based Exploration Efficient Exploration for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.400792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.400792Z digest=sha256:0083a0cfb1e3ba87f55a12e38a0c4c5f8a5c0a9411e6c3b1a4762b672d3f26f1

Observation 90b6bc5f-ed61-4879-a5b2-7aa9de428e58 · outbound

This paper cites Kto: Model alignment as prospect theoretic optimization.

Online Preference Alignment for Language Models via Count-based Exploration Kto: Model alignment as prospect theoretic optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.802452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.406030Z digest=sha256:85dc0bc1804b4e1fac92641e6cc904c8c987637071352a4e5dc1f6cd53ee7cc8

Observation bf25acec-14a5-45e5-96e7-ffe138407b89 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Online Preference Alignment for Language Models via Count-based Exploration KTO: Model Alignment as Prospect Theoretic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.411074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.411074Z digest=sha256:da0fa464913b3dd3a79011a73d311f4f3df55045bcafec7a3387f72793878e5f

Observation 70831cce-6fb2-4675-b1f0-a9132a2859b4 · outbound

This paper cites Scaling laws for reward model overoptimization.

Online Preference Alignment for Language Models via Count-based Exploration Scaling laws for reward model overoptimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.417149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.417149Z digest=sha256:6824f59238e04d8856e0a77b9d7c1e30774325fa61373a48ce94969cc5e93fb6

Observation c4d4090a-b281-4761-b526-efe9d7227328 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Online Preference Alignment for Language Models via Count-based Exploration A framework for few-shot language model evaluation, 07 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.422098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.422098Z digest=sha256:a582bf4383bfae3d1687dc21a2e9fccdb17c539a2d05f273334e89fe7a80591c

Observation bb455c0a-59fe-442c-8611-7e6a6b9419db · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Online Preference Alignment for Language Models via Count-based Exploration Direct Language Model Alignment from Online AI Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.426898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.426898Z digest=sha256:4a29f88dd9d8dc289a927a464da70c184912987c8dd6a3b786a6bad685141e40

Observation f617c9d4-9cf5-475a-98cc-be6e812296e7 · outbound

This paper cites Exploration in deep reinforcement learning: From single-agent to multiagent domain.

Online Preference Alignment for Language Models via Count-based Exploration Exploration in deep reinforcement learning: From single-agent to multiagent domain

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.765809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.431968Z digest=sha256:b4f2211054dc87c2bc72e1871987e261e0da5843901ce2950b90b4d649427e55

Observation 77bf73e0-b47e-4b04-abc8-3d37966d4aab · outbound

This paper cites VIME: variational information maximizing exploration.

Online Preference Alignment for Language Models via Count-based Exploration VIME: variational information maximizing exploration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.749626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.436780Z digest=sha256:31c34e476bbd453478afd12fc9bef8985fd6848a3acfef763256b403028deab4

Observation f3290652-12c8-4769-b09c-31b9aadc1879 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Online Preference Alignment for Language Models via Count-based Exploration Lo RA : Low-rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.441734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.441734Z digest=sha256:f4147db02e0cc100c95841da65ba1e94f87ab3c966d02d75f463b4ba4a1d1e44

Observation b2f05f73-cd26-4f7a-a356-b5d5ecccb49e · outbound

This paper cites Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback.

Online Preference Alignment for Language Models via Count-based Exploration Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.721788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.446933Z digest=sha256:2db93af044ad64582e1863b469037f1cb6b40e5f54a764658ce5a6adb27283ea

Observation c7e4530f-8b0b-4cb2-83e8-46d35f5a7a31 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Online Preference Alignment for Language Models via Count-based Exploration LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.451981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.451981Z digest=sha256:2630d6982ca3097cc3fac3e190d1badb231ffebbaceb1e4b4fe1b9ac8938acb4

Observation afc77d64-92f3-4511-810d-a71c1d8ed02d · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Online Preference Alignment for Language Models via Count-based Exploration Provably efficient reinforcement learning with linear function approximation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.705640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.457552Z digest=sha256:d08df1cd6a9d6a8637088cf54abd54eaac3e88bca20bcd8778798ba66fd8a41a

Observation 1f1c2440-5c48-4ebb-8250-3e39785ed575 · outbound

This paper cites Kearns and Satinder P.

Online Preference Alignment for Language Models via Count-based Exploration Kearns and Satinder P

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.688239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.463450Z digest=sha256:3d7b90b7816fa0908db7d977f6fd4f772acf87b4eac18dabc8f41b0ff7f193f0

Observation de4c7455-0eea-4f23-876b-64f9db55da2b · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Online Preference Alignment for Language Models via Count-based Exploration RewardBench: Evaluating Reward Models for Language Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.469030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.469030Z digest=sha256:9b04eb2d2c353798af834d56a1cc06a01309df6fb38bbb61904280a0be06614e

Observation 8b3b8d5d-bf58-40fc-ae56-2429ed449777 · outbound

This paper cites Bandit algorithms.

Online Preference Alignment for Language Models via Count-based Exploration Bandit algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.475812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.475812Z digest=sha256:5c3030ff5355c43da61aa0fff84acc9d8e466eb1b53a5f858c5afb4b0626486b

Observation dee82772-2af4-4232-8af8-0e9167470206 · outbound

This paper cites A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity.

Online Preference Alignment for Language Models via Count-based Exploration A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.662098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.481760Z digest=sha256:aec067c9abfb949ae09838db07a50ea8b871c056d6fc70808db992e6ec8c1cdf

Observation 0fcf4a93-0e7a-4cd1-9a7b-4f9298893c9b · outbound

This paper cites Aligning Large Language Models by On-Policy Self-Judgment.

Online Preference Alignment for Language Models via Count-based Exploration Aligning Large Language Models by On-Policy Self-Judgment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.487409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.487409Z digest=sha256:43e663859177a80dc3ca060182626eed5b8fb00bc663a06c56665fa884e84bc7

Observation 2708c191-c90d-48ba-9d50-7948c84594fb · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Online Preference Alignment for Language Models via Count-based Exploration TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.492624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.492624Z digest=sha256:b12cda37d20e287f8111c7c843e026838d9f44d1cdf36fec38b3fd7c3ade29a9

Observation 3cdb3c77-9ce4-4617-987a-4ab43bd70dce · outbound

This paper cites Flipping coins to estimate pseudocounts for exploration in reinforcement learning.

Online Preference Alignment for Language Models via Count-based Exploration Flipping coins to estimate pseudocounts for exploration in reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.498083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.498083Z digest=sha256:e87d70d004c22f9ef54fb0fc1a3329bf42a5151211143e9fabd9f27ef41e626e

Observation 45194ecf-3380-4c65-8ebc-83eff819febe · outbound

This paper cites The sample complexity of exploration in the multi-armed bandit problem.

Online Preference Alignment for Language Models via Count-based Exploration The sample complexity of exploration in the multi-armed bandit problem

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.634569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.502922Z digest=sha256:c2e70cd1fa54b8c4a1940fcb8bcdd105b1fed711ce25bd3c08f0837be309c028

Observation 0aa02464-c6fb-408b-9cb4-382731f3dc7e · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Online Preference Alignment for Language Models via Count-based Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.508364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.508364Z digest=sha256:5115646a775df71c3ac92475accdcb2f792f7349d6931fc58dfc680416a1229f

Observation 91a4e0ba-81d1-414d-8fee-113b0b2aebab · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Online Preference Alignment for Language Models via Count-based Exploration Introducing meta llama 3: The most capable openly available llm to date

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.513411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.513411Z digest=sha256:19b7b6d320b83ff455effac57880f5dedf1e3e0756f0664e5bb3dfddb2c30615

Observation 983d111d-4900-4e47-bba8-dc096b0a7fbe · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Online Preference Alignment for Language Models via Count-based Exploration Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.518948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.518948Z digest=sha256:964e5ef18c2bd37a50ba3a545564fcfb2b2e90aa888c338d8d4ea804f7a97c6c

Observation 3dadfa13-9fd9-4bbf-846e-7ebd4a259c5d · outbound

This paper cites Nash Learning from Human Feedback.

Online Preference Alignment for Language Models via Count-based Exploration Nash Learning from Human Feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.524062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.524062Z digest=sha256:1084418c85f76ba157e0938e546baf0560bcce9a93ddfad844b9fb23c946c9df

Observation bde05da0-a5d9-4f72-8cf7-223df294b920 · outbound

This paper cites Count-based exploration with neural density models.

Online Preference Alignment for Language Models via Count-based Exploration Count-based exploration with neural density models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.528746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.528746Z digest=sha256:6726520a71e26cfe5639128e6550cfa18c7f5e9bfd2b85c2cd67abf619f45d35

Observation b0cc0e5f-ebaf-4182-b347-dbc033d14cd2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Online Preference Alignment for Language Models via Count-based Exploration Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.533763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.533763Z digest=sha256:b672fd56b34163e5bc96ae8509cbab65a65498079702f027f9e408f422254b42

Observation b8995c09-eae0-4c6b-813d-d75ac4a71731 · outbound

This paper cites EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models.

Online Preference Alignment for Language Models via Count-based Exploration EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.539280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.539280Z digest=sha256:80dbc7e0c408aab0e26e096084181b9bc8d19a5f23e80ef5e6928fdfd75af5b5

Observation 925ec6a9-62f9-4792-bce0-e83dcb298f1f · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Online Preference Alignment for Language Models via Count-based Exploration Curiosity-driven exploration by self-supervised prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.545044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.545044Z digest=sha256:1e18fb19df1a4cb7d8cd58af25f16c775e362b0e801a108c6f69ba85d67c089c

Observation 4860ad6f-df33-4618-ad38-765b80d8bf23 · outbound

This paper cites Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning.

Online Preference Alignment for Language Models via Count-based Exploration Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.572349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.550680Z digest=sha256:7fcf6f1340d8adafa4d5264080bb5a18e393b3668dd1b5d2ab78c5dc606a6caa

Observation 3f86d261-1d25-49dd-b357-a0df6ebbcbdf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Online Preference Alignment for Language Models via Count-based Exploration Direct preference optimization: Your language model is secretly a reward model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.556057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.556057Z digest=sha256:d1232dc1f5e38034d51e5df5b8da592ddc85a80fdb6ee4d6a3491b88e860861e

Observation dc65dccf-c32c-4e79-840d-5f81612232a5 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Online Preference Alignment for Language Models via Count-based Exploration Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.560951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.560951Z digest=sha256:6193368733789cea14dc838bab5b4c671bc3750d7b0eab836af04a78e4c77475

Observation fd5effcb-c5e7-486e-9db5-38f54e5ee978 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Online Preference Alignment for Language Models via Count-based Exploration From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.566480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.566480Z digest=sha256:49ac766c2e55d9253516dde1653e2547e1131eb2352432afb2ccaa2c761acc66

Observation 0a117eff-091f-4b7d-a6bc-4bb6fe7826a2 · outbound

This paper cites Optimistic exploration even with a pessimistic initialisation.

Online Preference Alignment for Language Models via Count-based Exploration Optimistic exploration even with a pessimistic initialisation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.543956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.572108Z digest=sha256:d12fa858e5256fb0d5d53b467e558840a3309165f06de7a7462b4f864cbb9ad6

Observation f3b8509e-4cee-4633-a644-45c1cc7f43ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

Online Preference Alignment for Language Models via Count-based Exploration Proximal Policy Optimization Algorithms

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.578564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.578564Z digest=sha256:80eb4e21522a78d7b6e09bbe609d64b942154f639bf400fda9b2daab99076911

Observation 4974abf7-e451-4b7c-b3a2-52f85b3ce142 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Online Preference Alignment for Language Models via Count-based Exploration A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.583705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.583705Z digest=sha256:25b3a3bebb00e1e907cbe1a9e72227c72d7ba0b3525cd1cb3d4a66b764dd0b0a

Observation 83b75950-a1ae-4a1c-81e5-8a156403d141 · outbound

This paper cites D2PO: Discriminator-Guided DPO with Response Evaluation Models.

Online Preference Alignment for Language Models via Count-based Exploration D2PO: Discriminator-Guided DPO with Response Evaluation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.589047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.589047Z digest=sha256:08fb3053206a2d96c4ee455a04bc970bd3e9bc52925fe7783cfc22be4b22d30e

Observation 3eab34b1-641e-424b-a777-9405db12d272 · outbound

This paper cites An analysis of model-based interval estimation for markov decision processes.

Online Preference Alignment for Language Models via Count-based Exploration An analysis of model-based interval estimation for markov decision processes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.527064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.594628Z digest=sha256:9bc62ef9a2bbacd94ee8039d10780cfbd546c818977d372943a05f9fcab26fbf

Observation fd4f4167-2e88-4f61-850d-ab1fbbdf79d0 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Online Preference Alignment for Language Models via Count-based Exploration A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.599465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.599465Z digest=sha256:d379ee23031977f77022d4cf3dd1710a0fa04e1ef421ac15d546db7ab2b75347

Observation f8fd4d4e-54de-4202-9681-2119cceb56c2 · outbound

This paper cites \# exploration: A study of count-based exploration for deep reinforcement learning.

Online Preference Alignment for Language Models via Count-based Exploration \# exploration: A study of count-based exploration for deep reinforcement learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.604270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.604270Z digest=sha256:5f6b3915e0f4ee29f433fded3f01406f1c91c0ed760803234545a978766bee16

Observation d5d1e5d9-8723-4de7-97a1-bf97586244a1 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Online Preference Alignment for Language Models via Count-based Exploration Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.609795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.609795Z digest=sha256:e4e52aaef5bf8a44cb91b3bc032ca4e5cfc5fe8af3281f91d790b692e7525213

Observation 28198650-2280-43b9-b673-bd3300888985 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Online Preference Alignment for Language Models via Count-based Exploration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.614756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.614756Z digest=sha256:af8996a76edfa40a7e5da3487f9d813f6caf3823788cc41daace6e2fe8283bf9

Observation 11f1b9f5-4ff3-400a-9676-bd2544fa9720 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Online Preference Alignment for Language Models via Count-based Exploration Zephyr: Direct Distillation of LM Alignment

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.620383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.620383Z digest=sha256:773411878efd911943712def14d839304e558c1218babec3608913fa35ffd74b

Observation 90197c55-9f72-42fe-a34a-f099727d8e81 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Online Preference Alignment for Language Models via Count-based Exploration Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.626291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.626291Z digest=sha256:e865d7241abb1447a5d23f9f2f8ab2330f2a3e1eb5df481f519d7f3d401fa61e

Observation ff9d91df-bbb1-4686-bece-ff3fa3c263c4 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Online Preference Alignment for Language Models via Count-based Exploration Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.631290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.631290Z digest=sha256:1d48061be65c060c934aee877b9ccc5eed0df65ac28b11133ffc48539a1f36a1

Observation b68010ec-8d77-469d-8b74-cc3561b50960 · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Online Preference Alignment for Language Models via Count-based Exploration Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.488698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.636237Z digest=sha256:5dc12f3acf0a4b7142b4a3bf9feddc24df814d844e8dd8d5246761d02f3e484b

Observation a5c652fc-42fe-4a2b-8e5a-f41cf3cf9f65 · outbound

This paper cites Rorl: Robust offline reinforcement learning via conservative smoothing.

Online Preference Alignment for Language Models via Count-based Exploration Rorl: Robust offline reinforcement learning via conservative smoothing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.640741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.640741Z digest=sha256:4c9cbd49e67d3034fc177ac1c2fb1f4f4baf8d6b014a299a21e7e874c20a6b35

Observation 64c656ab-1840-42e8-b787-1f0d1cb7924d · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Online Preference Alignment for Language Models via Count-based Exploration Yi: Open Foundation Models by 01.AI

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.645376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.645376Z digest=sha256:7712b532fae6feb4a6c15e7c1dba2692fa08dd46d898f5d6533d588ef424facb

Observation da437e9d-fa08-415d-8c39-920c81aaccf4 · outbound

This paper cites Regularized Conditional Diffusion Model for Multi-Task Preference Alignment.

Online Preference Alignment for Language Models via Count-based Exploration Regularized Conditional Diffusion Model for Multi-Task Preference Alignment

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:58:02.864613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.649973Z digest=sha256:d76a5dbd744fd76f6bf0312e2fc0bc9ed41fa5a0c71491605cb13f09fb6fa55f

Observation 16e79d25-7dbb-489f-ba33-e2376f4b90c8 · outbound

This paper cites Self-Rewarding Language Models.

Online Preference Alignment for Language Models via Count-based Exploration Self-Rewarding Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.655481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.655481Z digest=sha256:bc3921734fe9aa8aa609692155eb094ebc4d871a80a5314b27895e8b764c224b

Observation 158014e7-edc5-41f4-84c0-7e149477fe99 · outbound

This paper cites Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control.

Online Preference Alignment for Language Models via Count-based Exploration Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.660431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.660431Z digest=sha256:84d23c5d10b0b0aad0c641ee57b38c71fe24b232ffe108f3e9d34ceb073afa78

Observation e93129e2-bb24-4e68-b060-75ec78a33bfd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Online Preference Alignment for Language Models via Count-based Exploration HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.666433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.666433Z digest=sha256:5fb38a37ffe2326f1d3d1da663b8a933015b4eb35976c4189b9927c2e3161205

Observation 8589d06f-41f9-4b44-9136-39958e26ad64 · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Online Preference Alignment for Language Models via Count-based Exploration Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.671392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.671392Z digest=sha256:1ff869f7e9b1b62a5bbbe1184e3e5b9a130b1292a141c95f86dc27390b332d7f

Observation ccd42246-84e3-40b2-ba18-505b2e9a2926 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Online Preference Alignment for Language Models via Count-based Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.677210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.677210Z digest=sha256:cb489ca1a885a24f01505d977c08f23e1562bcf8169c0feb7c9cd44eef9b70bd

Observation 219b48b0-ef9c-4a92-966c-d08c93b6e255 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Online Preference Alignment for Language Models via Count-based Exploration Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.682027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.682027Z digest=sha256:7f40ea49025931d2cf1582777995f182f9076c53246f89af158795b1384f5aa2

Observation f1210001-b081-4da7-b0eb-90c876f76849 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Online Preference Alignment for Language Models via Count-based Exploration Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.452536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.686792Z digest=sha256:8a365953349856b95a49d51a964a32e39d0bacc0a9f72efbcea1c91f9e7e5938

Observation 7c8604ef-bd9c-4599-9b57-b61fe2bbadc3 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Online Preference Alignment for Language Models via Count-based Exploration Fine-Tuning Language Models from Human Preferences

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.691569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.691569Z digest=sha256:4e841fc9254db000aa0babee3d42eaa012f9cbe217e65a596885812f651758cc

Observation 335f4395-a3a5-471e-99aa-b46c6a07584b · outbound

This paper cites @esa (Ref.

Online Preference Alignment for Language Models via Count-based Exploration @esa (Ref

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.696219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.696219Z digest=sha256:91bbc3d64f8d3371271212359e7f17dfda15452ab4c16a0ea10fda57e4d1854f

Observation 3f144e43-f8d6-4621-b292-9fe9213e1530 · outbound

This paper cites an unresolved cited work.

Online Preference Alignment for Language Models via Count-based Exploration Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.701890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.701890Z digest=sha256:9199a2d2e07e9e1dd264cc632ed756f94c26b1afcd7b7016fef36e87901903bb

Observation 34a2b941-c2a4-40be-922c-02facace7c73 · outbound

This paper cites a&!Ï "l=-BpEMUs J5ū ?bvtCy O (^rڜH - Y`J* aH/'V t@Ys ;ӓj(u B FBa 竑 6 ^mN OB Y>X 5 >D Q=h .+' A Ի, _|k P(qd/T) nV P C/ۿA+ڽ.W b< ydQx> gQ`U Ԡ < a.

Online Preference Alignment for Language Models via Count-based Exploration a&!Ï "l=-BpEMUs J5ū ?bvtCy O (^rڜH - Y`J* aH/'V t@Ys ;ӓj(u B FBa 竑 6 ^mN OB Y>X 5 >D Q=h .+' A Ի, _|k P(qd/T) nV P C/ۿA+ڽ.W b< ydQx> gQ`U Ԡ < a

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:58:03.416315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:58:02.706784Z digest=sha256:932603cc8142925843755070c040182575f668e762f8a4f6ff93f9bed3e206a2

Pith citing papers

Observation a61d8773-626d-46fe-b3bf-f7fa28628346 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Online Preference Alignment for Language Models via Count-based Exploration

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.448812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.448812Z digest=sha256:d8ee9b5a2cdd79b34c5a1cc82353c0f523024072bc3d3e504ee4d67d2755c6b5

Observation f0ebcc62-be89-4ba5-b101-e6d25f04a06e · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Online Preference Alignment for Language Models via Count-based Exploration

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.606163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.606163Z digest=sha256:b08abf68cfdcc3d2de5e4c73a44834c4286a27af6287ff39c73aaed4f23861a1

Observation b02aeed5-5e6d-4ec3-a9fd-2532a5403d52 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Online Preference Alignment for Language Models via Count-based Exploration

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.222940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.222940Z digest=sha256:00fd0e1c604b68572bc703243df3b31686110174886d305f7d984e1442c07e47

Observation 82d9f26c-c454-4aa4-90cd-cef9da877dcd · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Online Preference Alignment for Language Models via Count-based Exploration

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:47.450020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:5b14e7e78deb9cd1f6184b27e6507b41e0f6771bb04596571270be05c81527bd

Observation 7e4c3140-62f9-4b08-b63f-b3e808a90c04 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Online Preference Alignment for Language Models via Count-based Exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.459065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:93b96c3bf00934e0c0772d5822f0ed0a6e7f9951206fa338cee4e31c675d3d3c

Observation 6ae9a656-4640-495b-853d-d45c2a6cb363 · inbound

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization cites this paper.

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Online Preference Alignment for Language Models via Count-based Exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:37.299709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T14:02:05.833651Z digest=sha256:379a3ea0ab8c75fd4bbbd0635afc60099ba19620ba87dce8cdc6491bbd330105