Pith. sign in

Paper Citation Record · LEDGER

A General Theoretical Paradigm to Understand Learning from Human Preferences

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2310.12036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12036 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:05:15.089359Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b045bfd9-05d4-4f25-8c19-72a780e85c46 · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.769240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:fe027696f955934662e0208e6b8e3909beb736139f4a2ddf090d0ee29db93213

Observation 42e8f3a9-a770-4e32-93e9-97381c0226f8 · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.089359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.089359Z digest=sha256:97346eb2c8e625f944b252ddec79e8d8861bdc8584dcd69a05b4ba430719149d

Observation 267afe4a-96dc-445d-801f-0d2a3660b55c · inbound

Free Process Rewards without Process Labels cites this paper.

Free Process Rewards without Process Labels A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:06:17.842642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:06:17.842642Z digest=sha256:0917e78cd42de1dd1ff6b15cdd0b85fde48ca415936a988c64f110b41e9b7aeb

Observation 15dea84b-555b-4a49-98f8-50a8e061c7ca · inbound

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model cites this paper.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.822647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.822647Z digest=sha256:d8e7afd5f4ecb8900f6f2dd6599626139c2cfebeed45d29bb0ee0b8576a4f981

Observation a6df2947-5ffe-4330-ab02-943c86d44356 · inbound

Understanding the Logic of Direct Preference Alignment through Logic cites this paper.

Understanding the Logic of Direct Preference Alignment through Logic A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:20.892964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:24:20.892964Z digest=sha256:0a9852c7458104cc42104a43bc8a769e664c44b54887eaa2bb3a0355c044b47b

Observation d8f7dbff-b358-418f-935e-0c88108c1415 · inbound

InfAlign: Inference-aware language model alignment cites this paper.

InfAlign: Inference-aware language model alignment A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:57:54.202526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:57:54.202526Z digest=sha256:c88e30ba2a98ddfe1c1d92422258921f6d5f8f8692d55707809f893c16993a25

Observation 271cf759-74ff-4e8c-bbca-483ebcf9d94b · inbound

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking cites this paper.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.356697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.356697Z digest=sha256:af2f8c783a070908c654a7b7f8ee34a655a5afcd478e199c3883cec504a8cf9d

Observation 5aaa6f2a-ec60-4d7e-bc62-efde2d93a408 · inbound

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking cites this paper.

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.215406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.215406Z digest=sha256:e1f94528c4c89e76d8a315ae9be45bb67034b9ad87b7dfb96fe39861fd4cb495

Observation 7699aa7b-11f8-4f6e-8770-38fe3745ab72 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.577727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.577727Z digest=sha256:abb7f764a9a92f367c2468139fc92df7fd995a6cdf2b317dea11c4dc74af20a6

Observation 972de14a-91e6-49d4-94cb-6ccdaf5c574b · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:30.944982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:da4e16b70c4e71565be0cb564de598725c1cc2c628e8d516e517ef8c6864c2f5

Observation cf191c35-876d-4119-a643-6416b43f25a9 · inbound

On Fairness of Unified Multimodal Large Language Model for Image Generation cites this paper.

On Fairness of Unified Multimodal Large Language Model for Image Generation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T04:52:40.534383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:52:40.534383Z digest=sha256:1b84ddad34c63c5133a330cbe9680da342ae01e8bf5c771d0e062baa602318da

Observation b3153567-2340-453f-ad62-cc1375bc3e29 · inbound

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions cites this paper.

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:59.145850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:59.145850Z digest=sha256:8683823ebaba543ad05d1131d95e1475a9b7ba698d35240a204f35a07dc3dcb2

Observation 54a2c4d8-ace3-40d0-979d-4e89e722de36 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:45.997262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:45.997262Z digest=sha256:b5b516f7d6b770582ff427615c934a7109765fa72263a030c13491cdb5ebf752

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:22c449a924894b6af79ece20d83a19966bc79d00816bbd0f380a4d22b1b8c436

Observation c9abefb6-8aaf-4792-96b9-8cd2dca89181 · inbound

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising cites this paper.

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:42.228327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:57:42.228327Z digest=sha256:a294e06d4a69ac1682b4c1479c5adeeeb76fb0e32e56bb67dd0a14df06098ed1

Observation 38a7ec8d-1f7a-4786-babf-f21180eb9c63 · inbound

Reinforce LLM Reasoning through Multi-Agent Reflection cites this paper.

Reinforce LLM Reasoning through Multi-Agent Reflection A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:07.333606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:07.333606Z digest=sha256:82c81885d3bde61e26dfaf97cce31f9d2f94c54831387b631dcf2addc448fec8

Observation b2c59bc0-5896-4dd9-aef2-605091779e68 · inbound

DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization cites this paper.

DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:16:06.358214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:16:06.358214Z digest=sha256:763a170efe8301d0c5488658a1014d615ea92621c0012f986e6ce898e80e2972

Observation fe1636fe-aa28-4d0a-a15d-841dcc051593 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.283348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.283348Z digest=sha256:2acd9a13c6c8497a7401ddfd4affbf2e08294d8086cdbc11b168dc1e959ef250

Observation e1575096-4be4-46e7-b026-bd9c440d36d0 · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.607776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:f3dee8bb4f1fbe2ec4a780a482900d1c88db287361880ea3634b406ae83e22ba

Observation fec1937c-8fd6-4e10-b526-f01cd2d18c30 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:13.809159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:13.809159Z digest=sha256:dc3cb6d9f3ce5e404f5999093c122c23767b71343a8266e91e91c59935c48a7f

Observation 4e7e3696-975a-4dd5-9d5b-d36ed9ffd8d3 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.813225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:eab8531a329fff992d18ae6a136ad1d6ee6179c5b597f2a99d30d01c01737ce3

Observation 4deb9cf5-bfb4-4a4e-a885-a13706c67041 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:18.668425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:18.668425Z digest=sha256:847a9a248feedb0d8f6aa12535cd9e559e997d78a2bfe4fda3913ed66c099e93

Observation d058d0b3-5f42-4060-9235-c76e24e0b6db · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.502646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.502646Z digest=sha256:d9208ce66b3c4f5187e28a52f99715e78cb7387cd8f4c9d08a18fda26ece72d1

Observation 03710aaf-d50c-4e56-b8a4-670bcc2e4707 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.683804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:fc84e7f9701cfa42894a327eddfeec5e214b6ff1a5f8b7dc6cf433b5dcef4fbd

Observation ae04bdcb-ef5f-4f03-9cd2-af5fee1a1754 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.004199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:1a191d1669435f5a750b88e06501ae5d3fab0903c530a7596b0e012fb21678d9

Observation 2deb40d5-6e9a-486a-9ba4-13dc4827f5c1 · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:28.007161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:c9beea43c996dc8b2306f608f41a94650c20ce2e8763b1213ed7c5dac8e1f274

Observation 36ad412e-a310-41da-b084-7a7d3f8d7d7d · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.313009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:139c00d3f88bbb0f38deed087c5c53c6c7470ed5ec73ab9e62221e331286c16a

Observation 89c015fc-013f-4f68-85d5-8f127abb6416 · inbound

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning cites this paper.

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:32:19.327417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T05:29:42.381280Z digest=sha256:480acbdb54ec75c16bfdc83125aece2f70c4f11cb618c98cb5bacd616070e149

Observation 4b7e7949-a768-47e2-85a4-642bbdd8a483 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.181807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:cbf2c89b3c17396330b3caedff3f7e55af524b021bbbc0da16ff174a15dc9498

Observation 09a58498-7564-4110-a7c0-603eedba08a4 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.641433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:4d510dccfe2f40fa527a119757ba8db636d606717d771cd8b60cfc1f733913f9

Observation ce4112a6-da47-454b-a8cf-6dbba167bf8f · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.214389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T08:07:51.697353Z digest=sha256:adf8559fa1ddc838710a55d0ad293a65edfb8af42f5386537b5b394f7a1bbc13

Observation 735f9d5d-4baf-4dcc-be85-c42485d78836 · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.998440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T17:48:02.228941Z digest=sha256:edc924aac1ac3517f23558ca05940e269e79a9eceb01f129f1dcf8198bd6fa91

Observation 5b41e8e6-8bf3-4891-842a-8fd6f1a8a774 · inbound

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates cites this paper.

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:33:24.402718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T12:29:55.729913Z digest=sha256:de731e7f21e2245cbb2069c2dd343a473adf5bac91c9c2203a318d5d11f06367

Observation bb7add3f-651c-4f56-bc78-9cd92d032cec · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.512570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:2a5118f4b0f52cb01ada9baa43955f3e609a5a2362d417173cd08611ddfa937b

Observation 75c04942-f7a4-49e3-81b1-6bda7cfbccb3 · inbound

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization cites this paper.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.361855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:d49fb47734968201bf78dee62f11475e835055224bd744e9914eabdda9ea7f14

Observation 5b049f08-6e51-467c-9c92-342a0f7eb1a2 · inbound

Weight-Space Geometry of Offline Reasoning Training cites this paper.

Weight-Space Geometry of Offline Reasoning Training A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.385341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T10:26:28.702213Z digest=sha256:33c20c2c48b0e75acc46110803c7682c28edd6229207b46c08e6bf42c8e921c4

Observation a11da4f5-e823-4ed4-a425-56b5519e8320 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:47.161717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:86f2d92eec79fa81dc902095d47efebeee8946cdc95554e8310be7b45052d68e

Observation 38364724-f099-488a-a77c-71f5162f8b76 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.556729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:6173539525e6e53968ac2cca1b4abb502b9faa1d4ac0cb2768c671db5ed1654e

Observation a1437ddd-5690-4ed6-a8bc-fa94eab0e2ef · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.374927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:62ffec1882a678c13ce0d0b62c9a49649ab07e94f7546da4c6307095b261159e

Observation b6fc0ef8-98a3-4acf-9eb1-8386c9e100f9 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:16.169555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:16.169555Z digest=sha256:4056b6f7b658e806e1f0f50d2069bd29fa7dceb7f025f08143e62aa51a17a934

Observation a5e8bc7a-5eba-4e50-a3b7-20421d5bdd65 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:ead2cac0ce5748ffe32c821a314ac29815633b39088914ec9d199e9a0544d663

Observation a8fe755a-582d-4e88-b4ca-db7b229fe51b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:14ece03a2fe140883bda71d9c46f24768330ccf043db07ae865781043e430add

Observation b839b03e-a437-4a0d-a884-49efdc099734 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:51.116483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:51.116483Z digest=sha256:989be6919b929526fdc0b5d92ea36ccc2d7b29ff6d102ca88f5799f02c897a68

Observation 8899f132-d89f-4e05-bac8-94441a09aef5 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.021635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.021635Z digest=sha256:c7f3e235bc3f9a811791a1dff1b09a9ee19e96d93c3fdc37e8e9f6ff31c26f9b

Observation a51d4882-beae-4871-b73f-73d1a086b0f0 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:50.414984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:50.414984Z digest=sha256:e586d0f501ad4c342a47f81f05a6e82239cec8d632a931c02932bcdfb82e0928