Pith. sign in

Paper Citation Record · LEDGER

A General Theoretical Paradigm to Understand Learning from Human Preferences

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2310.12036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12036 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:59.145850Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b045bfd9-05d4-4f25-8c19-72a780e85c46 · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.769240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:4d3fab745fb84111c0e3ca610a8437bb1deccf2f57a9a1e81e0c38fb75f411e9

Observation 972de14a-91e6-49d4-94cb-6ccdaf5c574b · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:30.944982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:c789c53d9adfc47fc7dffc66d772171302c5d9edfa552d40b3f8fc6a9d7d2bd6

Observation b3153567-2340-453f-ad62-cc1375bc3e29 · inbound

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions cites this paper.

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:59.145850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:59.145850Z digest=sha256:e9e232b2a66c1f2a85d2be056855525bce74a301432ad2c304fbbdc5751feb6d

Observation 54a2c4d8-ace3-40d0-979d-4e89e722de36 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:45.997262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:45.997262Z digest=sha256:23d7261ab9019896ce6b4f36173d549c5beca4d74ad2dd1bc51ca9e210fbabf9

Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:47.923648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:47.923648Z digest=sha256:db4617b484de218cc4319a9372284a62391f5f7fd55ff2180ca38bc81a52fb3f

Observation c9abefb6-8aaf-4792-96b9-8cd2dca89181 · inbound

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising cites this paper.

Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:42.228327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:57:42.228327Z digest=sha256:10be6401dc6a47488423a98a740e230da2f349c9dc10cce08789a4a6c8dc66b2

Observation 38a7ec8d-1f7a-4786-babf-f21180eb9c63 · inbound

Reinforce LLM Reasoning through Multi-Agent Reflection cites this paper.

Reinforce LLM Reasoning through Multi-Agent Reflection A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:07.333606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:07.333606Z digest=sha256:a80a9cab857231005c14d16efa023c276a8ba72bbbba964c5da574cf2613fb50

Observation b2c59bc0-5896-4dd9-aef2-605091779e68 · inbound

DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization cites this paper.

DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:16:06.358214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:16:06.358214Z digest=sha256:7a9fac85d59db59f82b18625d8f33600fa9401ecd7b6e38d2fb45b6a48d76cad

Observation fe1636fe-aa28-4d0a-a15d-841dcc051593 · inbound

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation cites this paper.

CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.283348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.283348Z digest=sha256:e1b545217e275b2d1d589671504ba9842fa58def5d4056e3dff5d73b74d800f2

Observation e1575096-4be4-46e7-b026-bd9c440d36d0 · inbound

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap cites this paper.

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.607776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-05-21T23:46:24.208438Z digest=sha256:e557676b94fc520279f635bef917e41466514539be994c38b9b9265b84398c14

Observation fec1937c-8fd6-4e10-b526-f01cd2d18c30 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:13.809159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:13.809159Z digest=sha256:e54580a2489e4381bd6bee8ce868e1e99a69e82aad0f61e22927b8db483c96b9

Observation 4e7e3696-975a-4dd5-9d5b-d36ed9ffd8d3 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.813225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:417c049e1587942af6261344a01b7fdd8502db1a33e1964212624d2572a1111b

Observation 4deb9cf5-bfb4-4a4e-a885-a13706c67041 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:18.668425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:18.668425Z digest=sha256:3ea0e362f1f5b0568bc337b4498507cc8fe0eb87b0463c7605023b08ed48f671

Observation d058d0b3-5f42-4060-9235-c76e24e0b6db · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:19.502646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:19.502646Z digest=sha256:9ab2ccf93b7bc8e9d1cb2c09904f3236373a634d10e87b5ee4d75c0d49b70c35

Observation 03710aaf-d50c-4e56-b8a4-670bcc2e4707 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.683804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:daeab29828746cf88d267931e7bef6995e25091e325f7c87172e29a02d65d4f0

Observation ae04bdcb-ef5f-4f03-9cd2-af5fee1a1754 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.004199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:b7a8e8eee8169be2bd6027c11cff37afd73d7c0e1be612eefcb178079e0f696c

Observation 2deb40d5-6e9a-486a-9ba4-13dc4827f5c1 · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:28.007161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:58fa7ca4459534bbdd87ed871d37f1fc7f71b0a7326f64b28829bfa4ddd41cc0

Observation 36ad412e-a310-41da-b084-7a7d3f8d7d7d · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.313009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:bc3b96cb130e9b494227c7fd10037e18ea58926a1675bc3a8538cb892a3b514b

Observation 89c015fc-013f-4f68-85d5-8f127abb6416 · inbound

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning cites this paper.

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:32:19.327417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-13T05:29:42.381280Z digest=sha256:5c5d40f8928010dd2b461f66c267ea3136846ed25fd16cef547ee6973b6f600f

Observation 4b7e7949-a768-47e2-85a4-642bbdd8a483 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.181807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:4c692eddba6fb94ea1b6a9ecd93c650e7020bac20108b8a3354d08b7487a86aa

Observation 09a58498-7564-4110-a7c0-603eedba08a4 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.641433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:a17732f63dde01345666d7319ffb6fc78222679d06e1fd13f2087eec02db7e8d

Observation ce4112a6-da47-454b-a8cf-6dbba167bf8f · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.214389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-05-22T08:07:51.697353Z digest=sha256:63d7c91894857f0900ad0d93f623695c790d7d5b3d5f9376e219854e883e354e

Observation 735f9d5d-4baf-4dcc-be85-c42485d78836 · inbound

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models cites this paper.

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.998440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-30T17:48:02.228941Z digest=sha256:0d8bb5ed34fa5e07e437ff1531a31964654324f50fc9c1c23977a2567b6d4b9d

Observation 5b41e8e6-8bf3-4891-842a-8fd6f1a8a774 · inbound

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates cites this paper.

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:33:24.402718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-29T12:29:55.729913Z digest=sha256:721cea93be9fcb1e4f713ef5ec07de83577953cb3118ab02a8d5206a2e70e9de

Observation bb7add3f-651c-4f56-bc78-9cd92d032cec · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.512570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:39347b6f333e0d609c939670ad696471233b0461c639a8117df219b3a59d910b

Observation 75c04942-f7a4-49e3-81b1-6bda7cfbccb3 · inbound

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization cites this paper.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.361855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:772332e969ec2ee496029ad340b77ed879f7b89e224ba69b199be611984b6e35

Observation 5b049f08-6e51-467c-9c92-342a0f7eb1a2 · inbound

Weight-Space Geometry of Offline Reasoning Training cites this paper.

Weight-Space Geometry of Offline Reasoning Training A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.385341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=arxiv_source observed=2026-06-26T10:26:28.702213Z digest=sha256:a73a393d94c39dfe3a4c5ef10393c291d5c6ef0e3a5a7446c2c58748e55513aa

Observation a11da4f5-e823-4ed4-a425-56b5519e8320 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:47.161717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:535933acb0fd32e91a0eef6188f8e669fba86836c8d6997f6442ece22fb4838e

Observation 38364724-f099-488a-a77c-71f5162f8b76 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.556729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:6d8d86acd404659ca83a1865816664ea005fcda3c8181810b42d1ce6d5171245

Observation a1437ddd-5690-4ed6-a8bc-fa94eab0e2ef · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.374927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:2836deae9e2ff416d30e39017237363ca5b18e4c8f4d25ed6ff7f8e93b1994f6

Observation b6fc0ef8-98a3-4acf-9eb1-8386c9e100f9 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:16.169555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:16.169555Z digest=sha256:cb17e536f809f1394bdc1f757b68d2e48b01165876168dc9f2e354ca706508bc

Observation a5e8bc7a-5eba-4e50-a3b7-20421d5bdd65 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:8730a0a67fec72a8b47b0125e88e1333773538dec138a8b54e79e91b6820a794

Observation a8fe755a-582d-4e88-b4ca-db7b229fe51b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:f61b1a8c69eaa6d45688e0feaf31495a7428fda1d4f9b7d949c3ce2b343bf533

Observation b839b03e-a437-4a0d-a884-49efdc099734 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:51.116483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:51.116483Z digest=sha256:4977c880680c81180dc91ea15ade6db180dad7c166d59a2452f1df30ef392cda

Observation 8899f132-d89f-4e05-bac8-94441a09aef5 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.021635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.021635Z digest=sha256:26df775cfe9104ff5c51668d25137c0feb802a8d71a3e71e5f17f2b1aef08ae7

Observation a51d4882-beae-4871-b73f-73d1a086b0f0 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:50.414984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:50.414984Z digest=sha256:11be1dbf7d442dac67fb4862912cf32ccbd07c769a17c8f86de4b7e569b625cb