Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:37.117468Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2411.10914.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:37.117468Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:37:39.875926Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T06:12:07.140554Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0ec019e-8f45-4012-b076-58691db7ddeb · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Survey on Data Selection for Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3154b58-ab12-4254-8c5d-afc49900c419 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 470b6e01-bc63-4cb0-ac21-1cd26bac9118 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93c2bbe0-6ec2-4180-ab29-6b3781f3c3df · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45604d82-2d11-4c4a-84b6-6ef30e438afb · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2519ffd-99eb-4bb3-a109-e7e8707deb90 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5362d29f-2207-4fc8-b396-47188baba437 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b4ffa9-9d0d-4952-a528-318a1175749e · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38aa032c-7cac-4485-9382-77a7bd88c691 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62cacff-d5d7-435d-ad73-8a2573ce2ebd · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment KTO: Model Alignment as Prospect Theoretic Optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0024ca61-e16a-438e-874c-caa354ecc0a7 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Human-like Summarization Evaluation with ChatGPT
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e828a1d3-e343-4d26-bb7a-f2873e5832e2 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0713a90d-c316-4a3b-99b7-5bc41834f415 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LoRA: Low-Rank Adaptation of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949183c7-e3f5-4fc1-8f41-806b3bf59bd8 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef45d9f0-f162-4f56-90f9-db7d5747802a · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Calibrated Language Models Must Hallucinate
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7c0633-7a0a-4f16-b7f0-ee028ace3e3b · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87bc80b6-af2c-4e9c-bc38-ea8157032dd7 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Adam: A Method for Stochastic Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f963a1-910c-4938-9a68-620575c4cdb5 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7457da-6aca-4b07-8be8-af1734fffd41 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e6bdf45-8520-4bbf-8dcb-e76b1121a180 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b530734b-bf51-43c2-bb55-b2345dbb6d04 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dce8b1e-2382-4c61-a350-b2c549601292 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e7b513-cee4-435a-b139-692ff894730e · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74eb6b19-7226-4387-b346-8100999209c1 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Filtered Direct Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2439acd-c0e0-4144-bd67-41d12f5b2700 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4c1eca-3c78-4ff6-9d3d-acef119a17fe · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Text Clustering with Large Language Model Embeddings
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6616da-87e0-4958-b31f-57c03c22d5b2 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf408ba-8b58-4cb2-9a53-57592c1747f5 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c3a59b-ed83-4baf-9d37-c742de0a9c58 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b066ec-1053-4320-a8ee-bed034af49e5 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96594f66-9ded-4b99-969d-71a73a8165ed · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086b6d6e-edb9-4462-88db-ce08ad0f4fec · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42734856-6fdf-4619-a1cd-ca674320fcc5 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Large Language Models for Data Annotation and Synthesis: A Survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9151b43a-189e-41c4-95cc-b9cf93321c97 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 600b09c2-b0cf-4a80-84ee-201ea6d4c4f8 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d49670-edac-47e8-ae30-29b0204c216b · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7052df9d-d093-454c-86e8-712b6a15f902 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LESS: Selecting Influential Data for Targeted Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caff5bb9-a58b-46fa-8c55-8b9cfcb40e0c · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31669112-6165-4d6a-9289-87bc1809bbc9 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309eb1cf-e393-4d62-bbec-c2ccdf9dabdf · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9956b1d7-86f8-4c61-8146-b2d4e0f16354 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de50be09-0fa0-4baa-b30c-f4b0c619a46a · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afef957d-bbef-4f72-8fc9-050b3d87cb66 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aab5a24-06fb-42b2-8d4b-28dbd8fdceeb · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf83a48-672c-44aa-802b-c3476bd48e78 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1cca3cf-1002-453b-8444-bfbf03024b56 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 58f29db0-07ed-4283-88e8-8a00c4d0fe53 · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70199fa9-37cd-4829-99cd-5b4eeb78834d · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment online" 'onlinestring :=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f1c020-bf53-45b2-8905-2631d69a296f · outbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment write newline
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa35949-57d6-484d-b57f-992a1b64dab6 · inbound
From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32eca10d-4bae-43eb-851b-42e3a8ed6445 · inbound
From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.