Pith. sign in

Paper Citation Record · LEDGER

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2411.10914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10914 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:37.117468Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:37:39.875926Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T06:12:07.140554Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0ec019e-8f45-4012-b076-58691db7ddeb · outbound

This paper cites A Survey on Data Selection for Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.907006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.907006Z digest=sha256:27b3113ce9b9c497d567fa85354b0e26fcba19aa50bd2b950bff43edfcf4906c

Observation c3154b58-ab12-4254-8c5d-afc49900c419 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.911802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.911802Z digest=sha256:ef7d93611588c6f5f7ff877bf2c621b06a6096d9988337efe2905e82167e2cf5

Observation 470b6e01-bc63-4cb0-ac21-1cd26bac9118 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.643972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.916115Z digest=sha256:ade95b8ec387076a9ec41836fb9102e461c082de7140ba0e054467145f83297c

Observation 93c2bbe0-6ec2-4180-ab29-6b3781f3c3df · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.919782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.919782Z digest=sha256:10713861ce8fd91599caa265d8e8ed4f73996e4043e61270b4c7779f2d72cc74

Observation 45604d82-2d11-4c4a-84b6-6ef30e438afb · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.924430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.924430Z digest=sha256:b4c44426fc587226cee7dc18854aae4beaece6b2ac7b05231c00019473644562

Observation e2519ffd-99eb-4bb3-a109-e7e8707deb90 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.928235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.928235Z digest=sha256:6fabd3ea40fc58e415b29201b4b80452a82493c1dbf5de4e9587435a876a43ca

Observation 5362d29f-2207-4fc8-b396-47188baba437 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.932277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.932277Z digest=sha256:cdad3717e8576514012900b8ef573cae690679fd7cb39dd36464e90beca81ecd

Observation a5b4ffa9-9d0d-4952-a528-318a1175749e · outbound

This paper cites The Llama 3 Herd of Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.936125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.936125Z digest=sha256:62528f268f5cc46c927cdfe0f7398a93d9f321589f3eddc1733757aa1104da8d

Observation 38aa032c-7cac-4485-9382-77a7bd88c691 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.940399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.940399Z digest=sha256:0b6a09da061d3738970bfbe06df55f5ca6230c573588e5237aa4d63cca8b4ee1

Observation b62cacff-d5d7-435d-ad73-8a2573ce2ebd · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.944535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.944535Z digest=sha256:1f45428179d918237d9563acd89182fa4ce50e634dcdcb6e28dcd19ce9818c2c

Observation 0024ca61-e16a-438e-874c-caa354ecc0a7 · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Human-like Summarization Evaluation with ChatGPT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.948222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.948222Z digest=sha256:c90868be52222eca40d26b811b6b70ef728842bc4f10ba1d4b447c835d899db4

Observation e828a1d3-e343-4d26-bb7a-f2873e5832e2 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.611067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.952159Z digest=sha256:1354a264d39923f30fa18751cb87d76b6825394f6c951f68a62c142a58975bbb

Observation 0713a90d-c316-4a3b-99b7-5bc41834f415 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.955773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.955773Z digest=sha256:7a2d4c2d3394f1cac332658f941890728c6a9e538509b788f0ae09c51648fcbc

Observation 949183c7-e3f5-4fc1-8f41-806b3bf59bd8 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.598659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.959518Z digest=sha256:1ef473637a44ea937fb6fd58de1a8e58c4df21672c6b4457c8367ef47d749fd1

Observation ef45d9f0-f162-4f56-90f9-db7d5747802a · outbound

This paper cites Calibrated Language Models Must Hallucinate.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Calibrated Language Models Must Hallucinate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.963035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.963035Z digest=sha256:b458a3d839d199f560351b947d1465edb211f867093ea75f2d89004217840f71

Observation 3f7c0633-7a0a-4f16-b7f0-ee028ace3e3b · outbound

This paper cites RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.971420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.971420Z digest=sha256:2895a1ae0e48e1b3a599dc3d2a46a7033775d9e70f81aa003f74188b8cd63f91

Observation 87bc80b6-af2c-4e9c-bc38-ea8157032dd7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.975613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.975613Z digest=sha256:c5b5d4dd2fb051be7907072e894e48ee315938e9f33cac96e3b5d02bf2886e5e

Observation 14f963a1-910c-4938-9a68-620575c4cdb5 · outbound

This paper cites DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.979632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.979632Z digest=sha256:34cc7724ad4f73a44424a95b97d6c776f771a7bffa9b845a8a7b139dd70a2b71

Observation 6b7457da-6aca-4b07-8be8-af1734fffd41 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.983735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.983735Z digest=sha256:9863adcf012d76ddf700b889295ee594a5d80a01e2392ed0a9a80573852dc8d1

Observation 4e6bdf45-8520-4bbf-8dcb-e76b1121a180 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.987699Z digest=sha256:06af6357cccc254d1c198e307ba8c37996e797b7b355b3aedbf190ee1fa0d6b8

Observation b530734b-bf51-43c2-bb55-b2345dbb6d04 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.991382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.991382Z digest=sha256:564d909deaeaa11185c83fd98b8c64b254550c683602f4ebbdfee2a9b95c3f44

Observation 8dce8b1e-2382-4c61-a350-b2c549601292 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.995369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.995369Z digest=sha256:b00a915c18d1fa1966c1772e05052cd28c6d908463f43176bae978797a9e5bce

Observation 70e7b513-cee4-435a-b139-692ff894730e · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.999314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.999314Z digest=sha256:da4b6e698f1a3492329b3feeb43ac8c32380319c0dd5fa516ad0e31403b2b28f

Observation 74eb6b19-7226-4387-b346-8100999209c1 · outbound

This paper cites Filtered Direct Preference Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Filtered Direct Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.003269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.003269Z digest=sha256:202480da3ea00dfb9c04e011add1fe2dd48dd0f6c748de367cd7ba8100638070

Observation d2439acd-c0e0-4144-bd67-41d12f5b2700 · outbound

This paper cites Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.015126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.015126Z digest=sha256:84066bd84204bcf78b244db2849b5c485ac868886c09595fd2ceddabe16f6ca9

Observation ea4c1eca-3c78-4ff6-9d3d-acef119a17fe · outbound

This paper cites Text Clustering with Large Language Model Embeddings.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Text Clustering with Large Language Model Embeddings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.019154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.019154Z digest=sha256:2c3800d4dc4660492085b65e6d6fcdb488cb6c2a96ee5ccbd54f0e28bb754350

Observation 3f6616da-87e0-4958-b31f-57c03c22d5b2 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.023460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.023460Z digest=sha256:fbcab0f4ee5ea7cb1f540efdade8fb5ecae01f20ddd606dd2f053136df71b6a6

Observation ecf408ba-8b58-4cb2-9a53-57592c1747f5 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.027162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.027162Z digest=sha256:6ba75193c492b81fffe3de1e0eef9c2025129999e96960b97d4f166396906f2f

Observation f0c3a59b-ed83-4baf-9d37-c742de0a9c58 · outbound

This paper cites Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.031609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.031609Z digest=sha256:4cb4b4be39b8e8152a35acff3cef3d1ffb3938e47cec575e24e2d585932b5e62

Observation 47b066ec-1053-4320-a8ee-bed034af49e5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.035918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.035918Z digest=sha256:43381338e9e4786baabd4a49ce25dc4139952bd417af5061643cacd4a6e13e19

Observation 96594f66-9ded-4b99-969d-71a73a8165ed · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.040457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.040457Z digest=sha256:f55f6382913d270e8250adb57b57585dce3c6d8f90307031cc7bfb4fffda1d6f

Observation 086b6d6e-edb9-4462-88db-ce08ad0f4fec · outbound

This paper cites Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.048328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.048328Z digest=sha256:509256709e584ea1b7dded94307a9cf6064f504499320b8d7750813dc7632052

Observation 42734856-6fdf-4619-a1cd-ca674320fcc5 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Large Language Models for Data Annotation and Synthesis: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.051944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.051944Z digest=sha256:a8d02e3e75bc3d0aaceab8a36bff9d9eea8d4d7ab1702366702b184f07ed1b25

Observation 9151b43a-189e-41c4-95cc-b9cf93321c97 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.055697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.055697Z digest=sha256:774ac7b07e18f6ed9ea9fd8c53a011032c19afb79c08b32facfaa365677296b9

Observation 600b09c2-b0cf-4a80-84ee-201ea6d4c4f8 · outbound

This paper cites Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.059688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.059688Z digest=sha256:8e566bec00cfac8873f9156e784d1d9f859c555a8b2842418b44366711a3ba1d

Observation 68d49670-edac-47e8-ae30-29b0204c216b · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.063986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.063986Z digest=sha256:83cf39863b8523c97fba6d8a5aad862533ad81978cd2d2880e51a0dd09211845

Observation 7052df9d-d093-454c-86e8-712b6a15f902 · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.068426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.068426Z digest=sha256:31e513300f5e0d310fdfb4ebd6915a68ea146aa41a92caa08060602045f18341

Observation caff5bb9-a58b-46fa-8c55-8b9cfcb40e0c · outbound

This paper cites Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.072615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.072615Z digest=sha256:e3a69f8e95e79515ef744256a8a573c023b5fcf677639f74f6b122f0abafd216

Observation 31669112-6165-4d6a-9289-87bc1809bbc9 · outbound

This paper cites Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.076717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.076717Z digest=sha256:8d0b1b36857561eb254bb04e4b0c31fc07cf58bd540d556c323dcb4dd8bb38dd

Observation 309eb1cf-e393-4d62-bbec-c2ccdf9dabdf · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.080793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.080793Z digest=sha256:c8daa0a0b7d14c16cf0e1daec082900a70adba58f42e5ccbd74ea4525f9fea3b

Observation 9956b1d7-86f8-4c61-8146-b2d4e0f16354 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.084737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.084737Z digest=sha256:7973a741a6af8450119db3eb939f7b37eb19e7858de5f8d1ccfce50dc27abf5f

Observation de50be09-0fa0-4baa-b30c-f4b0c619a46a · outbound

This paper cites A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.088657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.088657Z digest=sha256:5446251d33681a8baeca78aaa22412d488acc4336873965148ddb434fbfdd738

Observation afef957d-bbef-4f72-8fc9-050b3d87cb66 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.092585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.092585Z digest=sha256:e08ea8424ea27a106b78da272d43ef37d9408ac1ede6b4b7a5e63ea9f28cac7a

Observation 9aab5a24-06fb-42b2-8d4b-28dbd8fdceeb · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.096279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.096279Z digest=sha256:1211ef8cac4d95c300f79bb94e66f6f12249b003af81b0ef6d489bc96e1cbab4

Observation 4cf83a48-672c-44aa-802b-c3476bd48e78 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.100583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.100583Z digest=sha256:22b27fb67aebf3d6e4b225bdc1ecace19c53f1a8c67510b996da762fed7365f0

Observation f1cca3cf-1002-453b-8444-bfbf03024b56 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.542635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T19:14:37.105579Z digest=sha256:2dd20a2a7dd2100da192cc57184dafe5f3f84c4198275d70ed2e07e5930ba316

Observation 58f29db0-07ed-4283-88e8-8a00c4d0fe53 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.109155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.109155Z digest=sha256:e955298e963bf2aef7850d776b23dd0f17c90ea2dd4c92a5c62566b55aa5cf3c

Observation 70199fa9-37cd-4829-99cd-5b4eeb78834d · outbound

This paper cites online" 'onlinestring :=.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.113077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.113077Z digest=sha256:f67717f1980638bb00e615b3a002da9ba5b6067e401a015ec92fdc0400772bd2

Observation 92f1c020-bf53-45b2-8905-2631d69a296f · outbound

This paper cites write newline.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.117468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.117468Z digest=sha256:1b68781feb5ef0307d45a8ecfaf9f6cc989fb175081e4cf2427dfc568a5178a4

Pith citing papers

Observation bfa35949-57d6-484d-b57f-992a1b64dab6 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.142754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:ff5cd4c5864a48e1ae80587f97c6bc7b9844c29b4a6b351d31a6a419f3e2120c

Observation 32eca10d-4bae-43eb-851b-42e3a8ed6445 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:39.875926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:39.875926Z digest=sha256:68097f2258fadd41a7182d9c53a08c82ba3b2de2fe1b45c377e352188e0dd6fe