Pith. sign in

Paper Citation Record · LEDGER

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2411.10914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10914 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:37.117468Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:37:39.875926Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T06:12:07.140554Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0ec019e-8f45-4012-b076-58691db7ddeb · outbound

This paper cites A Survey on Data Selection for Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.907006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.907006Z digest=sha256:e503797292d3c7003fc5ec2b5733df37b0496606f0bda375af16cae8e7bcc1ce

Observation c3154b58-ab12-4254-8c5d-afc49900c419 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.911802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.911802Z digest=sha256:8bdccf5b09db2efc5a4528d89b485b0b54c567a7a055e59c7081e36ea5a0336d

Observation 470b6e01-bc63-4cb0-ac21-1cd26bac9118 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.643972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.916115Z digest=sha256:e866e8eeaf14de0c81b100367a25ee3fd7ecbc5bc06106dd68f1bbdda61d76f1

Observation 93c2bbe0-6ec2-4180-ab29-6b3781f3c3df · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.919782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.919782Z digest=sha256:4a28bc0b7770c6e34492ddad07a604fff556bfe5142c905007066d347a4fb6b3

Observation 45604d82-2d11-4c4a-84b6-6ef30e438afb · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.924430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.924430Z digest=sha256:52f7e6f337e37a228c4b60a83b0417216f430c4c459543282f0c2145d630dd9b

Observation e2519ffd-99eb-4bb3-a109-e7e8707deb90 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.928235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.928235Z digest=sha256:3bd2073129bbad37a2e9742f3d2a37f0615e780fbc5a64fca327a0e5cb790145

Observation 5362d29f-2207-4fc8-b396-47188baba437 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.932277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.932277Z digest=sha256:11fe95e2fcdcaf30b4d0387ff1d0aa25c05e30521dd0bc3d053c287429bec9ca

Observation a5b4ffa9-9d0d-4952-a528-318a1175749e · outbound

This paper cites The Llama 3 Herd of Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.936125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.936125Z digest=sha256:b7f4ba350cfb3a81f52f2da02d3f8dc871318527ac9e531066d8675c8bf24f5d

Observation 38aa032c-7cac-4485-9382-77a7bd88c691 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.940399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.940399Z digest=sha256:89595c10a45f127debf6fcaff0f4091863a313493494859cd46c7a502e304a9e

Observation b62cacff-d5d7-435d-ad73-8a2573ce2ebd · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.944535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.944535Z digest=sha256:a936904649ac3b36062b9785e00b3b96a32c837f21bba4fae69af0fe192d5440

Observation 0024ca61-e16a-438e-874c-caa354ecc0a7 · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Human-like Summarization Evaluation with ChatGPT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.948222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.948222Z digest=sha256:111d29d6eedffeeca3c1be054e592b50859e1134c276f815e159d76428172cfa

Observation e828a1d3-e343-4d26-bb7a-f2873e5832e2 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.611067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.952159Z digest=sha256:3cbb20e28625dc6c5a0488c740bf1e6e121e9a3268077fdb46addb93e7c7f3bc

Observation 0713a90d-c316-4a3b-99b7-5bc41834f415 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.955773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.955773Z digest=sha256:3c87db5af0b6ea8a6d1365772e1f552435af7a9bd148bbae80a8e22e6504bb9f

Observation 949183c7-e3f5-4fc1-8f41-806b3bf59bd8 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.598659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T19:14:36.959518Z digest=sha256:3b5d48e7594f17a5026a7256b56b1b7b7b60cc9b81783b98fe83b0f3d0758f54

Observation ef45d9f0-f162-4f56-90f9-db7d5747802a · outbound

This paper cites Calibrated Language Models Must Hallucinate.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Calibrated Language Models Must Hallucinate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.963035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.963035Z digest=sha256:52340d744933ea5ec870d6b9ec1b6ecae4754e6a1c493ce1e7ad7c8069ca5cc9

Observation 3f7c0633-7a0a-4f16-b7f0-ee028ace3e3b · outbound

This paper cites RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.971420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.971420Z digest=sha256:ef61a959d8bc3896407a01d42b62e10950d0660b8b4df23c679ccec58c17d5eb

Observation 87bc80b6-af2c-4e9c-bc38-ea8157032dd7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.975613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.975613Z digest=sha256:4715b9b2b18455f0d2bd7d0e5051683a44e0795595e31c81b2e4985817a17f97

Observation 14f963a1-910c-4938-9a68-620575c4cdb5 · outbound

This paper cites DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer's Disease Questions with Scientific Literature

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.979632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.979632Z digest=sha256:82539f155b6e514fb19707ac72592d94d14aee42f4a1c1dfd58f645a35803c8a

Observation 6b7457da-6aca-4b07-8be8-af1734fffd41 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.983735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.983735Z digest=sha256:e9702f7248c1a3bfeac1b1c629c7a1bc5da558b918d5dd541dab50c60221aaaa

Observation 4e6bdf45-8520-4bbf-8dcb-e76b1121a180 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.987699Z digest=sha256:190112c953f8a591fb4a8280126ddc26f820fd0d57305eb4b32ed8b6c1803118

Observation b530734b-bf51-43c2-bb55-b2345dbb6d04 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.991382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.991382Z digest=sha256:f142c55bf6c79e92f36dc267afe9db6a1a9ce0722bcdaf743c4e347c3ddf6f29

Observation 8dce8b1e-2382-4c61-a350-b2c549601292 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.995369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.995369Z digest=sha256:2601bb5cb6f217ef1244702781717ea114a919a9c532db55e9759bac6bb43b86

Observation 70e7b513-cee4-435a-b139-692ff894730e · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.999314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.999314Z digest=sha256:69d04157d82c1ad2fc63bac68fde488b7a75d366ebd76cf5bdfe9c7bbcf5fc60

Observation 74eb6b19-7226-4387-b346-8100999209c1 · outbound

This paper cites Filtered Direct Preference Optimization.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Filtered Direct Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.003269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.003269Z digest=sha256:ff7eb736975d07883b05c8b28382bb89a11f5ca03b69976b17ab842f595dea4c

Observation d2439acd-c0e0-4144-bd67-41d12f5b2700 · outbound

This paper cites Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.015126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.015126Z digest=sha256:12f7089702b58f1fec305913ebb5a3aedaaac8468227735e78e9fbf6203275d0

Observation ea4c1eca-3c78-4ff6-9d3d-acef119a17fe · outbound

This paper cites Text Clustering with Large Language Model Embeddings.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Text Clustering with Large Language Model Embeddings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.019154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.019154Z digest=sha256:45548256b88ed9bae4a9148b02c6e691f0b8535315efecfb524de74a022e23a5

Observation 3f6616da-87e0-4958-b31f-57c03c22d5b2 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.023460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.023460Z digest=sha256:a7047b845557bd69af51e9dbe71cfb67ff8232c95ad331cca10fcc7b3fcb9516

Observation ecf408ba-8b58-4cb2-9a53-57592c1747f5 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.027162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.027162Z digest=sha256:d3e9428e8b676f5d6d1ceac41bd01d4a91ac1c2be72313a3e2be4ceb621a9799

Observation f0c3a59b-ed83-4baf-9d37-c742de0a9c58 · outbound

This paper cites Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.031609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.031609Z digest=sha256:3234fba8dfae2a2a39a6b8e5ae0f95441e33fa701f62545194a20719b89d14dc

Observation 47b066ec-1053-4320-a8ee-bed034af49e5 · outbound

This paper cites Proximal Policy Optimization Algorithms.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.035918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.035918Z digest=sha256:9a6d4aa3fb5b9140ff5e8911359839879c17347bdd23000eb16e7d32a3a99579

Observation 96594f66-9ded-4b99-969d-71a73a8165ed · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.040457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.040457Z digest=sha256:eaab7a9da8381032f195a487dbddf2b8fc1500637bfb18f64668152ad06f0367

Observation 086b6d6e-edb9-4462-88db-ce08ad0f4fec · outbound

This paper cites Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.048328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.048328Z digest=sha256:6dcefa6eca0a74ee7cbd5d3a1ec4ea43bcb873caf2687c31458bdc9e4dfd59a4

Observation 42734856-6fdf-4619-a1cd-ca674320fcc5 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Large Language Models for Data Annotation and Synthesis: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.051944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.051944Z digest=sha256:2654fa080cb058c8dc8664bf770eaa94fb678311cd9334aea77562fae46f3df4

Observation 9151b43a-189e-41c4-95cc-b9cf93321c97 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.055697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.055697Z digest=sha256:e03a6aca48e6b9613788302e27c7804460d5ac2f83e2b7e41e7f427c33f13bdc

Observation 600b09c2-b0cf-4a80-84ee-201ea6d4c4f8 · outbound

This paper cites Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.059688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.059688Z digest=sha256:8342d47c750cee8cbdad260da287d12664690bdc3300dd000a72d54a2ca073fc

Observation 68d49670-edac-47e8-ae30-29b0204c216b · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.063986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.063986Z digest=sha256:61d799dedb9d6a6b2430361e232ab075fe81d67d32e41c74ddce12a3a5df8ae8

Observation 7052df9d-d093-454c-86e8-712b6a15f902 · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.068426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.068426Z digest=sha256:f6634a3784d8203af5c9d1795008dba24bbaf1c7b12baca0a5e1f0572ebc0263

Observation caff5bb9-a58b-46fa-8c55-8b9cfcb40e0c · outbound

This paper cites Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.072615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.072615Z digest=sha256:fbab1b386998fcb0d645abb8066d1a0a6dd8727c962fae559c26830a82fded55

Observation 31669112-6165-4d6a-9289-87bc1809bbc9 · outbound

This paper cites Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Balancing Speciality and Versatility: A Coarse to Fine Framework for Mitigating Catastrophic Forgetting in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.076717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.076717Z digest=sha256:296552c89f5434f24a4c5a0b4a9a0447f6f6a11b3a76fc29719fd75dc97c0dfb

Observation 309eb1cf-e393-4d62-bbec-c2ccdf9dabdf · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.080793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.080793Z digest=sha256:38c6c4e7fa0395c84ddbb13dae9be408193ac1223be9365c31eb09e13f9ac114

Observation 9956b1d7-86f8-4c61-8146-b2d4e0f16354 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.084737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.084737Z digest=sha256:a8c0a68d3594415bf07234a87f198610b0088a408dab5aa284e0ab513cb4b558

Observation de50be09-0fa0-4baa-b30c-f4b0c619a46a · outbound

This paper cites A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.088657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.088657Z digest=sha256:e5bcdc13cfc9df5b941cfbb81a2d195a85ddac18aff9b7aa96c7887bcb001a42

Observation afef957d-bbef-4f72-8fc9-050b3d87cb66 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.092585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.092585Z digest=sha256:c62f2c4cbe16198d7c208f88f476378e002f166604a462ab069b415134caf729

Observation 9aab5a24-06fb-42b2-8d4b-28dbd8fdceeb · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.096279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.096279Z digest=sha256:7b1f09169680be9d80ddc9a4b780b4c3bea6afc16dde663b46258249cd2022bd

Observation 4cf83a48-672c-44aa-802b-c3476bd48e78 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.100583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.100583Z digest=sha256:7e023e3adf05fd8442a9199d757c2c7b8a2a1ad3d9caf5af2663f3474e760e8c

Observation f1cca3cf-1002-453b-8444-bfbf03024b56 · outbound

This paper cites an unresolved cited work.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:14:37.542635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T19:14:37.105579Z digest=sha256:08839d2ee70bd5fc33a09c1d39c14b4522afa89282e6e5d1edeca1cef39a7d35

Observation 58f29db0-07ed-4283-88e8-8a00c4d0fe53 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.109155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.109155Z digest=sha256:b1e77e13d081f2f7ea6042c56e9f8020f197ea5974420105fa5d4c90a185b29d

Observation 70199fa9-37cd-4829-99cd-5b4eeb78834d · outbound

This paper cites online" 'onlinestring :=.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.113077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.113077Z digest=sha256:7a2460fb7e031164cde5d9b5a2f9cd6b57612856b6755ed81a2f2c862790846c

Observation 92f1c020-bf53-45b2-8905-2631d69a296f · outbound

This paper cites write newline.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:37.117468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:37.117468Z digest=sha256:f1c827fe7b9da499dc45d096c709b9794e6819ff26ac695fbb733dd05230363b

Pith citing papers

Observation bfa35949-57d6-484d-b57f-992a1b64dab6 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.142754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:130f5b924ddcc44cf77570842539655459e0528d922c4cb27c24ae8637d83586

Observation 32eca10d-4bae-43eb-851b-42e3a8ed6445 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:39.875926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:39.875926Z digest=sha256:bf929b5591802b55e478664a0f732d0a7ece2ce4be13735a3a46f551ac3d9733