Pith. sign in

Paper Citation Record · LEDGER

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2509.25148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.25148 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.258118Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:18:58.373323Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f85890b8-4fab-4908-b811-2d4220c38e01 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Harnessing the power of llms in practice: A survey on chatgpt and beyond

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.792700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.061817Z digest=sha256:f87c47b12a3e3749de14bb828b3dab7c9934791cd89446b8353676946afcb975

Observation 7e4e84ab-65ee-43a4-9e33-dee86a050071 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.065373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.065373Z digest=sha256:177d53d5418ead39d9d54eff67c0c15d37a503fd96ad24b1f2cc56354c12d509

Observation d9048366-9d1e-4557-b7db-57d92e51a4dc · outbound

This paper cites \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:49:58.605246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.068948Z digest=sha256:6839e852f84dcc3832e3eb2b4dc685f8c61da0589ab60a0bafea4ad4b28a94a1

Observation 3e67ee99-bf16-4aeb-b499-937ff5aa3806 · outbound

This paper cites A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.785712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.072181Z digest=sha256:d25fd168a9ff87b06b1e1a5a5ca764e8d7bb38abeb6944a73cc15378f629a181

Observation 4fdda5c9-fb6f-4973-82ea-31fa4c32784f · outbound

This paper cites Training language models to follow instructions with human feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.075991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.075991Z digest=sha256:53ed08e48ec0509d92f6d2d3434cc15a4ccf84370fe3bd12453ffb1e7c0aa9b4

Observation 4f8a0d64-a97c-4051-83e3-f8f4f95bf819 · outbound

This paper cites Deep reinforcement learning from human preferences.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.078529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.078529Z digest=sha256:7804bfa843ca2ecbc42a066a4ab842259c01e8234487489827c6f3d7a30b165e

Observation d98b61c7-d0b4-4444-8999-b0f0948a7d50 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.081745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.081745Z digest=sha256:1600543171b5f972d96fc4d691133679aabd06ca7b58fbcac9d2d21136216343

Observation c8aec489-bfc7-4e0d-9795-013e3887e282 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.084732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.084732Z digest=sha256:056b220f4a1b2db412bb1f764b69ab9c2b942012adc8aa6481e6045011776462

Observation 56686d20-3a5d-490d-bb73-6df25066f822 · outbound

This paper cites Curriculum offline imitating learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Curriculum offline imitating learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.768174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.088393Z digest=sha256:7cac6502d26a81984ef291b319329c5949c70e498f24592dfefa38c40fa6c0f8

Observation 11144b59-b2cf-4452-a379-532cd0493a36 · outbound

This paper cites Offline imitation learning with suboptimal demonstrations via relaxed distribution matching.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline imitation learning with suboptimal demonstrations via relaxed distribution matching

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.761477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.091500Z digest=sha256:85f3b48355fed9b9274131cddde37450b8e7a047cdb19855a6ceb4819db1cb38

Observation e38fa1cb-b58c-4008-acaa-fc7a3123f3b3 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.094264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.094264Z digest=sha256:292d70b0f94b81883b5de360cff45af28d97960b16aa2a275e3493e494305f1c

Observation 1cd45998-58e6-4d1a-b03b-681575012957 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Grounding large language models in interactive environments with online reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.097622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.097622Z digest=sha256:ab02b5dc3aff6bf5eee17b719846dbe85bb9b2ed8ad3d3af442dc890de9e3f73

Observation a34bb80f-a6e5-47ef-b26c-5de71dbef4dc · outbound

This paper cites Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.751138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.100869Z digest=sha256:2e817c7c9b6e3e3433ed26dd37211b668d8dc1ab60f39bb5cb2ab040405799cb

Observation 12dcd533-12d3-4efd-8ae2-85c5e8b1fe88 · outbound

This paper cites Preserving Diversity in Supervised Fine-Tuning of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.103306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.103306Z digest=sha256:cc6ba3aa90bc202fee2e8a0c31448d7798b8e977c141e485e94fd70798cded77

Observation 070ec0c1-259d-4d4b-8a41-30d2908456ab · outbound

This paper cites Regularizing Neural Networks by Penalizing Confident Output Distributions.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Regularizing Neural Networks by Penalizing Confident Output Distributions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.106522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.106522Z digest=sha256:0820ef323cd0101b0d58c2c35b753dcfb2ea3a759a529bd054aca89a36329633

Observation b801e931-70b3-47d3-9822-e7b9710433c2 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.743340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.109290Z digest=sha256:8dfb6531cb3625dbb588a5e46d38201bb4134ffd3dbc28a3af5708bf704c3b60

Observation 2b49f18d-da45-43a6-b2be-234497283dd4 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.111627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.111627Z digest=sha256:0683b4ee822403e4f3d6cdd6350a4cb79b7d000b7a4f6e09ecbf1cc21341fdcc

Observation 2316bf61-624d-421a-83bc-f0cf68cc97f7 · outbound

This paper cites On the diversity of synthetic data and its impact on training large language models,.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the diversity of synthetic data and its impact on training large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.736182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.114383Z digest=sha256:a9224cb014787bfcb7ed5154d3a3b9375fb1fb5cc58210a629a808d2d7abcde3

Observation b8b1e8c8-d5b7-413e-9e0c-7d90f7b60170 · outbound

This paper cites Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.728110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.119357Z digest=sha256:c53edd2213e0be6e5a04ef5816dc54c70a5025db9fda86ea2a420a5e96e94743

Observation 938e0ef4-961e-4a9c-9218-bd996afa1a9b · outbound

This paper cites Idgen: Item discrimination induced prompt generation for llm evaluation.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Idgen: Item discrimination induced prompt generation for llm evaluation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.721234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.123053Z digest=sha256:d9640f7156e58d6ef46ced004fac1963b2ba48d3140edc956b023d65f3aa9bf3

Observation d27af825-ab83-4627-be9a-6228d708508d · outbound

This paper cites Qwen3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.126327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.126327Z digest=sha256:57f4423eab9f90f66facd9a6512850e779bf995d33bf0f8050a39496d5293b8b

Observation 08e3b7df-4394-48c2-85f5-199226c1483e · outbound

This paper cites DeepSeek-V3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.129064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.129064Z digest=sha256:d5095765a51a4e6b4626379c12a4e67cb502a348486864e392c9e6d2a94b5e24

Observation 4fcddce0-5c59-4725-bca2-b15566824e6e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.132125Z digest=sha256:f2d4c93d2c3b4cdd01492dc452272e66f8e32b60a61c64b577ce26537073f642

Observation 98c886ce-718d-4454-8c6e-ccafe91ea064 · outbound

This paper cites GPT-4 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.136271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.136271Z digest=sha256:bd5b0148b464064bae96cce8ae046b237189d1775a34b3d82c12ac28f1f1d037

Observation 23866afd-1047-4d4b-b9ec-a3a32e138b27 · outbound

This paper cites Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.138798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.138798Z digest=sha256:578db389ebf0ce8a3064a1a787df0b6df835661e73413ec2ab8a4ce91e243f95

Observation d439eeeb-0b36-4daa-8288-a0f1779a44bd · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.142255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.142255Z digest=sha256:51f65f7cfd6dc9db7169ac497e6c5d9ce97e7aca519dfb6a43c1b39d41a5b3ff

Observation df333a4f-9aea-4408-801c-1d20fc8e213d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.146538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.146538Z digest=sha256:0921003871578ca5ffb8acf6b1226175cf3aa1f3a199da4e6f822f41c5c0f216

Observation 7c7ec438-faeb-4e0b-ba3b-1a95d3c8c691 · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.150276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.150276Z digest=sha256:64f4762768a910ccfe0974a718d2af340ac289bdee422a8b014569a305043d96

Observation 028c0ad8-5942-4d92-9c1f-aaa47d435aab · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models AutoGLM: Autonomous Foundation Agents for GUIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.152922Z digest=sha256:3df6e1cad636bb694a783e76752a6fd55c95c556cbae1d40b99faeaa416e5ce6

Observation 7ebdcc37-7fe3-49c2-b0a6-2d60a82ad4c6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.155753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.155753Z digest=sha256:66c5009025f7239d432acd95b338f400cfbe32e975e648cbeef206087c55e436

Observation a7faba00-97d3-403c-a438-df0d71050587 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.158558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.158558Z digest=sha256:7b335d68792802b6a3294a2df7fa593c0f9f1c66aeb562b812dcd37f44b55d12

Observation 04c91798-d270-49d5-8e48-98e58360054b · outbound

This paper cites Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.162168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.162168Z digest=sha256:2a2af59f6b15f7662a13c52e494e763cf2c8c2401aff51bad38fa36e823b7df1

Observation 98ef0e3a-b85d-4bb8-9e5d-8a6ec1c0af87 · outbound

This paper cites On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.165261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.165261Z digest=sha256:01d5a9552317d2f6a46255f81b56c97036f2391b89e6b83bbb177b5415afc54c

Observation b3ecce42-6552-4fbe-b36c-1906c1d830e9 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Learning to Reason under Off-Policy Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.169948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.169948Z digest=sha256:e4c3dce5e9c997b946cfb0bf26c136f5d34f0cbdca0873bde56f40c539ff26ce

Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · outbound

This paper cites BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.173572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.173572Z digest=sha256:bd37392ee66ac113b33d99c891ad87c3dace64e12a817e66fdeb1068df3be052

Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.176602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.176602Z digest=sha256:2754bc3286d0efe0fff8a5d9b15c7b953f6630ff9dc2cd026b2f969aa33e1dba

Observation 875a2f1d-728a-487a-90bc-606c0c1d27da · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.179658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.179658Z digest=sha256:e7e99e9e70151cc63879205418abb88e79dd2a6c0becb11ad8d2085bb33271a8

Observation a198c5bc-aa3d-440b-bb9a-8d46304380f5 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.183190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.183190Z digest=sha256:ed94fc2e5ee82b71ec1a6ef52244c19a709959959337ef73e2fbe35fe517ea32

Observation 54df15d6-af5d-4049-993c-1bde54bbe74c · outbound

This paper cites Generalizing verifiable instruction following, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Generalizing verifiable instruction following, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.185865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.185865Z digest=sha256:fe4944ef5f2186d7416711b97b99670875d451cc6007eb148a40731de939cc4d

Observation a762c342-dd82-4265-8ff4-5eedfda017f5 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.188396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.188396Z digest=sha256:f70a2f874147561f2ffd487ae71c251a9216eab95dd19b5da109ef56f3223405

Observation 92da0776-98c2-4f65-8985-74d5e47917c4 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.190607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.190607Z digest=sha256:c17172759007faac91b340a14c00088f5355f87212e8a2cb104d5f20467ece7c

Observation f85c6f8f-bac3-47e1-b658-f4e938b68f5b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Measuring Massive Multitask Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.193411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.193411Z digest=sha256:f76d82872025eed1fdd87ec013af92030284ce2a6bdb4482848ec0a4a01b510d

Observation 9ee6b06f-52d5-4c6c-b909-a4c1087c866c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.196925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.196925Z digest=sha256:f25fd04c571be47e00ca829ea8d9804fe88eace8157b7f96090e3549019d7ee9

Observation 804d1d4c-4045-4555-a32d-91e48d1258b3 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.199271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.199271Z digest=sha256:865889bd7cbf107d3fcedc7556774426b92ee0a72f912929e5a340cbfabe9742

Observation b947f471-b0f1-400f-af49-1deb956d0934 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Evaluating Large Language Models Trained on Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.201429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.201429Z digest=sha256:bbf3bb0f0597ff50c11aad3a79235e8f827e5be0f85a799af1be39c8ad6535d3

Observation 0979e746-9fab-4609-beee-71afeb3e1c52 · outbound

This paper cites Program Synthesis with Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Program Synthesis with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.205422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.205422Z digest=sha256:673a9d1fd9d9b4a035d67122d0b95f59cbf692c74a17aba12c9eb9e59751d606

Observation 6005055c-4816-4fd3-97ff-ab9624d64ae7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.209138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.209138Z digest=sha256:5b2156c91b94b78dcd90a91d45c97f9601a29f834475e3cac58a2db015ad7a1e

Observation b7ace46a-de22-4611-9396-116de9b0e6f1 · outbound

This paper cites Let’s verify step by step.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Let’s verify step by step

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.212019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.212019Z digest=sha256:f7f45d8be1064b9c2a5b37bc8bb64001bf3686668cf9e2916a8cff8f0253f09e

Observation 82d5c9da-f8f5-4afa-8ba3-be80aee45ae7 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Theoremqa: A theorem-driven question answering dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.699052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.214895Z digest=sha256:bc3b8ebc7bb3f6e12ba571e5c650a848da88e842e5fe4399e14cc523252de643

Observation 62e6de16-ca27-4c2d-9a22-65beb8e7d1d8 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.217412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.217412Z digest=sha256:19205e32403e96cb97548dd56115ff8c02bdb2894e9254fe557fa10047e5d37e

Observation e8fcaa28-bd44-49ef-b07e-4760195a12b7 · outbound

This paper cites C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.688924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.219886Z digest=sha256:b6611c6c70c7de650f055f747782a72fc24193678dd1447b4702455d904b78e9

Observation ebb3d37d-99fb-4248-8f04-9148819eb483 · outbound

This paper cites Pre-trained policy discriminators are general reward models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Pre-trained policy discriminators are general reward models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.222765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.222765Z digest=sha256:d67612c410e268992b793016ba3e69d71d8559742506d03a9e6fb830258a464f

Observation 324a8893-47d2-4c6a-a407-5dbb8abfa522 · outbound

This paper cites Swift:a scalable lightweight infrastructure for fine-tuning, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Swift:a scalable lightweight infrastructure for fine-tuning, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.681698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.225967Z digest=sha256:c8ee41b1ebfa1be64e790ea5a5ae12a5747002e1ad41d91901f225c7711ac196

Observation ac153536-d7aa-48c5-9f68-870b934ff8ad · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Hybridflow: A flexible and efficient rlhf framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.228336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.228336Z digest=sha256:6e8e0e979b1386b5a827c219221bf2de004986198b10734c9a9d27dbbb2a9c14

Observation f6522c83-15cd-4dc4-8a16-7f7d405c00b9 · outbound

This paper cites preference.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models preference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.671519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.230938Z digest=sha256:d3b085af3b31aae965dddd605e019796db2e64d541d3192f25f56c611d488e7b

Observation 5c6a21b2-369d-4393-95a4-8a8e50d8b392 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.663350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.234960Z digest=sha256:aa2b81027086aba66587d9b22b52fefdc3fd142b21c8640c5a5691333df22deb

Observation b4ef34fe-e198-4a5d-950e-33621b0407d7 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.655942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.238143Z digest=sha256:2c10a3c916515e6331a9a9c17efff7b5119751b88a782df62f72bf6e8a3f9cff

Observation 35491afa-7d8e-49b3-b244-848509f433bb · outbound

This paper cites Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.646748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.240961Z digest=sha256:a790adacad26f2107e33a47367db429807f9dfdef37010a7f0030f9756b716a1

Observation a999e21e-f008-46c7-bd41-fb96a7093c29 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.638677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.244479Z digest=sha256:b18af07dac6cfa54fb1739ce89c7158b9edcb57f80f3864666ad10b2ebe124f1

Observation 64913ba6-5505-4e22-8bbf-9bb3998ba328 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.632107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.247254Z digest=sha256:48c872868fe32abaa3137faf133572d10dfc900ac939a6610382422ff0c5a532

Observation 9bd36c44-6590-42cb-863f-e34b69f0c504 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.625194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.250724Z digest=sha256:978b5b9ccf890eec97d1708947a08d70eb2506700962db254f95092c9b8b522e

Observation a35edc2e-255f-43c8-a8bc-2e8c9c63e1c1 · outbound

This paper cites This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.618467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.254776Z digest=sha256:7ab898a8a16e10e9dfdfb974906c46b05988c7f957277210ac89a677593d4d58

Observation d01692d6-8b7b-4e18-935f-1834b976261b · outbound

This paper cites winner" (preferred) response andyl is the.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models winner" (preferred) response andyl is the

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:49:58.330875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T15:49:58.258118Z digest=sha256:bef6921d54b14eb4b09b56fdfc29c9ef6575701f014dd37c75742caf300f6eb2

Observation 829c96dc-3465-4610-bb3e-b7248a0a1296 · outbound

This paper cites On the Diversity of Synthetic Data and its Impact on Training Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.116911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.116911Z digest=sha256:47681e04ce05604a06857b31f33714a9609fa5146f9651699523416fac7c8b7a

Pith citing papers

Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.374761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:e366b36ca7ec436a1cb23d54cabb43cf8ea0919838d4a88b0fc31017c7e4278b