Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T21:22:36.970101Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2408.15339.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T21:22:36.970101Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75e67232-a357-44c1-9fe4-90ff766f4b75 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e57c9ca4-ccc4-4766-9abe-a640002eceeb · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022a
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c6e65cd-25be-4678-8647-54495b3e903c · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aedee061-b3dc-493a-893f-f2657d3ce1f8 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ab1016e-dcb4-40b5-b68f-a0f1c667a2ef · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Measuring Mathematical Problem Solving With the MATH Dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc63e7ba-1101-480c-a20d-a745edab1d73 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types ORPO: Monolithic Preference Optimization without Reference Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95e49e35-dc8e-4cd3-b648-8af0b9c58686 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LoRA: Low-Rank Adaptation of Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f32f8a9f-0d18-48f9-87c8-7c7e16123b04 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Mistral 7B
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ebd1d91-ecbf-45fb-bfd3-e5688bb9c815 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LiPO: Listwise Preference Optimization through Learning-to-Rank
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 526fd740-085f-4992-a9f9-b3d9e4eebd52 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nash Learning from Human Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3328669-db3d-4546-a3b9-058deafc9d92 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nemotron-4 340B Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 710e60f8-dc94-4cbb-bbe3-e01e1a82361b · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b9ffa27-5a71-4bac-ad08-adbca58e3303 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93c2f85c-31fb-4c8a-b53e-f10c6b3e8cd1 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fca72ce2-2f53-4be1-be10-2196cf86329d · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Offline Regularised Reinforcement Learning for Large Language Models Alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6003053-d699-432c-98ff-7e1f2f765e5d · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0894764-c8cd-48c4-9789-8e388b060228 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Preference Ranking Optimization for Human Alignment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 391ef5c6-cc14-465a-8d28-91bc4e34b19e · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3de95e35-1551-479b-9295-867dec474dbb · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Gemini: A Family of Highly Capable Multimodal Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbd8afec-9ea4-41ba-b087-9325b32ea148 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cac8fab4-4a54-4d55-a42e-1558f0187b6c · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Play Preference Optimization for Language Model Alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3107b263-5fce-43a2-99a7-30228e15d883 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7d6c5fd-dc09-4622-953b-957c1f47e449 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Rewarding Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8fa9f457-f4b5-46d5-b3bf-c65a6b2866e0 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 078580d8-a1a5-4ca6-b1db-d8247483d942 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Token-level Direct Preference Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 737e8445-6ce5-4896-920a-334a477100e6 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0c7040b-6999-42a1-9f27-9b5f1973a6a8 · outbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Instruction-Following Evaluation for Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.