Pith. sign in

Paper Citation Record · LEDGER

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2408.15339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15339 v4

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T21:22:36.970101Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact21
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75e67232-a357-44c1-9fe4-90ff766f4b75 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.397187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:668e56845b5c528bca5a3bd644eb7fbd1c9311869a5c88a4da2df1a1de8e2e1b

Observation e57c9ca4-ccc4-4766-9abe-a640002eceeb · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022a.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T21:23:28.456889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:2d1a48d1aa9f9665c968b61ce306b9c58e333200004dee6c394456cf82c358be

Observation 4c6e65cd-25be-4678-8647-54495b3e903c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.366013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:357b6433e9d5ea7506202956101cac423e53065ad762916e2fb0929b7a6abb6b

Observation aedee061-b3dc-493a-893f-f2657d3ce1f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.371734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:f87e1ef4c5efec7524718ca7f546aebc5386b22aa4ccf355955e31c272f613f9

Observation 6ab1016e-dcb4-40b5-b68f-a0f1c667a2ef · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.424359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:8c52985bb953694bb1c8187c9a19003647e1bff6b6d397ac8a0f0888929cd78a

Observation dc63e7ba-1101-480c-a20d-a745edab1d73 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types ORPO: Monolithic Preference Optimization without Reference Model

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.403390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:b7c62968e2966ab1a71971b10a7aeee892b6685b4b2f1bf3b24f37bf5692b9f5

Observation 95e49e35-dc8e-4cd3-b648-8af0b9c58686 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.500846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:047aab83391c8f62624fb55a47cb7635c5fdb8ebd22d8cfa9bf03aef0e293fda

Observation f32f8a9f-0d18-48f9-87c8-7c7e16123b04 · outbound

This paper cites Mistral 7B.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Mistral 7B

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T21:23:27.359297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:181b30a4caf25c90f57024e89ba8ed9b2d6d139e3a9d96eb394e4e6bd6cffe70

Observation 2ebd1d91-ecbf-45fb-bfd3-e5688bb9c815 · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.417555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:8af35486e41696af6285a62928ff3b68c21d8aaa56d5f097bdb50c87b46399d9

Observation 526fd740-085f-4992-a9f9-b3d9e4eebd52 · outbound

This paper cites Nash Learning from Human Feedback.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nash Learning from Human Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.384210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:708c31e7f8499bd12ca2210dd9406605f45a885ec4523d73e4e59e8fb660b803

Observation c3328669-db3d-4546-a3b9-058deafc9d92 · outbound

This paper cites Nemotron-4 340B Technical Report.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Nemotron-4 340B Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.437406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:7f4cfcac3b525a9b33566911546b135e15387771b5c170157f59f7bb094b3249

Observation 710e60f8-dc94-4cbb-bbe3-e01e1a82361b · outbound

This paper cites PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.465151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:deeb68330293b4fffab7f9a6354a14a89cc83468b6b7a1b015acf5c56e2c8d28

Observation 3b9ffa27-5a71-4bac-ad08-adbca58e3303 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.431660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:ae2987af7aeb3cc859cc132d73f57fca2a32f686e243b2963224e2054959d680

Observation 93c2f85c-31fb-4c8a-b53e-f10c6b3e8cd1 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.442405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:b460c08f3c3db0378d03627a4412d9fc8aaaf19d884125c604879019a4d19b2e

Observation fca72ce2-2f53-4be1-be10-2196cf86329d · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.507382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:d26cfd0d9435f1295af7aed1bea628af49983d92fc919d3ea507b2d1c215fb3f

Observation a6003053-d699-432c-98ff-7e1f2f765e5d · outbound

This paper cites an unresolved cited work.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-23T21:23:28.461031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:024f1e2cf7f7bd7eb3affdc8b49fc5c5b358c04f545f4197d72ccf38c132182e

Observation b0894764-c8cd-48c4-9789-8e388b060228 · outbound

This paper cites Preference Ranking Optimization for Human Alignment.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Preference Ranking Optimization for Human Alignment

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.476519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:b04170605864f655e50bf97195794a67ea2e432ae2f3c8068ad7674962b5f48f

Observation 391ef5c6-cc14-465a-8d28-91bc4e34b19e · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.495114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:2d178dfd8d73c18964af3d87257447f4922c1535b0aca86bb65a5a5f6c0b764e

Observation 3de95e35-1551-479b-9295-867dec474dbb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.452991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:6958298f5ff3de99936110fc4b4f5a6f2da4f028b1557a71d8f48fd6dab99959

Observation cbd8afec-9ea4-41ba-b087-9325b32ea148 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.390797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:9767ab71b04ffb520a29e50b27b39dfda8ba2f301cd9a7bb61751c3199085fac

Observation cac8fab4-4a54-4d55-a42e-1558f0187b6c · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Play Preference Optimization for Language Model Alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.410682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:315931206f9c2abaa2c68441ea8124db888d54f2f583d566abd8702d3005a3ad

Observation 3107b263-5fce-43a2-99a7-30228e15d883 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.459223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:acee1a2a47b93c878a382ab527f022e0bd985e9546c2f470ba56e20ece74d886

Observation d7d6c5fd-dc09-4622-953b-957c1f47e449 · outbound

This paper cites Self-Rewarding Language Models.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Rewarding Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T21:23:27.377474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:b7729d328d52fc539955859fcf13faa5256a6bceff6b8cee9ea5f6931d6b194d

Observation 8fa9f457-f4b5-46d5-b3bf-c65a6b2866e0 · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T21:23:27.482146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:c4abc31398f11cf53e8d92ecc6089ea4e93b27e62df3a83bcd83678168f8f669

Observation 078580d8-a1a5-4ca6-b1db-d8247483d942 · outbound

This paper cites Token-level Direct Preference Optimization.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Token-level Direct Preference Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.487999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:5c48405701b3733972a582ffe3582f6a7ed4e8596a28245b229b41bbb9332374

Observation 737e8445-6ce5-4896-920a-334a477100e6 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T21:23:27.447863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:18652e27acb29cae40cd2c9dc8fc7495604b4ba6901ea1203a0fbd4d1bb1dc81

Observation b0c7040b-6999-42a1-9f27-9b5f1973a6a8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.470716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:42ffcfe1e44b516ce485aad9cd765e5258224098bdf4993aff4cc1c6a25a6128

Pith citing papers

No inbound Pith citation observations are available.