Pith. sign in

Paper Citation Record · LEDGER

Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.18571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.18571 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:31:57.941537Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T07:59:50.977785Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c1c3830-a08e-4dee-aea8-068e398d2bfe · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.929415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.929415Z digest=sha256:4feb62b174fd6aea3014ee6ca5485f04e861e5b859aa091e88e0faed30a1136b

Observation da1bb913-e86b-4ec3-a51c-8e5b0c49a762 · inbound

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes cites this paper.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.118721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.118721Z digest=sha256:b759734ef0fbed5bfd79587ba722f48545263897449f9cec6ce4c40bf78f9bf1

Observation f269a302-d70b-4d9f-af17-0cb8f115d568 · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.638227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.638227Z digest=sha256:9e1b287ac4cfa86d5baf792948c4b69e4aea3220ed351e56963641c24f808cc4

Observation b5c40558-960e-416e-b402-753a034c8944 · inbound

Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants cites this paper.

Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T11:20:48.489366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:20:48.489366Z digest=sha256:0f5877fe7950f2311f596f934ef472a8ec0b15e0be67816054c03bfc5b6019ee

Observation 8c8b266c-63c4-4646-98d7-6e4e80215d26 · inbound

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment cites this paper.

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:31:57.941537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:31:57.941537Z digest=sha256:be914bfe81005fffe0993ac01f6669efaf0196e79568b57baf8423ae17b49697

Observation 63933835-d9dd-44a5-b42f-57788a1a3070 · inbound

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models cites this paper.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.923655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.923655Z digest=sha256:a5242cfb0ba3b503dc58dfe07457a6b403f2e2333d2ce56c30ba7b2c721245ad

Observation 2c8814b4-ac09-42a0-8df4-2e87ae403b07 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:47.967756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:47.967756Z digest=sha256:d7faba1c5ceafe44830f55a4beefc905d6304e92746fd76efb56a512108fef7f

Observation e7c8015e-3c13-4bf1-bfbc-7df13e0bfdea · inbound

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment cites this paper.

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:18.299705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:18.299705Z digest=sha256:809bf9785a889fdb4c2a9ba9742a1196ee280f84557f24359d4ec838aef3c019

Observation 455ce814-39ff-4b23-ac00-f1f8cb8a6233 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.859644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:2ddf12a31e5d102e4ba9e93253524391ecf2c65ae0528e2d48cc26c50b76456d

Observation 6f690dbf-7c73-41e2-8399-c55da0741e86 · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.114202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:28e2f9565b7781c0dc2d72e140884ad9bb9bacdfae46e001abf6d74ac2f073f5

Observation b3524346-517f-4318-8abe-cf6999948d15 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.808702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:30ed3de207eba7527e15532e5ec5b40fe21349c57c7e361a28eb849e27f96762

Observation e1eb872b-f9b3-4eca-b37c-7cdbd9eec21d · inbound

CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization cites this paper.

CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:50:54.299657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:16:28.593349Z digest=sha256:fc33dde75bb9c30642ca140d983f3c5a80383c8d3caa4c0a05c2d0786e0c064b

Observation f5a28b24-336f-4925-9cc1-7de381f28271 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.526224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:4926262b665e66ac34f1a78deef0cc1f1baa49b277e6fc8738eefb3719fe5a1a

Observation 7675384e-1375-4dae-9231-c096f722abec · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.887870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:02373f1fb6557a630b5750d1779309da63f2e1e5b5be035dc24f1277d9267164

Observation 1de03b3c-9dc8-4e29-bbd7-4c767030b6f2 · inbound

MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization cites this paper.

MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:13:05.313131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T06:09:56.684622Z digest=sha256:749f99a2807b2842c4f6dab1d3cd80a74fa11931ba58905ce03950d9be7eb996

Observation bb1526a0-e8e7-49a9-8bfd-b38fd5a1973e · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.979793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:8e203623ce65b0ce25a6684587675cbd7c2e045d8214f490d88b73fb1eb875ae

Observation dfaca053-4c67-49e6-b91a-07cf2c2eab44 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 148

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:d821628b084291c49aa1519160f8526bdfcb17c965fff5f9e2ad42ce5aa7b190

Observation 364a638b-056d-4e49-bf1e-b4ceb9c665a2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:48.709771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:48.709771Z digest=sha256:6ab674569945828531ff9c586a835262b0fea8a17967f6762d08beca24957e94

Observation f5a5977d-de80-4375-9be3-5169aed0a61b · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.520968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.520968Z digest=sha256:b38ef920825a3fd991fd27fa88dd58a6416d48630cf147ec4b4367f4defcdaaf