Pith. sign in

Paper Citation Record · LEDGER

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2412.13998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13998 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:38:43.203572Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.399277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:34:26.163405Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43689ceb-c541-4db3-b210-86441f99eeb6 · outbound

This paper cites write newline.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.374266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.374266Z digest=sha256:97b6af8e675de27783fe843bab64704a529bfabefaddb0a2f0088f0d20f4f574

Observation 250908ab-3120-4f44-8c85-56326a1db72a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.420860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.420860Z digest=sha256:47e3c5c1a443ed3a3d47801528f72eb72a875ad87485f35daf0cd05abc52a238

Observation ad1e4c42-4b68-4d93-ade1-8f6705b801c9 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.432089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.432089Z digest=sha256:61c3ea4b54c98db6f5de87c0555132e751b23e04667b42400d20aa4f63db4638

Observation a26b12ce-55ee-4cc0-ac18-8e93e5019c61 · outbound

This paper cites Pareto-Optimal Learning from Preferences with Hidden Context.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Pareto-Optimal Learning from Preferences with Hidden Context

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:38:43.254158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T12:38:42.442531Z digest=sha256:0965a86915ee288c1a5e7e07c4ef64ef60f2282370ebcdb06af1d8b018990aa4

Observation c9dcdbbf-3471-4d72-9556-16611bf5b0a1 · outbound

This paper cites an unresolved cited work.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.449622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.449622Z digest=sha256:3705abcc2e8991a31898641fe2c42b0ae4420d252ca785c2ddb358c8f2b8e845

Observation c046b64a-76df-4ac3-becd-caba3cc8f1f9 · outbound

This paper cites V., and Syrgkanis, V.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes V., and Syrgkanis, V

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.455651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.455651Z digest=sha256:3afa88cfd4f0c6449c610126b1a9947ed8f3e0e0a2e9b0fe93d2a7334eb0c3e0

Observation 27a85868-da9c-4cbf-b000-666680618db7 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.460776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.460776Z digest=sha256:56ad073a77e6f605323aeecfe19cb93f5c07e6e636e2bb17ea1e38ee4329e15c

Observation 9966b605-cc5a-40fd-a080-678a3f0bdd6d · outbound

This paper cites W., Rezende, D., and Eslami, S.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes W., Rezende, D., and Eslami, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:38:44.145821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T12:38:42.466288Z digest=sha256:70d74dbcc71734a6ac2107c0b0edde798d78eaa4b55a5036de38ac5bf8c2f155

Observation deb0e917-eefd-4d4f-bc5a-1d50bafec7a7 · outbound

This paper cites Neural Processes.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Neural Processes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.542225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.542225Z digest=sha256:ca3dc2791ff03b0bab12ff656515d8856473f8c19d80198a6ade9cd67d5915cf

Observation 528954e4-c12f-4004-b160-bd3384c4d079 · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.606208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.606208Z digest=sha256:b7a28f66369595996bec83c1a8ef0f4a53d98a7a8c29611e47a9b7c1bfd830a4

Observation ba85029f-7f41-4750-bb58-e62d39f7e4e9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.687841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.687841Z digest=sha256:78c199745fe23c258a566167bea0e7a85f78789ae80c76dc669c6c5537566d28

Observation bdb0deea-a249-478e-8cc3-65d4c56f5484 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.693809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.693809Z digest=sha256:d70dd6cef7d78050537a94f6b36ce61e37417ecdd48a42133c72c9b78088420b

Observation 9259effb-283c-4e17-a72e-1a043393b3f2 · outbound

This paper cites Preference Transformer : Modeling Human Preferences using Transformers for RL.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Preference Transformer : Modeling Human Preferences using Transformers for RL

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:38:43.954124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T12:38:42.699493Z digest=sha256:afb9b4685f3fecb6319b636245fac2d9a3dee117c05dd6dbb1d74f4cd55807a3

Observation 649ec6cd-7114-4f2c-893e-5699022ca650 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Adam: A Method for Stochastic Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.704373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.704373Z digest=sha256:0304eb53e61ec573b722a583938de6b31511db66f8ebd573c6f436b4ddd466d0

Observation f0e15f9e-bd09-49f0-92b1-4f4d3adfcc70 · outbound

This paper cites Training language models to follow instructions with human feedback.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.709208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.709208Z digest=sha256:5330c32f8444325da21509c372f25b7705a937b620be8b83a3b99d282529d71d

Observation d39d24ee-4f7c-4713-96aa-8e4d92794d4b · outbound

This paper cites Automatic differentiation in pytorch.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Automatic differentiation in pytorch

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.715496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.715496Z digest=sha256:23c5eb7dce5d6f57fa8a9186a86e5b8c3dc8f368a8e7cb927bd33072e3c9b2f8

Observation 640a6259-da4a-4a4d-ad6f-e33fabf67bcf · outbound

This paper cites FiLM: Visual Reasoning with a General Conditioning Layer.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes FiLM: Visual Reasoning with a General Conditioning Layer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.720440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.720440Z digest=sha256:d160538d0a37769fc4f277088bd52d0e8f111d581b69fc082debf2f8adb3458c

Observation 025fb091-50e1-4470-a2b2-e333053699b1 · outbound

This paper cites Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.725901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.725901Z digest=sha256:562f964e08da7cf312cbf1a529c246088c1348538296b2127143bc1fbb1df2d9

Observation 3aca3f66-cddd-4ac6-95f2-b1d7bc3ba8f0 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.732968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.732968Z digest=sha256:ceb873caaa3b360c20f6c72d92fa92a801b5fb6c932d73aa196907f49f1bdaff

Observation 381a1218-56d0-47fc-bfa7-658bb07a7b81 · outbound

This paper cites Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.778618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.778618Z digest=sha256:9e553b33340e53e1d9a1b7f7b44672fb34d5f45ed13858fb3a1bfb39ad7d3f87

Observation 162becfb-7cf4-4aae-9f36-4588d7d45ddf · outbound

This paper cites DISTRIBUTIONAL PREFERENCE LEARNING : UNDERSTANDING AND ACCOUNTING FOR HIDDEN CONTEXT IN RLHF.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes DISTRIBUTIONAL PREFERENCE LEARNING : UNDERSTANDING AND ACCOUNTING FOR HIDDEN CONTEXT IN RLHF

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:38:43.874000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T12:38:42.865823Z digest=sha256:c84d60f1ccef255a02457ceed2174b35011877535ad0c3a44636087ca06ed772

Observation 07f9abf7-96ec-4007-a446-4ed10da6cfdc · outbound

This paper cites A Roadmap to Pluralistic Alignment.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes A Roadmap to Pluralistic Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.918505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.918505Z digest=sha256:8e1f0a7efdcc99d9c58233af37b7a2f6523148e5668239d33ead0d4d368fdc1e

Observation e27b4d01-c81a-416c-b870-d071856d62b4 · outbound

This paper cites M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:38:43.854661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T12:38:43.031870Z digest=sha256:f95c6748c219782a8aff5aff83c40f36a2493dfdff95efbdbf2994ad10bf5e6d

Observation da1bb913-e86b-4ec3-a51c-8e5b0c49a762 · outbound

This paper cites Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.118721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.118721Z digest=sha256:cb3213899d629715880830cacc6151351d40459e2e966e7382d41f0138ef8246

Observation 7225a9a4-7606-4b12-81ed-c638cc3613ac · outbound

This paper cites Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.187939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.187939Z digest=sha256:cb8dfcc518a61c3245263f9211cd33395abcd3f6f0ff691d006765cf28e7233e

Observation 286c9f84-ce67-4444-8e5b-41ba393b1719 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.193545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.193545Z digest=sha256:b9849f6522db5be2ba3ca29e91f8e00477076b70cb9598371d4bddf39b8c5ef7

Observation a7dc0ca9-b69e-47a1-b330-91bd90b61821 · outbound

This paper cites Deep Sets.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Deep Sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.198497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.198497Z digest=sha256:46ddb4b15f851dd0596f9f89514586426785397fbabcd90fe6d948f112dd21f3

Observation bd68763f-dc41-4f9f-aaed-8695034a5913 · outbound

This paper cites Panacea: Pareto Alignment via Preference Adaptation for LLMs.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Panacea: Pareto Alignment via Preference Adaptation for LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:43.203572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:43.203572Z digest=sha256:c4cb9ae98d9b350077b13ee14be6afec0769aeb71c6aaf26d80247f4a91ac0a3

Pith citing papers

Observation 8acdb000-0686-4932-9d12-8726a2b6d43c · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.399277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.399277Z digest=sha256:98d87d431e1a970b556832c233edaa5209466355cbd21d928a0580eb09f4ad81

Observation 4e12ffca-9b85-4c25-ba3e-61cd0bbcb210 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:34:26.167583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:34:25.069706Z digest=sha256:25b21f15e89d2e1c18d1fae1355a4f8770781137c6512bef92241f25ad7e07b8