Pith. sign in

Paper Citation Record · LEDGER

ARGS: Alignment as Reward-Guided Search

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2402.01694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01694 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:28.390947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.378731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3258aed-6092-456c-8e05-fe4a1ca085de · inbound

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time cites this paper.

Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time ARGS: Alignment as Reward-Guided Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:20.211826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:20.211826Z digest=sha256:c5a79ef2033458e370f793f21b479804cbced2cbb24464fe535b32b3c528f215

Observation 9eabd63f-8d6d-4fa7-a945-faa52e58f4af · inbound

BiasFilter: An Inference-Time Debiasing Framework for Large Language Models cites this paper.

BiasFilter: An Inference-Time Debiasing Framework for Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:28.390947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:28.390947Z digest=sha256:29a78d95fe72b9f9b13df57412b1abdc3b7cf0a81b35746a81a6c81dfb472208

Observation 8ae2857e-1b3d-4a3a-bc91-16c3ace39994 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.597106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:d3488215de8152f77f4882e0e7d3be27bbd1126cd26d79001ac1490f6787af43

Observation 55bc5fd5-0f8d-40b5-8ddd-d4c87946db01 · inbound

LLM-ML Teaming: Integrated Symbolic Decoding and Gradient Search for Valid and Stable Generative Feature Transformation cites this paper.

LLM-ML Teaming: Integrated Symbolic Decoding and Gradient Search for Valid and Stable Generative Feature Transformation ARGS: Alignment as Reward-Guided Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:57.257974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:57.257974Z digest=sha256:208b514cd8de5ff0410156c2305d59814a81f394b87969863609f5ee3efceefa

Observation a3223f1b-673c-43ab-ae62-693610b0990b · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:55.771105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:55.771105Z digest=sha256:5570e6d5edda786b65279e89c76c1e78f4b28b43bce84f374ab1d009fed12df7

Observation 041852b0-1b03-4c76-b9c6-1bf402807f70 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling ARGS: Alignment as Reward-Guided Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.872167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:a32a44a24deb615237f4c5d15bf27cceaa8f82be688a7851be47ae443b368a0b

Observation e6a67e93-2552-4e32-97b0-f01c49869cd0 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary ARGS: Alignment as Reward-Guided Search

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:32.704726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:32.704726Z digest=sha256:e275242ad1128fd4da1f96839f2a26bf55915503df6d6f4de4804608d8af2464

Observation 50e73708-3b7d-4fc0-90bc-a99fa3c9ceb4 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities ARGS: Alignment as Reward-Guided Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.063586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.063586Z digest=sha256:384391c748cef6d86bb7ccaf401b5ec2ca5f001fb795621d74c35e385f917592

Observation 0b960119-1126-497f-8a76-8dfbd9da11f7 · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:42.488047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:42.488047Z digest=sha256:37b7109ae840d632b8b0e476edb3e8d4f43d345acdfddbbe6061518fbb55437c

Observation 58f9ff3c-6405-491a-bbab-d6af7d01fcbe · inbound

Virtual Agent Economies cites this paper.

Virtual Agent Economies ARGS: Alignment as Reward-Guided Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:09:37.443627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:09:37.443627Z digest=sha256:1416f425ba6eae0ac9e1389ff7acbf0ecc2b8114044b7850df790816c5a68112

Observation 94dfb535-c47f-463f-9aa0-cb9d3c44c5c4 · inbound

T-POP: Test-Time Personalization with Online Preference Feedback cites this paper.

T-POP: Test-Time Personalization with Online Preference Feedback ARGS: Alignment as Reward-Guided Search

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:11.128628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:11.128628Z digest=sha256:16ccadf6602cae2225638cc31e4e0218ef5816d7e92347e8d8985ec20bcc2b73

Observation b8888d26-c4bb-474f-8d6e-a3042504f5fd · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards ARGS: Alignment as Reward-Guided Search

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:44.167888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:44.167888Z digest=sha256:5bd1120da969c753d5c77f219bde3694ec71b7eb19e3f7f7d2fe162a87cfceed

Observation 80061e44-3d5b-4f76-975c-ae0241157149 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training ARGS: Alignment as Reward-Guided Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.535591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.535591Z digest=sha256:6039b46d0123cb89b44f32066f5824a7022790697a0f79a0f144fb5c9158c60d

Observation 0ce5fda3-b4e0-4e8f-b7b2-94d90910bdd8 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective ARGS: Alignment as Reward-Guided Search

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:23.362311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:23.362311Z digest=sha256:6dc50bd5f04c5f08ffcfd6ec4376f7fea2b0c8a22121cb6f8384c78fe8f685d4

Observation 63eda618-2b8d-4180-ae03-67eaa70e4fea · inbound

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning cites this paper.

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning ARGS: Alignment as Reward-Guided Search

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:05.309215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:06:43.463303Z digest=sha256:2ceb88c8828f740af942344cd756fc0cbca3373a659d6fb5b996207bd4a6a7ce

Observation 8b261bb4-3bc0-405d-8574-37ed6cabbf66 · inbound

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control cites this paper.

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control ARGS: Alignment as Reward-Guided Search

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:03.967708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T02:31:07.932802Z digest=sha256:de02fc79c5d03d1c3ffb3265d68426477aa4c24560859b5b97325b5a855dcf04

Observation 68bba25e-0882-401e-95c8-7161af716b3f · inbound

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing cites this paper.

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing ARGS: Alignment as Reward-Guided Search

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:19.384158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T06:28:14.378602Z digest=sha256:92464c12f647f1ae029c3e6947e8c150b65b993bcc9aa80d05bfdd13273f7645

Observation 0fdb043b-ea26-481c-b498-e96acf9505b8 · inbound

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement cites this paper.

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement ARGS: Alignment as Reward-Guided Search

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:23:48.490580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T22:19:33.582024Z digest=sha256:19c43c203816b35264b3b69016b5b368e448cebe8ff17726276024f730eed358

Observation 6e3afd58-137b-463b-ad8d-35702c7b44c6 · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment ARGS: Alignment as Reward-Guided Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.948373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:2a403962a5fd6aca2b7a9241435549d4e8dd0fbb0b4874e04b9e11e2157595b6

Observation 7b6cdbd5-0343-4c1f-ae14-6ef2bd6f0982 · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.008018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T08:10:55.720464Z digest=sha256:fbe4a6d7d644e53cd49d4587bdf28ba537f94cd9b92a373c5d54e2896ec1d83b

Observation 3bbb5d27-7c5d-49d2-b8c2-2b40751759b5 · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning ARGS: Alignment as Reward-Guided Search

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:24.337367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:47:33.925413Z digest=sha256:57a9142ba61cc9fca87a7a23f24f7e0feb634ca1efbb9b27705ba1057b48cfde

Observation 2daf7f82-38aa-4714-815e-050469ebd7b7 · inbound

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control cites this paper.

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control ARGS: Alignment as Reward-Guided Search

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:56:44.380450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T07:18:51.066384Z digest=sha256:081fdd2ba3dcbd861757f40c48ee2e66b2217e4c8ac4460e254f53579625fa2c

Observation 2aec05c3-0fba-4c6c-9285-d088541337c5 · inbound

Safe Inference-Time Alignment via Lagrangian Reward Augmentation cites this paper.

Safe Inference-Time Alignment via Lagrangian Reward Augmentation ARGS: Alignment as Reward-Guided Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T07:05:47.150308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:05:47.150308Z digest=sha256:60f589cbc739777fb4b534ef058a281ce125f10ddb2e47a33adee57c6963f06f

Observation 71560dde-6950-480f-8a10-0553bd785f97 · inbound

IFHierBench: Hierarchical Instruction Following for Large Language Models cites this paper.

IFHierBench: Hierarchical Instruction Following for Large Language Models ARGS: Alignment as Reward-Guided Search

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T23:14:22.758776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:14:22.758776Z digest=sha256:0332b04ff1031aac6b0d59686f86bba7a8aefbf1c442447587893c68e653de74