Pith. sign in

Paper Citation Record · LEDGER

SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.05344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05344 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:25:11.024505Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:43:25.917797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8909a967-efb2-4504-a40f-28bcc70c8c3d · inbound

ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation cites this paper.

ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:25:11.024505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:25:11.024505Z digest=sha256:ef3d67fa1002fca2d742d13a2387d123b9ba736c88c4379f44f14c63a5c848da

Observation f470421d-ab73-4e1c-b06d-bb39dce0e019 · inbound

Data-adaptive Safety Rules for Training Reward Models cites this paper.

Data-adaptive Safety Rules for Training Reward Models SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.113800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.113800Z digest=sha256:b90d57ef1920fb3f4e056dfe48625378a0fb959bc89f20fa8d67d4722543af36

Observation 6eac74c1-acf7-456a-abaf-aaef16daf23b · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.691866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.691866Z digest=sha256:2e60f8154e7d06fa6b18a0fb1774b934a012adb7beddf04c9132b662e121c222

Observation 9026ddb6-4b52-45c5-b7a9-6d578d66c114 · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.231708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.231708Z digest=sha256:5635fa466a807a9e3c7bb0f376e2df10f1087b69ddec1daa83d28a9ac7e2cd76

Observation 1bc69fa1-00f7-492d-a817-92d8f5655d7e · inbound

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment cites this paper.

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:16.759470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:10:16.759470Z digest=sha256:b38e7d80017ecbccd107cebed5577c493d05da24edab2c88cc25c36ff2be4cce

Observation 95ddb21d-32d7-4cd5-84cf-81fee0122848 · inbound

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost cites this paper.

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.919250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T12:36:18.979940Z digest=sha256:11f4210bea00d08441f5f94b85cb811cf084c984242c3b5274261da193446e78

Observation 3bb022fa-2cd9-4970-96fc-1dc62aa41205 · inbound

Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences cites this paper.

Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.645528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T23:21:35.338959Z digest=sha256:c26a76583063964e121f81d7f716f76e8c9f83b11098db8f65e96c3ab6442745

Observation 75b356a4-3b6e-4008-a853-ca5d470408ca · inbound

Step-Level Preference Learning for Generative Agents in Social Simulations cites this paper.

Step-Level Preference Learning for Generative Agents in Social Simulations SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:25.187877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:25.187877Z digest=sha256:7e85b21a96ade6d3a5d69dbe4d797fd304716a47e6314de4db250294c2d96c73