Pith. sign in

Paper Citation Record · LEDGER

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

As of 13 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2510.06096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.06096 v3

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:17:36.764467Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a4c187e-5c6c-4a98-9c82-14f8146df2af · outbound

This paper cites Uncertainty quantification in fine-tuned LLMs using LoRA ensembles.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Uncertainty quantification in fine-tuned LLMs using LoRA ensembles

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.682867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.682867Z digest=sha256:76f1fe66ec70a33c36b0abdf52954e5ec7ca23bb41ed310dc4bccf317f9da1b5

Observation b2174ce8-2c66-4898-a84a-1f6dcd1579e8 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.700516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.700516Z digest=sha256:db57b356a6c29f051566404a96f76b16613032330cb8c215952b66e0b164b3d1

Observation 8e0b2168-f3dc-4494-9489-80b906e6c9ba · outbound

This paper cites Toy Models of Superposition.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Toy Models of Superposition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.706487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.706487Z digest=sha256:b2f3970ea4f9b364979116c17121b7f4611a22f07bb313aa9349c0837565f186

Observation b9c2b000-f501-40e1-85e4-21a600ba3d39 · outbound

This paper cites Alignment of Language Agents.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Alignment of Language Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.718355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.718355Z digest=sha256:9d837b5935db194e5f2aac662a65ed576f866084523be437a2b480987fc1d0cf

Observation b8c570dd-d43a-4150-b0d1-544d1c28b749 · outbound

This paper cites Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.724299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.724299Z digest=sha256:c6e5358fc0a05f839e4ba95355517bb4703c49e950dac1337b82044f1fcd3225

Observation 7206c980-9ad0-4ebd-b63c-4191a50e06af · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.730234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.730234Z digest=sha256:d0f48eb994d52ff7268d0ebaf2a162d5b9e3a66a68a522938d9dcfab8a9e3e1f

Observation a8960377-10e3-4407-a63f-80af2ab5a6f5 · outbound

This paper cites Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.748639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.748639Z digest=sha256:c4572396c0da8c64915fd37cd1c77b001183ccba4fb33991471b25b259f66491

Observation 4440dea5-165d-4523-8693-b37bfd7be1bc · outbound

This paper cites Bayesian Reward Models for LLM Alignment.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian Reward Models for LLM Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.758976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.758976Z digest=sha256:91044fb34980b7f9b3285c8997fa1396c15212a61f9155f785562f5a12bf8d6c

Observation 6feb97e1-b97a-4e3f-a7d9-b00823355b0e · outbound

This paper cites toget back at fuckboys.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives toget back at fuckboys

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.764467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.764467Z digest=sha256:587ada2efea32e66d4d6d943671839992152d9334c2877f8842db1a1a5eaeea3

Observation 86eb37f9-d4fa-4adb-8efd-97e1592fb523 · outbound

This paper cites Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.735976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.735976Z digest=sha256:f3dbcd9ea630cbf5ab92fb37a2fe4c2d20bba21ee84988f3c9cdefe978fade12

Observation 69cb4738-ee36-4eff-815f-889f1f09edca · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.712466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.712466Z digest=sha256:ca7ed54f91a14df52c56a5049e201c2484f93e56c60db7e4fa3a410f8052543c

Observation 58ce233a-7689-4a88-977a-8ab27b2358d0 · outbound

This paper cites Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.742733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.742733Z digest=sha256:c1f9f3f8169934439cb4b0ef5e12db2a8466b823c4991acfcfcb4047463f395c

Observation 4e9a5758-72b3-4456-9766-bff79fea411a · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.688704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.688704Z digest=sha256:5bdeb24c0e3dcec2bdc7248d94827a5cb2d9723909ca85c6e52744cace0e1fff

Observation a7416fa0-5e01-4096-8f7b-9953471000e4 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.676213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.676213Z digest=sha256:732f510b563cae45621e04a9991fc8d00bd3690bb5f0d845b46b2de615bdfd5b

Observation cd63256c-b9d3-431d-8e54-c3cdc8957097 · outbound

This paper cites Ethical and social risks of harm from Language Models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Ethical and social risks of harm from Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.753954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.753954Z digest=sha256:54380cb0ecb0c07e8f64635fbbbb7743697771ca20509eb03484d52bd184916c

Observation ce3d0d09-b88d-4b46-a180-347335720062 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives On the Opportunities and Risks of Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.694949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.694949Z digest=sha256:d01a6ea019395f4f516d5a55753318f53ed481a2d7c36276856625ae686c1b22

Pith citing papers

No inbound Pith citation observations are available.