Pith. sign in

Paper Citation Record · LEDGER

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

As of 7 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2510.06096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.06096 v3

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:17:36.764467Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a4c187e-5c6c-4a98-9c82-14f8146df2af · outbound

This paper cites Uncertainty quantification in fine-tuned LLMs using LoRA ensembles.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Uncertainty quantification in fine-tuned LLMs using LoRA ensembles

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.682867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.682867Z digest=sha256:52d231eb673bf946a5dfe049ca024970190da24c7701ceffb6e9e31362194b88

Observation b2174ce8-2c66-4898-a84a-1f6dcd1579e8 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.700516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.700516Z digest=sha256:615a81148496cd98ef92b5e08f5d01963e9d6f693ddb49b4b33de6624eda1f53

Observation 8e0b2168-f3dc-4494-9489-80b906e6c9ba · outbound

This paper cites Toy Models of Superposition.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Toy Models of Superposition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.706487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.706487Z digest=sha256:8d53ebf5b524b638e6432d565194457145a9477687a3e3450f98813a3a5dac88

Observation b9c2b000-f501-40e1-85e4-21a600ba3d39 · outbound

This paper cites Alignment of Language Agents.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Alignment of Language Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.718355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.718355Z digest=sha256:3519ce064a387bfe0f852b0e2d1dc5bef515746d9c780994673a30f33c333546

Observation b8c570dd-d43a-4150-b0d1-544d1c28b749 · outbound

This paper cites Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.724299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.724299Z digest=sha256:a9145dd7087f83a77ad7a37fae35c29c750d0b3c046c6fb6aa88912dcb572004

Observation 7206c980-9ad0-4ebd-b63c-4191a50e06af · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.730234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.730234Z digest=sha256:5e8ab86ad7bf5b5dd1a4a0828092045e533791a7c9a22492dd1211f654d9c943

Observation a8960377-10e3-4407-a63f-80af2ab5a6f5 · outbound

This paper cites Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.748639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.748639Z digest=sha256:797417a3664d59b0eb5289dcf7e783d823dd1f2a863fece17cf461b8995013b7

Observation 4440dea5-165d-4523-8693-b37bfd7be1bc · outbound

This paper cites Bayesian Reward Models for LLM Alignment.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian Reward Models for LLM Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.758976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.758976Z digest=sha256:7927a25544f8a14809edc13eab0a19217cb6ba0bfb47f3d70811cba40e42b998

Observation 6feb97e1-b97a-4e3f-a7d9-b00823355b0e · outbound

This paper cites toget back at fuckboys.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives toget back at fuckboys

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.764467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.764467Z digest=sha256:55a7d2f5891bde5112d7a22f36d5400548cf10977e675dff8d7c7d2535a12c68

Observation 86eb37f9-d4fa-4adb-8efd-97e1592fb523 · outbound

This paper cites Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.735976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.735976Z digest=sha256:f7ccf5849e42e311183936aaccfb9d7b2676d9bc9977581cb5efea9271a25e91

Observation 69cb4738-ee36-4eff-815f-889f1f09edca · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.712466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.712466Z digest=sha256:2d566b7b1d4f4ad17fc43fc43e8d6a4224f7e2da7ee3250da621140f1e4901a0

Observation 58ce233a-7689-4a88-977a-8ab27b2358d0 · outbound

This paper cites Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.742733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.742733Z digest=sha256:7a83625011656cb86f8b80db0136aedf297f0e18fa4780b36bbf536b5c3b5932

Observation 4e9a5758-72b3-4456-9766-bff79fea411a · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.688704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.688704Z digest=sha256:8b4a970a05db19b1fbc9df50855938bae9feb220b73b1d2567149aa144efb98d

Observation a7416fa0-5e01-4096-8f7b-9953471000e4 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.676213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.676213Z digest=sha256:0202eca110236b957c937ab2b4bffc4a48c821e222ed604e53f74cb4ef875eeb

Observation cd63256c-b9d3-431d-8e54-c3cdc8957097 · outbound

This paper cites Ethical and social risks of harm from Language Models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Ethical and social risks of harm from Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.753954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.753954Z digest=sha256:6774d848f1248da365189ec916bec80c8b57bc080177b578d3342c6891f4c4c8

Observation ce3d0d09-b88d-4b46-a180-347335720062 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives On the Opportunities and Risks of Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.694949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.694949Z digest=sha256:a4cf511f3928d1bd32fb616eea6b5d03699387c69ee91ad3896f1c1065235c3b

Pith citing papers

No inbound Pith citation observations are available.