Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:18:17.886907Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2412.11145.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:18:17.886907Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ee5ceee-70c8-4a54-9594-ff6f9165c7e9 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Large Language Model Alignment: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4a4bff-479e-4ce9-a2de-474314a45475 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Statistical Rejection Sampling Improves Preference Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1c205b-8504-432f-98d1-6245523cdc74 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models KTO: Model Alignment as Prospect Theoretic Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32a8254-b368-4624-95fa-9520e269f644 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b50d46-0267-4805-a866-16ef8e1779be · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Supervising strong learners by amplifying weak experts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28e9c60-0409-4a0e-ade6-96397527604a · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0ea86e-646b-42c1-8dd0-c87bb1a3d15c · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd39d37-f285-4817-9a59-e0666b752e76 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models AI safety via debate
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ef8334-a65e-445d-b030-20965e2b20c8 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Training Verifiers to Solve Math Word Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66629ff-e189-4614-bdf9-3f3a95ab5386 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models LLM Critics Help Catch LLM Bugs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c67b252-b1f5-4a07-bc7a-bbc985a2c63a · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2708c8af-192b-4a68-93db-24b239e5a1e5 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models AutoDetect: Towards a unified framework for automated weakness detection in large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ecd67083-f9ee-4f72-baab-447e193754a4 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Language Models Learn to Mislead Humans via RLHF
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1047eca7-37bc-493c-b977-30c9fb9dcd0b · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db461c9-03b7-44bb-9755-b95547d1b5d9 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Unveiling the implicit toxicity in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2d76fa22-507b-429e-a715-bcfd64e27215 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Improving Reward Models with Synthetic Critiques
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9196234-2fa9-4c5c-9a18-9db170704218 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Reinforcement Learning for Generative AI: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b3b56d-02ac-4c2f-a0a9-7bef81be63a4 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Learning to refine with fine-grained natural language feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 81eac04b-19eb-454d-b5f6-73eb1e5e90f8 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Improving Model Factuality with Fine-grained Critique-based Evaluator
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db97a81-6283-4170-99ba-26eda00bbc34 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Bayesian calibration of win rate estimation with LLM evaluators
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 717fbb25-ced1-4ed4-b551-15a149fe3a90 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbecd98-f46b-4f7b-92fb-2dc7d38a88b7 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f631249e-2027-428f-a169-0f54032ef4b2 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Self-critiquing models for assisting human evaluators
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28700c0-7ecd-4477-8c5e-6f7da6145206 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Lm vs lm: Detecting factual errors via cross examination
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9b7f6cc6-638b-4e6f-8427-d88ade9038f5 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Panacea: Pareto Alignment via Preference Adaptation for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c2a3476-5879-499e-9109-0c8b9ff6b1a9 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Evaluating Large Language Models Trained on Code
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c7ea3fc-7624-491e-8f72-99b0ad5f281e · outbound
The Superalignment of Superhuman Intelligence with Large Language Models A Survey of Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ced059-cdeb-46d0-bcca-3f7ffa1b1e4b · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Proximal Policy Optimization Algorithms
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9719a814-22c5-47cf-b1ad-ab859268a7d5 · outbound
The Superalignment of Superhuman Intelligence with Large Language Models Model evaluation for extreme risks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.