Pith. sign in

Paper Citation Record · LEDGER

Towards Reliable, Uncertainty-Aware Alignment

As of 16 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2507.15906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15906 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:18.315140Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d809cf6f-ee7c-4ffd-892d-e7588102b78c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Towards Reliable, Uncertainty-Aware Alignment The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:21.156461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.176395Z digest=sha256:aefd9569a28d08f7679700c9000bb2c61868299a814f468905260f593d9f42cc

Observation 47697553-f704-4bb3-82e3-bfc825ef34c7 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Reliable, Uncertainty-Aware Alignment The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.266168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.266168Z digest=sha256:40d753a0bc8f84ed110da0346cd50a5109ebf6836b9ca357f6ebe7c71d1ec1b8

Observation ee77b394-d87e-49f6-8280-3c14d712c9cb · outbound

This paper cites Concrete Problems in AI Safety.

Towards Reliable, Uncertainty-Aware Alignment Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.333386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.333386Z digest=sha256:51e4255b2ce3e68a484a4e2c94915fe0f443c00001525aac95e92556817e38c1

Observation fbb27280-7859-4c32-8d4f-a42d30e3c2f0 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Towards Reliable, Uncertainty-Aware Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.417979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.417979Z digest=sha256:73d3a908179bff8c6cde5f6076a3367ece63600c9affb2d0b913412bc2f1f292

Observation 48ae5dfa-eb9f-465e-aa15-b6cc4542ab59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.494528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.494528Z digest=sha256:65da800c0a510965b9419adb05722c727e1abea0ba23be3f58312951e0a1c058

Observation 1728c136-ef99-4e4d-a6ff-0be5e299c464 · outbound

This paper cites Deep reinforcement learning from human preferences.

Towards Reliable, Uncertainty-Aware Alignment Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.621778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.621778Z digest=sha256:1f47b8010b38d00d02af190dd54ba25f4a815734d533db9e12128d691d5dc425

Observation d81661bf-1992-4d60-91f9-a40889259d1e · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.721341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.721341Z digest=sha256:72111b57277d481dffdd7a05d8788e7d5b15b4f1558f01a7134e7bec9d06b6cf

Observation 08da191c-c502-4312-bed6-4226835d001c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.818475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.818475Z digest=sha256:0ab7ebee80bd308b997b984c78164bd14bdd113f11b5ae7fbbd2fbb7e56eeb9a

Observation 1d279d30-0dcf-4f19-bfaa-89c952a0a89f · outbound

This paper cites Suphavadeeprasit.

Towards Reliable, Uncertainty-Aware Alignment Suphavadeeprasit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.952919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.865665Z digest=sha256:9ce633bcacfb65415aec865272d22a9bd622bb60e72ef5032d8d2452a7419e18

Observation 5e10c0f0-fb17-4bfe-9a8f-44419e76630c · outbound

This paper cites Gemini 2.0 flash: Next-generation multimodal ai model.

Towards Reliable, Uncertainty-Aware Alignment Gemini 2.0 flash: Next-generation multimodal ai model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.783795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.944596Z digest=sha256:68cf4e2e895495e45b0bb34a86963cbb3ccaf31843e39a10ac3b5c7a08d6000a

Observation 1361e871-f6d0-4c4d-9c6c-b5e0025049e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment DeepSeek-V3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.010645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.010645Z digest=sha256:0609b760b1c1f46f7dbae77a154c0f2d4e2ed53ee8c2cfc4c089408d79c1737e

Observation 97e0d881-e718-487f-a5bb-1e108dcacce4 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

Towards Reliable, Uncertainty-Aware Alignment RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.132426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.132426Z digest=sha256:b61b82bb40b44da643c6983eb0d7a6294e5d953ac30fff150bfb7325e3ba6205

Observation e91f2dfc-d7f7-4e9b-8792-7f4fef976c8f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Towards Reliable, Uncertainty-Aware Alignment RLHF Workflow: From Reward Modeling to Online RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.208038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.208038Z digest=sha256:ed7536d8618f77881e841ded49d636b301e241f5d6348f466f76e4a6d12e8712

Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.278132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.278132Z digest=sha256:b57c16a51a9cbef8f4767f999a0c03285bed746bca853fe563ad432d2b550626

Observation 0e5c7de4-926c-4d9f-8825-509c26dcdb16 · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Towards Reliable, Uncertainty-Aware Alignment Understanding dataset difficulty with v-usable information

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.612471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.350372Z digest=sha256:cf10de0dbb6a69695ba4c029daf87434940f8f2f6d285d463153c7d45a255946

Observation cd467f85-11b8-4976-a8fe-6efab80bbead · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Towards Reliable, Uncertainty-Aware Alignment Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.401191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.467970Z digest=sha256:669b2ea66557bd65a542301dbe38d14cd94f964638bf21589aadfba6c8fee6eb

Observation 4115c2db-0a6e-4527-b18e-31c2da76f903 · outbound

This paper cites Scaling laws for reward model overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.648398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.648398Z digest=sha256:70e8b9568a786b2f3fdd0d74e9ef9be4c5612dfe918c7bd9bc63c93818bc3b5d

Observation 8f8c4ae2-e1e8-4d32-b9a1-69738b2f53c3 · outbound

This paper cites https://huggingface.co/google/gemma-2b, 2024.

Towards Reliable, Uncertainty-Aware Alignment https://huggingface.co/google/gemma-2b, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.250009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.807037Z digest=sha256:124fbbb72a947569819d541d46be235a7e709dc1b8fc66a9cbc6daf2fd08696d

Observation 1dc032e9-a02f-4dce-b96d-6b068474f4da · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Towards Reliable, Uncertainty-Aware Alignment Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.973348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.973348Z digest=sha256:fcd6041907b076a09a9db62c7aa9c356957dbe2fdb1bc8564a5d9b07c0d282d2

Observation 0130163d-ee01-4d80-8e01-84e6c0938b5b · outbound

This paper cites Mistral 7B.

Towards Reliable, Uncertainty-Aware Alignment Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.092213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.092213Z digest=sha256:e051a3d2adf35713e50a6c32a8d4ce363fa818fd93b0b6f7c2a307242afd1402

Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · outbound

This paper cites Reward Design with Language Models.

Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.257918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.257918Z digest=sha256:25fefd7502feefa56f4c43d75208c8c9869c0a166493aef1678976835ae6da82

Observation 61587e0a-0251-4e40-b1ba-8cc8fee53523 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Towards Reliable, Uncertainty-Aware Alignment RewardBench: Evaluating Reward Models for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.368188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.368188Z digest=sha256:83df770b6cbe7d3241e95b7b14dff1ba676547f40f94ed5f73c9e5342d66637d

Observation 24b201b5-75dd-4e28-9733-9bee5051bfbc · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Towards Reliable, Uncertainty-Aware Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.487856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.487856Z digest=sha256:797709c5253320f71fc89cd5a39466f063a6af67043c7f40268af3deb9a7581f

Observation 6f0b80df-145c-4ef6-af21-d3a6aedfcba2 · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Towards Reliable, Uncertainty-Aware Alignment Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.024665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.623570Z digest=sha256:3082d70e54036175a2ef98f51f80cf9b2bc8a554dcc17363c42cee47b153e3c9

Observation 1544c899-73e4-4a7b-9a41-91ea72ae8d17 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

Towards Reliable, Uncertainty-Aware Alignment Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.689847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.689847Z digest=sha256:cdec1d9e5cc7e8eba1c8ea5a689f2385b218ec3dc6bdb845fe23c456d04c965b

Observation 0babacde-4ada-4f0a-adcc-94e46862e9e8 · outbound

This paper cites Iterative prompting for estimating epistemic uncertainty.

Towards Reliable, Uncertainty-Aware Alignment Iterative prompting for estimating epistemic uncertainty

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.849773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.762639Z digest=sha256:1c66a3a4729d02451206cb6c1bdd2ff70e83b60903b0629ab725625255d648e2

Observation 902799f6-fb92-4983-98a6-59f7a3fff674 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.880022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.880022Z digest=sha256:c5be207a8bb10a314fcc56d7cb820cc58b3b60d0f06b8f1235a455bbb7ecb4e3

Observation 28c4d9ce-4dcc-4d52-85c6-fc6d793041e9 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Towards Reliable, Uncertainty-Aware Alignment Introducing meta llama 3: The most capable openly available llm to date

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.957662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.957662Z digest=sha256:2e0262405842a270791b08cca11d9522d2fb92c589bf5adb4e4593d9675a137f

Observation e439b574-48dd-419d-a3bb-64f9c38ba7af · outbound

This paper cites Montgomery and George C.

Towards Reliable, Uncertainty-Aware Alignment Montgomery and George C

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.638664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.020519Z digest=sha256:f50a125c8be1d91fef087508e11fb7018f95416c49b7751a70732c9386d88db3

Observation 757574e9-7c31-43a4-abd1-614f6e78f258 · outbound

This paper cites GPT-4 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.098615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.098615Z digest=sha256:8ce4997d287c1e0c279b428eb09f353b3f0f963fba1cf67faaedc7665e1a5f9c

Observation 91246783-23e2-48a4-8de0-3e8e6fa149ca · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.193548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.193548Z digest=sha256:b460ceece6c01683558e02214b360a2d23d6d61e9f5df3e53c3bda847bea3da7

Observation 18c522b9-122d-4a4a-adbd-92b17f02a131 · outbound

This paper cites Language models are unsupervised multitask learners.

Towards Reliable, Uncertainty-Aware Alignment Language models are unsupervised multitask learners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.285137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.285137Z digest=sha256:2717b8ea607bc1d87866dca5ffb78d2aa78c78e2ddd0a02c9c5371132019f276

Observation 57c3ab16-89df-48eb-a27d-8adab523beee · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Reliable, Uncertainty-Aware Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.368788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.368788Z digest=sha256:80abd0fed4a94aef5ee3438b4bc02d36cde30d259e8b9d21cf539635aba3804a

Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.470462Z digest=sha256:cf435ff7abb4f2caba9ff7f374311dd882cf54558c0779f61f1a532e458292f0

Observation 9b210cb0-f49a-48d1-aea4-6c6df8c36850 · outbound

This paper cites Trust Region Policy Optimization.

Towards Reliable, Uncertainty-Aware Alignment Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.538614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.538614Z digest=sha256:aa2bca633c3de0dd5166943ce116997d1e718bc39efa2778fea8aa7ed70d95e4

Observation 58d43223-87e6-472e-b055-18d7a48a7cdb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Reliable, Uncertainty-Aware Alignment Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.600627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.600627Z digest=sha256:dee8ba1e51e544ee25ed729412c7ce1d8c8b88437f99e0a99868c6aebb71bcea

Observation 98871c98-ad6d-461e-90d4-40083a555d8e · outbound

This paper cites Mutual fund performance.

Towards Reliable, Uncertainty-Aware Alignment Mutual fund performance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.401125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.680687Z digest=sha256:602e129cf01abf0a898802edb54ea96f68cb394a7fd95ff11d790be4820d7e32

Observation a399d956-22d0-4734-b28b-5c7ad4a851cb · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.751394Z digest=sha256:0f1b5994cd967b276bd36f5b5e391d49f3381bd3a1daea245a96fb7ea2a0f329

Observation 9aeebd62-11a7-44fa-94b3-9d09552a9746 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Towards Reliable, Uncertainty-Aware Alignment LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.846785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.846785Z digest=sha256:4c97b14a744df6cfe78a1c08dade94b3a89ebd35c8aecc9d1563dc3e2bb6125b

Observation cc34c0a9-eb6a-439b-851c-fc50b32b7883 · outbound

This paper cites Learning to summarize with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Learning to summarize with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.908542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.908542Z digest=sha256:9ec8e97bad2d0098ae5879b44566d26eafa32bfd0934b6f8c4ec9e08218bf52c

Observation 362f149e-6940-42e8-8316-2a92bb0a8b50 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Towards Reliable, Uncertainty-Aware Alignment Policy gradient methods for reinforcement learning with function approximation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.999910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.999910Z digest=sha256:761d8c874b299b0530902a49f8c7005d3b6ed3f4b359ed8ba8d8805c66958e80

Observation b75730e9-da5d-469a-ad94-d67b210d359e · outbound

This paper cites Quantifying Uncertainty in Natural Language Explanations of Large Language Models.

Towards Reliable, Uncertainty-Aware Alignment Quantifying Uncertainty in Natural Language Explanations of Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:40:18.689752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.037690Z digest=sha256:9cf985f838848da0522f9b94bb97a1b4fcb440b9a189d1c891dc6ab5b7f56edc

Observation 99d1eca3-0d57-44c2-89f2-f0e50ca4e355 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Reliable, Uncertainty-Aware Alignment Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.125727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.125727Z digest=sha256:ba5c13552d0e3dd13a6f2df1cb7a21da8f06928e55be86cd62947bcd9141f29f

Observation 6b251672-a983-4cec-8bd6-84eb6017d14b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Towards Reliable, Uncertainty-Aware Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.194495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.194495Z digest=sha256:f0544dfdcc75a115c09b1e706ea32a7cc9116fce11679bbb695d76d729c24a06

Observation 8d6d2705-1af5-4aef-a24c-771fdf608aa3 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Towards Reliable, Uncertainty-Aware Alignment Gemma: Open Models Based on Gemini Research and Technology

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.315138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.315138Z digest=sha256:928c83e171650f9221226e3b30a63e81fafe708f62d787a58c4e5f26cb71a86a

Observation 4b7ec90c-cf1c-40a0-b378-fd8e214e9552 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Towards Reliable, Uncertainty-Aware Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.350383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.350383Z digest=sha256:88b4c82307c9de53525b3b9b500b339165b6906157fac6948f73bfe92279a3ec

Observation 98e4ed2e-64d5-47b2-b840-901b19b509cf · outbound

This paper cites Trl: Transformer reinforcement learning.

Towards Reliable, Uncertainty-Aware Alignment Trl: Transformer reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.443886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.443886Z digest=sha256:aaaba84bcdb5caf501fcd1fc929fdf772383b8998ebcf50dc840149cdd7f30bb

Observation c8a35bd8-9eeb-4310-b713-fbe3f304dbc7 · outbound

This paper cites HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM.

Towards Reliable, Uncertainty-Aware Alignment HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.507422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.507422Z digest=sha256:b7b8bfe75d61314bb067397df73b09f68bd1a2ead8a0a3df7f1bd51bc68f730f

Observation 91d61d80-9a94-43c0-bbbd-640ab06fd754 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.562762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.562762Z digest=sha256:5edcd27322456f7ed54af9a0a1dbf81693ce70b95149d486a60c1a97a952d3c1

Observation d0847c30-0dc7-4bb2-b24e-876bd2a7af18 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.663383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.663383Z digest=sha256:6873c0969061b250b9373805843c93bf369792a969d3e2ce2bd9c31bc6fc2720

Observation 86902715-0dcf-4e6c-ac7e-5e21f4300593 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Towards Reliable, Uncertainty-Aware Alignment Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.718260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.718260Z digest=sha256:d255ac3d11836e247f4e056568069cc834832d15d5f318ba7b0a46955ce462f4

Observation b9e8f8b0-26ef-4a1d-ab22-ede3d8ba02d1 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Towards Reliable, Uncertainty-Aware Alignment Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.197653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.830385Z digest=sha256:abc3365bc8c5b73a5e6bae8b53f4d5c2fe0caf2007dfb35319284bdbcd793581

Observation 9624b5bf-a1d1-4fa5-8d09-f4b118e216e4 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment Qwen2.5 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.864241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.864241Z digest=sha256:c088dfe369c4b26a7a03f7d62cb8654e50dec4e9781b59efcc71bb1c7996923c

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:01513fea1631286d304fa43c20095f0dbf595b53bfbd40fd73982963c869111b

Observation 07acfe67-c226-46b8-9be1-2d768943a270 · outbound

This paper cites Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.065331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.065331Z digest=sha256:42ccb8c09b29d0cbd8aa7ac16bc8bcc9ce9fc2ee1d121fdf99623e1b8b0bda80

Observation 1cc00e7e-3488-4566-9cd4-acba6e84c90a · outbound

This paper cites Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models.

Towards Reliable, Uncertainty-Aware Alignment Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:18.999138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T15:40:18.152809Z digest=sha256:967cb6c75073d8d7489054fdc9c5b7bebb64a587b25fd60474135a5a6ac0c0b9

Observation ac20c0ed-0aa3-4bb5-bf93-3e9d82819382 · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Towards Reliable, Uncertainty-Aware Alignment Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.178863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.178863Z digest=sha256:63b0dfad9eee9e102a1f1cfce5ac90ab668b492967927e9f6482010aae381296

Observation 5790b6ad-18f7-4d2b-96cb-4127ecc14bb7 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Towards Reliable, Uncertainty-Aware Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.225082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.225082Z digest=sha256:913297c37f55a35ae9d9bec51054f63e3bbd0a4158701e821ad29a9aedc7fe28

Observation 6d9b15d1-25f5-428b-b2c2-6e41fc92c2a5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Towards Reliable, Uncertainty-Aware Alignment Fine-Tuning Language Models from Human Preferences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.315140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.315140Z digest=sha256:ae1825c035b1fe59216d59b0827af4598415f908234a0af039383c95a9b1a2e3

Pith citing papers

No inbound Pith citation observations are available.