Pith. sign in

Paper Citation Record · LEDGER

Towards Reliable, Uncertainty-Aware Alignment

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2507.15906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15906 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:18.315140Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d809cf6f-ee7c-4ffd-892d-e7588102b78c · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Towards Reliable, Uncertainty-Aware Alignment The claude 3 model family: Opus, sonnet, haiku

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:21.156461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.176395Z digest=sha256:94e2fda510f31eeb0b76df52593e5e1021431c3964b6bed6b4422c966c6a73f5

Observation 47697553-f704-4bb3-82e3-bfc825ef34c7 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Reliable, Uncertainty-Aware Alignment The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.266168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.266168Z digest=sha256:40d753a0bc8f84ed110da0346cd50a5109ebf6836b9ca357f6ebe7c71d1ec1b8

Observation ee77b394-d87e-49f6-8280-3c14d712c9cb · outbound

This paper cites Concrete Problems in AI Safety.

Towards Reliable, Uncertainty-Aware Alignment Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.333386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.333386Z digest=sha256:51e4255b2ce3e68a484a4e2c94915fe0f443c00001525aac95e92556817e38c1

Observation fbb27280-7859-4c32-8d4f-a42d30e3c2f0 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Towards Reliable, Uncertainty-Aware Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.417979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.417979Z digest=sha256:73d3a908179bff8c6cde5f6076a3367ece63600c9affb2d0b913412bc2f1f292

Observation 48ae5dfa-eb9f-465e-aa15-b6cc4542ab59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.494528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.494528Z digest=sha256:65da800c0a510965b9419adb05722c727e1abea0ba23be3f58312951e0a1c058

Observation 1728c136-ef99-4e4d-a6ff-0be5e299c464 · outbound

This paper cites Deep reinforcement learning from human preferences.

Towards Reliable, Uncertainty-Aware Alignment Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.621778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.621778Z digest=sha256:1f47b8010b38d00d02af190dd54ba25f4a815734d533db9e12128d691d5dc425

Observation d81661bf-1992-4d60-91f9-a40889259d1e · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.721341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.721341Z digest=sha256:0a298f48de40fb189e4313b487781fc7f02f21745812ffbbce12ba1bd2acd222

Observation 08da191c-c502-4312-bed6-4226835d001c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Towards Reliable, Uncertainty-Aware Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.818475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.818475Z digest=sha256:0ab7ebee80bd308b997b984c78164bd14bdd113f11b5ae7fbbd2fbb7e56eeb9a

Observation 1d279d30-0dcf-4f19-bfaa-89c952a0a89f · outbound

This paper cites Suphavadeeprasit.

Towards Reliable, Uncertainty-Aware Alignment Suphavadeeprasit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.952919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.865665Z digest=sha256:8a56a1b49860c497597b93e1141a3b3e659d591f72c923ed276a7853c4a45dba

Observation 5e10c0f0-fb17-4bfe-9a8f-44419e76630c · outbound

This paper cites Gemini 2.0 flash: Next-generation multimodal ai model.

Towards Reliable, Uncertainty-Aware Alignment Gemini 2.0 flash: Next-generation multimodal ai model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.783795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:13.944596Z digest=sha256:4bfadeaf11ecb74784fbc85a8f65e5f68b05fa61891fdd6dbe9d413b23ccdcfa

Observation 1361e871-f6d0-4c4d-9c6c-b5e0025049e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment DeepSeek-V3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.010645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.010645Z digest=sha256:0609b760b1c1f46f7dbae77a154c0f2d4e2ed53ee8c2cfc4c089408d79c1737e

Observation 97e0d881-e718-487f-a5bb-1e108dcacce4 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

Towards Reliable, Uncertainty-Aware Alignment RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.132426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.132426Z digest=sha256:b61b82bb40b44da643c6983eb0d7a6294e5d953ac30fff150bfb7325e3ba6205

Observation e91f2dfc-d7f7-4e9b-8792-7f4fef976c8f · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Towards Reliable, Uncertainty-Aware Alignment RLHF Workflow: From Reward Modeling to Online RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.208038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.208038Z digest=sha256:ed7536d8618f77881e841ded49d636b301e241f5d6348f466f76e4a6d12e8712

Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.278132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.278132Z digest=sha256:b57c16a51a9cbef8f4767f999a0c03285bed746bca853fe563ad432d2b550626

Observation 0e5c7de4-926c-4d9f-8825-509c26dcdb16 · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Towards Reliable, Uncertainty-Aware Alignment Understanding dataset difficulty with v-usable information

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.612471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.350372Z digest=sha256:9f15d68089e6b7f84b9cdeb0689f98d3ef777abab62dd5f1de1ae4666c008f97

Observation cd467f85-11b8-4976-a8fe-6efab80bbead · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Towards Reliable, Uncertainty-Aware Alignment Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.401191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.467970Z digest=sha256:db18c1ef18f3458612a2c8bfcbc31e7dadd9744ebfde1d50dddd9b8a185bb7a7

Observation 4115c2db-0a6e-4527-b18e-31c2da76f903 · outbound

This paper cites Scaling laws for reward model overoptimization.

Towards Reliable, Uncertainty-Aware Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.648398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.648398Z digest=sha256:70e8b9568a786b2f3fdd0d74e9ef9be4c5612dfe918c7bd9bc63c93818bc3b5d

Observation 8f8c4ae2-e1e8-4d32-b9a1-69738b2f53c3 · outbound

This paper cites https://huggingface.co/google/gemma-2b, 2024.

Towards Reliable, Uncertainty-Aware Alignment https://huggingface.co/google/gemma-2b, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.250009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:14.807037Z digest=sha256:5860ceed788bb0f2b356203e12fb05724dbe3d6976fe53296bccb6cc8f248fee

Observation 1dc032e9-a02f-4dce-b96d-6b068474f4da · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Towards Reliable, Uncertainty-Aware Alignment Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.973348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.973348Z digest=sha256:fcd6041907b076a09a9db62c7aa9c356957dbe2fdb1bc8564a5d9b07c0d282d2

Observation 0130163d-ee01-4d80-8e01-84e6c0938b5b · outbound

This paper cites Mistral 7B.

Towards Reliable, Uncertainty-Aware Alignment Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.092213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.092213Z digest=sha256:39fadfdf193a669bd78eab1b9b9250977fad7c2eb4bee7ba966bdbab8322ab4d

Observation 33e0d3f5-5f7a-4d3d-a8d7-a3ec2544814f · outbound

This paper cites Reward Design with Language Models.

Towards Reliable, Uncertainty-Aware Alignment Reward Design with Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.257918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.257918Z digest=sha256:25fefd7502feefa56f4c43d75208c8c9869c0a166493aef1678976835ae6da82

Observation 61587e0a-0251-4e40-b1ba-8cc8fee53523 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Towards Reliable, Uncertainty-Aware Alignment RewardBench: Evaluating Reward Models for Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.368188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.368188Z digest=sha256:83df770b6cbe7d3241e95b7b14dff1ba676547f40f94ed5f73c9e5342d66637d

Observation 24b201b5-75dd-4e28-9733-9bee5051bfbc · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Towards Reliable, Uncertainty-Aware Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.487856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.487856Z digest=sha256:797709c5253320f71fc89cd5a39466f063a6af67043c7f40268af3deb9a7581f

Observation 6f0b80df-145c-4ef6-af21-d3a6aedfcba2 · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Towards Reliable, Uncertainty-Aware Alignment Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:20.024665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.623570Z digest=sha256:9a6a5999b27d0d3ce8bb767fe5fe28fe1af895d29e2340cabef33d3836362c91

Observation 1544c899-73e4-4a7b-9a41-91ea72ae8d17 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

Towards Reliable, Uncertainty-Aware Alignment Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.689847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.689847Z digest=sha256:cdec1d9e5cc7e8eba1c8ea5a689f2385b218ec3dc6bdb845fe23c456d04c965b

Observation 0babacde-4ada-4f0a-adcc-94e46862e9e8 · outbound

This paper cites Iterative prompting for estimating epistemic uncertainty.

Towards Reliable, Uncertainty-Aware Alignment Iterative prompting for estimating epistemic uncertainty

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.849773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:15.762639Z digest=sha256:bc0a3613f411c969e1d7d5aad6d6810571146cfaf9ef5c7eaaa18e85b56f7649

Observation 902799f6-fb92-4983-98a6-59f7a3fff674 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.880022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.880022Z digest=sha256:c5be207a8bb10a314fcc56d7cb820cc58b3b60d0f06b8f1235a455bbb7ecb4e3

Observation 28c4d9ce-4dcc-4d52-85c6-fc6d793041e9 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Towards Reliable, Uncertainty-Aware Alignment Introducing meta llama 3: The most capable openly available llm to date

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:15.957662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:15.957662Z digest=sha256:2e0262405842a270791b08cca11d9522d2fb92c589bf5adb4e4593d9675a137f

Observation e439b574-48dd-419d-a3bb-64f9c38ba7af · outbound

This paper cites Montgomery and George C.

Towards Reliable, Uncertainty-Aware Alignment Montgomery and George C

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.638664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.020519Z digest=sha256:e6786b70f77a363923625a15ae768ab6b15624a68e0835cbd54d728af0ddca44

Observation 757574e9-7c31-43a4-abd1-614f6e78f258 · outbound

This paper cites GPT-4 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.098615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.098615Z digest=sha256:54bdee8c67262e83b7d388c57135420b8f98e1c1069574d66a37f326bf6b9629

Observation 91246783-23e2-48a4-8de0-3e8e6fa149ca · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.193548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.193548Z digest=sha256:b460ceece6c01683558e02214b360a2d23d6d61e9f5df3e53c3bda847bea3da7

Observation 18c522b9-122d-4a4a-adbd-92b17f02a131 · outbound

This paper cites Language models are unsupervised multitask learners.

Towards Reliable, Uncertainty-Aware Alignment Language models are unsupervised multitask learners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.285137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.285137Z digest=sha256:2717b8ea607bc1d87866dca5ffb78d2aa78c78e2ddd0a02c9c5371132019f276

Observation 57c3ab16-89df-48eb-a27d-8adab523beee · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Reliable, Uncertainty-Aware Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.368788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.368788Z digest=sha256:80abd0fed4a94aef5ee3438b4bc02d36cde30d259e8b9d21cf539635aba3804a

Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.470462Z digest=sha256:cf435ff7abb4f2caba9ff7f374311dd882cf54558c0779f61f1a532e458292f0

Observation 9b210cb0-f49a-48d1-aea4-6c6df8c36850 · outbound

This paper cites Trust Region Policy Optimization.

Towards Reliable, Uncertainty-Aware Alignment Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.538614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.538614Z digest=sha256:aa2bca633c3de0dd5166943ce116997d1e718bc39efa2778fea8aa7ed70d95e4

Observation 58d43223-87e6-472e-b055-18d7a48a7cdb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Towards Reliable, Uncertainty-Aware Alignment Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.600627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.600627Z digest=sha256:dee8ba1e51e544ee25ed729412c7ce1d8c8b88437f99e0a99868c6aebb71bcea

Observation 98871c98-ad6d-461e-90d4-40083a555d8e · outbound

This paper cites Mutual fund performance.

Towards Reliable, Uncertainty-Aware Alignment Mutual fund performance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.401125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:16.680687Z digest=sha256:b623d7afcb4e98d0570fb2ae9f442c540e40536c40ab135795dae6dcbcc0c980

Observation a399d956-22d0-4734-b28b-5c7ad4a851cb · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.751394Z digest=sha256:0f1b5994cd967b276bd36f5b5e391d49f3381bd3a1daea245a96fb7ea2a0f329

Observation 9aeebd62-11a7-44fa-94b3-9d09552a9746 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Towards Reliable, Uncertainty-Aware Alignment LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.846785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.846785Z digest=sha256:4c97b14a744df6cfe78a1c08dade94b3a89ebd35c8aecc9d1563dc3e2bb6125b

Observation cc34c0a9-eb6a-439b-851c-fc50b32b7883 · outbound

This paper cites Learning to summarize with human feedback.

Towards Reliable, Uncertainty-Aware Alignment Learning to summarize with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.908542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.908542Z digest=sha256:9ec8e97bad2d0098ae5879b44566d26eafa32bfd0934b6f8c4ec9e08218bf52c

Observation 362f149e-6940-42e8-8316-2a92bb0a8b50 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Towards Reliable, Uncertainty-Aware Alignment Policy gradient methods for reinforcement learning with function approximation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.999910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.999910Z digest=sha256:761d8c874b299b0530902a49f8c7005d3b6ed3f4b359ed8ba8d8805c66958e80

Observation b75730e9-da5d-469a-ad94-d67b210d359e · outbound

This paper cites Quantifying Uncertainty in Natural Language Explanations of Large Language Models.

Towards Reliable, Uncertainty-Aware Alignment Quantifying Uncertainty in Natural Language Explanations of Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:40:18.689752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.037690Z digest=sha256:a270aa058825488b7645355567c45d5677580f9b010ef7a165ff162ec816f478

Observation 99d1eca3-0d57-44c2-89f2-f0e50ca4e355 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Reliable, Uncertainty-Aware Alignment Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.125727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.125727Z digest=sha256:ba5c13552d0e3dd13a6f2df1cb7a21da8f06928e55be86cd62947bcd9141f29f

Observation 6b251672-a983-4cec-8bd6-84eb6017d14b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Towards Reliable, Uncertainty-Aware Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.194495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.194495Z digest=sha256:f0544dfdcc75a115c09b1e706ea32a7cc9116fce11679bbb695d76d729c24a06

Observation 8d6d2705-1af5-4aef-a24c-771fdf608aa3 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Towards Reliable, Uncertainty-Aware Alignment Gemma: Open Models Based on Gemini Research and Technology

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.315138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.315138Z digest=sha256:928c83e171650f9221226e3b30a63e81fafe708f62d787a58c4e5f26cb71a86a

Observation 4b7ec90c-cf1c-40a0-b378-fd8e214e9552 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Towards Reliable, Uncertainty-Aware Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.350383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.350383Z digest=sha256:88b4c82307c9de53525b3b9b500b339165b6906157fac6948f73bfe92279a3ec

Observation 98e4ed2e-64d5-47b2-b840-901b19b509cf · outbound

This paper cites Trl: Transformer reinforcement learning.

Towards Reliable, Uncertainty-Aware Alignment Trl: Transformer reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.443886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.443886Z digest=sha256:aaaba84bcdb5caf501fcd1fc929fdf772383b8998ebcf50dc840149cdd7f30bb

Observation c8a35bd8-9eeb-4310-b713-fbe3f304dbc7 · outbound

This paper cites HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM.

Towards Reliable, Uncertainty-Aware Alignment HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.507422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.507422Z digest=sha256:b7b8bfe75d61314bb067397df73b09f68bd1a2ead8a0a3df7f1bd51bc68f730f

Observation 91d61d80-9a94-43c0-bbbd-640ab06fd754 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.562762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.562762Z digest=sha256:5edcd27322456f7ed54af9a0a1dbf81693ce70b95149d486a60c1a97a952d3c1

Observation d0847c30-0dc7-4bb2-b24e-876bd2a7af18 · outbound

This paper cites an unresolved cited work.

Towards Reliable, Uncertainty-Aware Alignment Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.663383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.663383Z digest=sha256:6873c0969061b250b9373805843c93bf369792a969d3e2ce2bd9c31bc6fc2720

Observation 86902715-0dcf-4e6c-ac7e-5e21f4300593 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Towards Reliable, Uncertainty-Aware Alignment Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.718260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.718260Z digest=sha256:d255ac3d11836e247f4e056568069cc834832d15d5f318ba7b0a46955ce462f4

Observation b9e8f8b0-26ef-4a1d-ab22-ede3d8ba02d1 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Towards Reliable, Uncertainty-Aware Alignment Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:19.197653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:17.830385Z digest=sha256:11cb51911e4d60e7f46c278aa82342b487a8efbf271ce526d502716d142f3a5e

Observation 9624b5bf-a1d1-4fa5-8d09-f4b118e216e4 · outbound

This paper cites Qwen2.5 Technical Report.

Towards Reliable, Uncertainty-Aware Alignment Qwen2.5 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.864241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.864241Z digest=sha256:c088dfe369c4b26a7a03f7d62cb8654e50dec4e9781b59efcc71bb1c7996923c

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:bc21bff09b4a8d467b455da654c5f41ba520fc4eb2777476b2717a97f423b92c

Observation 07acfe67-c226-46b8-9be1-2d768943a270 · outbound

This paper cites Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles.

Towards Reliable, Uncertainty-Aware Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.065331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.065331Z digest=sha256:42ccb8c09b29d0cbd8aa7ac16bc8bcc9ce9fc2ee1d121fdf99623e1b8b0bda80

Observation 1cc00e7e-3488-4566-9cd4-acba6e84c90a · outbound

This paper cites Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models.

Towards Reliable, Uncertainty-Aware Alignment Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:18.999138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T15:40:18.152809Z digest=sha256:24ffccac81d4b0dce0b00ae3fab7e07aaa4e0a26df8cafa97efa1c689f5cdc7e

Observation ac20c0ed-0aa3-4bb5-bf93-3e9d82819382 · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Towards Reliable, Uncertainty-Aware Alignment Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.178863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.178863Z digest=sha256:63b0dfad9eee9e102a1f1cfce5ac90ab668b492967927e9f6482010aae381296

Observation 5790b6ad-18f7-4d2b-96cb-4127ecc14bb7 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Towards Reliable, Uncertainty-Aware Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.225082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.225082Z digest=sha256:913297c37f55a35ae9d9bec51054f63e3bbd0a4158701e821ad29a9aedc7fe28

Observation 6d9b15d1-25f5-428b-b2c2-6e41fc92c2a5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Towards Reliable, Uncertainty-Aware Alignment Fine-Tuning Language Models from Human Preferences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:18.315140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:18.315140Z digest=sha256:ae1825c035b1fe59216d59b0827af4598415f908234a0af039383c95a9b1a2e3

Pith citing papers

No inbound Pith citation observations are available.