Pith. sign in

Paper Citation Record · LEDGER

Aligning LLMs with Domain Invariant Reward Models

As of 13 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2501.00911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00911 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.452310Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec3479bc-16d1-42ff-a4ff-c498c8b31ee3 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.434617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.164501Z digest=sha256:5ab5010736b6eb10aeb7a24248bcf22a716f1807f337ba52e0ee662174f10dd8

Observation 4b844550-bd51-45df-8c11-613835eda028 · outbound

This paper cites https://claude.ai/ Claude.

Aligning LLMs with Domain Invariant Reward Models https://claude.ai/ Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.417746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.170691Z digest=sha256:0d495f5e8d23475feb707ca244dbf6cdc3b8fc20b663d7a7e998c3c6bf3c4cef

Observation 4a3f310d-bb93-4017-8622-985e59cf9656 · outbound

This paper cites Wasserstein GAN.

Aligning LLMs with Domain Invariant Reward Models Wasserstein GAN

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.176006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.176006Z digest=sha256:42eaef17d8a7a85c63cd5e3767d3e5f277908797526e3f0d9b8d00d4490386b2

Observation 22a944a4-4d6f-4d3f-8f29-81a8644bf16b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.181779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.181779Z digest=sha256:1dcf1cafc13d3097d8ed82ccd1b66d647f51b95bf691642f05b7be59568a75b9

Observation 4a0b8ffb-f07f-4273-8905-0b6072b4c0e3 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Aligning LLMs with Domain Invariant Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.187680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.187680Z digest=sha256:1a417ce1f8a9178ced42a8de9246808e2aaf6db7b4cbd3febc04374e69f765ed

Observation 9fd5e676-b502-475b-96b3-aa5f68fe30fd · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.390203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.198606Z digest=sha256:579e7dc7c8e353035c1615fc07d3128ca2f116d533623e2c7522327105f32639

Observation b63d2dc6-dc34-44d1-b8f6-9dc536e35640 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.370858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.204272Z digest=sha256:2c745922e0c8fbd16795f5c2c363c66ae72fcdc551803d43edf7a68bf2c0649d

Observation f22e4843-f4f6-4675-a54d-bd31bcd59cb4 · outbound

This paper cites The Llama 3 Herd of Models.

Aligning LLMs with Domain Invariant Reward Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.209009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.209009Z digest=sha256:d3571da0af2a2673266e3f3ca0841c0e20335b689c7a90bc7503b6d95e2cadc0

Observation 51e8f850-9bfc-4a31-b938-e890656a2175 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.215495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.215495Z digest=sha256:7d9f826a02f10ca5bb91100f6a3e763fca3aa12c832524808a1202ac68c23c45

Observation 3beec7c8-733f-454b-8434-ebf1fdfb028b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.339535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.220549Z digest=sha256:918eb46394141c4c311ee6e4c9d0e6a9f8f70297e7f40b408f2fd3527dce743d

Observation 80268a2a-c2ba-4c37-bdda-e06be7ef805e · outbound

This paper cites Ustinova, Hana Ajakan, Pascal Germain, H.

Aligning LLMs with Domain Invariant Reward Models Ustinova, Hana Ajakan, Pascal Germain, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.321934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.225250Z digest=sha256:e18a130317c4bfa8323c4fcea3507d3b4f7b01f4c7a7d1e7c4b381807217e20a

Observation 0eff83a4-e1f8-4ef7-ba63-5a00b1299d15 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Aligning LLMs with Domain Invariant Reward Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.230638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.230638Z digest=sha256:2ced0e748bba2526afd69db7517cc35932a5eb59ffd9e2d6c4143e687e948c3a

Observation 2d8c36a4-6021-4a39-bfe8-28126657e856 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Aligning LLMs with Domain Invariant Reward Models Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.235972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.235972Z digest=sha256:8492dd928d33e84122c6c2d0bdabd709a3b94925d237ec9cca767dde441f5045

Observation 7d03db23-832b-4021-977d-f7bf771db709 · outbound

This paper cites Improved Training of Wasserstein GANs.

Aligning LLMs with Domain Invariant Reward Models Improved Training of Wasserstein GANs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.241129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.241129Z digest=sha256:df78ed1444c5117e8804ceaa6a446c8a3e99373a680f992fde8e5726e2fbc768

Observation 525eec75-07ad-4631-9a9f-da21d0e2f561 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Aligning LLMs with Domain Invariant Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.246781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.246781Z digest=sha256:312d5a3c51d2c1db8e5617eba1610c5d5d851e324593a0b0f814a7fbe95ae0ba

Observation 9715775d-83cb-40a2-8121-0d37b211048c · outbound

This paper cites The Unreasonable Effectiveness of Easy Training Data for Hard Tasks.

Aligning LLMs with Domain Invariant Reward Models The Unreasonable Effectiveness of Easy Training Data for Hard Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.252155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.252155Z digest=sha256:4a969032ec076e3aa2dc43151aecd8d99b27ab9ec117d4048201a4a0d2b69fd4

Observation 0482d9d0-e854-4409-83b1-de60c72f6608 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.302982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.257783Z digest=sha256:613fa9ce11ed8ffb19d23c2d7c3eb2f61b6069d3cf582db93646869957163028

Observation 1b571a44-facb-4844-b079-229b3f4fc963 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligning LLMs with Domain Invariant Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.262601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.262601Z digest=sha256:be0a114ec3d39dc71434bd517368e0228e564fbd4eb3179c69755a325037e004

Observation 1c2111e6-485b-465c-b7c6-982e37b8167e · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.285434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.267497Z digest=sha256:267f2957e693a34a02f562b6b12903a01f5f5f891eb3f81e54c05656bf1deb75

Observation f719bc80-b061-4a16-9be3-dd5ca1ac8290 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.268035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.272533Z digest=sha256:ad83439c0627de63b7650d80d703fb2c35e9eecd737c3bc48768e356090e456d

Observation 5aa80a67-2fb3-40b9-af20-44ae63ebf13b · outbound

This paper cites UDALM: Unsupervised Domain Adaptation through Language Modeling.

Aligning LLMs with Domain Invariant Reward Models UDALM: Unsupervised Domain Adaptation through Language Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.277915Z digest=sha256:a289af81cbf9e7c67d3b0632c37ace568e3577c6a63d74cad7dd931fc1174f60

Observation f121deb6-27f1-4e45-a3cb-39abe0b41336 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Aligning LLMs with Domain Invariant Reward Models Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.283131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.283131Z digest=sha256:31ee3558a3a327b3339523923de8bdf7e1298234681200c03e06333cda8de398

Observation 399c1371-c5ea-41d9-900a-eb1a7ec50e34 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.247219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.288497Z digest=sha256:b14fc64b5d97ef7f2de906145881cbe057b388beff3f2f620bc72a426495e9ee

Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.293458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.293458Z digest=sha256:c0cd85f6e0dd5a16755b94ba4ab5956207ddfe006a325ff87555d9e79bff9cf9

Observation f67971a3-3e53-47bb-bb5f-339ef22d38be · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Aligning LLMs with Domain Invariant Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.299196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.299196Z digest=sha256:5568dc850f78e335d9ff43b88aca474acf9243fb3148992c038049ca579f7437

Observation 2855cfac-44af-460d-9951-9d95e4373c8c · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.304282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.304282Z digest=sha256:359a4176e7187f104de4229b04d6dc48cadeab497b40a8932be007e11b3f7c21

Observation ad49babe-b993-4957-9fee-352854386105 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.228994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.309736Z digest=sha256:27db43ddd9ebdeb27ee0d6bb3aec1b287f039bdb4d906b65c1f6624a0ea5fc16

Observation adb34e46-df8e-40f3-9fda-79761d91fc62 · outbound

This paper cites Preference Tuning For Toxicity Mitigation Generalizes Across Languages.

Aligning LLMs with Domain Invariant Reward Models Preference Tuning For Toxicity Mitigation Generalizes Across Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.315612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.315612Z digest=sha256:678683283fc7c11e99937321d7e16bfa6fd401c0661c3b20e48c5a412d4b84ab

Observation 73d1d84f-8ece-4bb0-8598-380f9625dd49 · outbound

This paper cites Decoupled Weight Decay Regularization.

Aligning LLMs with Domain Invariant Reward Models Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.322212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.322212Z digest=sha256:f7071714ae84bf847388639e1533d577e73c56a4a44f15d612e058d8493fd51b

Observation df829618-1316-49a3-ba2b-111a0c5aa761 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Aligning LLMs with Domain Invariant Reward Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.326988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.326988Z digest=sha256:dc840b343ff4e6255838a031d4356beb2ca51b76af6bd2a2f9317d5251258a02

Observation c9293c1f-65b5-46e0-81f3-3e1e4d1441dd · outbound

This paper cites https://chatgpt.com/ Chatgpt.

Aligning LLMs with Domain Invariant Reward Models https://chatgpt.com/ Chatgpt

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.211159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.332430Z digest=sha256:735736218c26fd7a2274f3491e4710a1ad10862896ac6ab52b1d3b894646883d

Observation aad28e62-e728-4eea-bb76-3be7c1245df8 · outbound

This paper cites Training language models to follow instructions with human feedback.

Aligning LLMs with Domain Invariant Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.338619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.338619Z digest=sha256:d91d03994d58ea763af084a9d1944319348e875b347e5cacf3a9dbc2f13b992d

Observation 8eb05ed5-6849-4bc4-ad4b-7687b0f1e684 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.344286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.344286Z digest=sha256:4b8a8e48ea9bb8b8586cf91d0886b94fee039170931ca8bb00b4b79f0c252c22

Observation b2c82572-4120-4f94-971a-a25ff1d87f10 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Aligning LLMs with Domain Invariant Reward Models Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.349688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.349688Z digest=sha256:f0666d81290f59ba5b7b04df1af497077d01daee4b4c805aef4b837cb870220f

Observation c68407d2-bd69-4615-b5b9-025746aa746f · outbound

This paper cites Aligning Language Models with Demonstrated Feedback.

Aligning LLMs with Domain Invariant Reward Models Aligning Language Models with Demonstrated Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.355146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.355146Z digest=sha256:b2d78418c1d9ce56ea474f8b79ed2597068dd1fd9a2d03cbddc532671db193e8

Observation bb8d7799-1e80-43f4-8d2a-29b83943b9e6 · outbound

This paper cites Wasserstein Distance Guided Representation Learning for Domain Adaptation.

Aligning LLMs with Domain Invariant Reward Models Wasserstein Distance Guided Representation Learning for Domain Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.361165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.361165Z digest=sha256:8f81796939ae84ab23151a1b92c533069977fb486da0850d54fe87b35cf9d4a8

Observation 9d92ae9d-cbe4-41d3-96c7-d9cd1c25a657 · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Aligning LLMs with Domain Invariant Reward Models Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.367288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.367288Z digest=sha256:5460bdc9553643e515593c911eec0f0ddb912dc115d746afd0f9cd9f1bcb6051

Observation 2de6c780-9bdd-4dee-835c-7e2d8618b558 · outbound

This paper cites Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment.

Aligning LLMs with Domain Invariant Reward Models Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.372847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.372847Z digest=sha256:4f7e413836856ffe13b25a7bb05033215ff4d1d704b799fdaa3612e0ddbceb3c

Observation ef303f7c-fb2a-4fb1-8fde-c11ed3a7cf1d · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Aligning LLMs with Domain Invariant Reward Models Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.378181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.378181Z digest=sha256:181dbfd8abd0b3ccf03ba0afe843bfd89e5c8d0a036b5f871702c7c411c06a44

Observation 15eaf064-48e0-4a6e-8d5e-c6c2f796e857 · outbound

This paper cites Chernova, and Dhruv Batra.

Aligning LLMs with Domain Invariant Reward Models Chernova, and Dhruv Batra

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.181683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.383414Z digest=sha256:c335f675146e12c1d54d684a61752abc455b4207ac2b0b081cbcd82edf11177e

Observation d9e7e8de-fa6f-4e8d-a26d-4cccde6dc901 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Aligning LLMs with Domain Invariant Reward Models Deep Domain Confusion: Maximizing for Domain Invariance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.388655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.388655Z digest=sha256:e4f1c2a131d34981b226d69de14a5f5c97a2e034d9b33d4b216afc5bc8237b3f

Observation 9e561a12-24a5-4b02-b30d-94e811486f54 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.393710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.393710Z digest=sha256:993bf30b18db011d149d04b47e83b3c909108a043d592af7825e1370f0ecd553

Observation 6adc45c0-0464-4e51-a507-fb1f30304374 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.161418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.398638Z digest=sha256:167b5e8ab8f3820f5dfdfbc0e99d1a388ccf1316d166f6e917ad21974c84a964

Observation a05d9238-5b9f-42d9-b0e6-ff3ebf2d2d76 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.141030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.404542Z digest=sha256:9957f9e96a0e25a8040f64e2aea974f9dc0a4b9392266c6469c31090d4d2b9d0

Observation 1340b498-31a0-45e7-b283-1ec86018401b · outbound

This paper cites Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment.

Aligning LLMs with Domain Invariant Reward Models Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.593278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.409531Z digest=sha256:98dd608d3f3fc67e6da452bc4eb3cb8c023ce9f36f627df1154a0e5b460f2c24

Observation 97b509e2-18b3-48b6-9e74-8dfa8fc741ee · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

Aligning LLMs with Domain Invariant Reward Models CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.414979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.414979Z digest=sha256:6a7c54d93c70f159b811bf6eb72eff420809c739ec4b4455d0af8db1d56e5c8a

Observation 99232992-40d2-465d-94e1-27471d260e93 · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Aligning LLMs with Domain Invariant Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.420151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.420151Z digest=sha256:99800ffb8222b4de38decee1a674fb7daa8ff461995cb38786896b80eac737b9

Observation e724335a-a9ae-42b5-9421-b8240cf682a8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Aligning LLMs with Domain Invariant Reward Models Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.425370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.425370Z digest=sha256:7efbe5968574ef6a6762107a307f915ed6ec8d541ec565e2a59a4f166d62db38

Observation 0496f569-b2bf-4586-9b6e-075691c6365c · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.431407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.431407Z digest=sha256:064e8818b0acfac455d466bc5717183cc0d3d29f8e9bcfe641f53d4b32d57e1c

Observation 4a406a80-2238-4f13-bd0a-8ba6f59e3246 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.123721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.436817Z digest=sha256:bbb6c024d07754559f5edcb5658ecad3f665d4bedbe6852bbfe51ce2fb8b8704

Observation f39061c4-0a81-46a2-bb04-17b0c58d1b40 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Aligning LLMs with Domain Invariant Reward Models Fine-Tuning Language Models from Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.441631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.441631Z digest=sha256:72f38f86b679a0f6e5bc977123e6c194aa69d8afa58a3c2e5ee42844a35c06fd

Observation b527aeef-d4b6-463d-9df6-3949ecd8a411 · outbound

This paper cites online" 'onlinestring :=.

Aligning LLMs with Domain Invariant Reward Models online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.446752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.446752Z digest=sha256:afd2feb1265d10b29727533f9691e9d827ba37c3ff2f17356384a4e788f76bac

Observation 1adb606a-3623-4c4a-a5c1-dfb6f422aee9 · outbound

This paper cites write newline.

Aligning LLMs with Domain Invariant Reward Models write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.452310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.452310Z digest=sha256:90a624e6f1caa4cedd0fd4c11203666ed8deead23cedd7db3a9658cad44774e1

Pith citing papers

No inbound Pith citation observations are available.