Pith. sign in

Paper Citation Record · LEDGER

Aligning LLMs with Domain Invariant Reward Models

As of 13 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2501.00911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00911 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.452310Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec3479bc-16d1-42ff-a4ff-c498c8b31ee3 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.434617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.164501Z digest=sha256:18d101a5d099719b98d9eea7d2e0de9d30c0894cfb769704eb0eb4684067684e

Observation 4b844550-bd51-45df-8c11-613835eda028 · outbound

This paper cites https://claude.ai/ Claude.

Aligning LLMs with Domain Invariant Reward Models https://claude.ai/ Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.417746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.170691Z digest=sha256:fd9afdfeb5b22745443ed7058dc9d07733c12dbbf6556cb14e217c2d72cd8958

Observation 4a3f310d-bb93-4017-8622-985e59cf9656 · outbound

This paper cites Wasserstein GAN.

Aligning LLMs with Domain Invariant Reward Models Wasserstein GAN

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.176006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.176006Z digest=sha256:d5192cab71a015f0620e09fc70e5e9e4940687064240f78c1ddc2c4506636a67

Observation 22a944a4-4d6f-4d3f-8f29-81a8644bf16b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.181779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.181779Z digest=sha256:8ec327b82bd8ecc225593f46c516b214e40f16da128374be6f9bb05400e90c75

Observation 4a0b8ffb-f07f-4273-8905-0b6072b4c0e3 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Aligning LLMs with Domain Invariant Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.187680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.187680Z digest=sha256:06e4be10e1cb9fefaaa06bd5a7334224f200a4b692854974769b3bb007052c91

Observation 9fd5e676-b502-475b-96b3-aa5f68fe30fd · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.390203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.198606Z digest=sha256:e53fd9297bc742bc41f3a0bdede935dd14f17830b40eb8415c3af562b86169f3

Observation b63d2dc6-dc34-44d1-b8f6-9dc536e35640 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.370858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.204272Z digest=sha256:4ab5c76797adf4b36b7a0ee55d0cb2691d9dd5ac05358eaaedd5163fde5e3c6a

Observation f22e4843-f4f6-4675-a54d-bd31bcd59cb4 · outbound

This paper cites The Llama 3 Herd of Models.

Aligning LLMs with Domain Invariant Reward Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.209009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.209009Z digest=sha256:9355817ab782018906893a441442c76fee3482ebf222f2e29e117dcdcedc2197

Observation 51e8f850-9bfc-4a31-b938-e890656a2175 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.215495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.215495Z digest=sha256:83f67cc5f24a7566444d619ebd813c3f84aedaa7cae638614cf9a98dc8bd25de

Observation 3beec7c8-733f-454b-8434-ebf1fdfb028b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.339535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.220549Z digest=sha256:d2d7934010dfcb60ff1f55300f521887b619b9b01a7fa560461dfa4530126205

Observation 80268a2a-c2ba-4c37-bdda-e06be7ef805e · outbound

This paper cites Ustinova, Hana Ajakan, Pascal Germain, H.

Aligning LLMs with Domain Invariant Reward Models Ustinova, Hana Ajakan, Pascal Germain, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.321934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.225250Z digest=sha256:4d1814e94a607a88c23ac027f8933cdb1a97e6ace5d197ab744049bc86a97806

Observation 0eff83a4-e1f8-4ef7-ba63-5a00b1299d15 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Aligning LLMs with Domain Invariant Reward Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.230638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.230638Z digest=sha256:c434cdbb1743b358563c2175cca982e890852858d126f8e1dea48b1c8f94fe11

Observation 2d8c36a4-6021-4a39-bfe8-28126657e856 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Aligning LLMs with Domain Invariant Reward Models Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.235972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.235972Z digest=sha256:d4d9f053a23f2238f2d3619d0114640f49225c0f1dc47c57d4933d96e5acc321

Observation 7d03db23-832b-4021-977d-f7bf771db709 · outbound

This paper cites Improved Training of Wasserstein GANs.

Aligning LLMs with Domain Invariant Reward Models Improved Training of Wasserstein GANs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.241129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.241129Z digest=sha256:418b22bcbc5f92fed71480f4aa18c175892bc778690d396174fa19fbb7b37d40

Observation 525eec75-07ad-4631-9a9f-da21d0e2f561 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Aligning LLMs with Domain Invariant Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.246781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.246781Z digest=sha256:30352f95a0dae2844b5750d4e93e231a0429c5b3330f6c6df0db35c0aedea1e5

Observation 9715775d-83cb-40a2-8121-0d37b211048c · outbound

This paper cites The Unreasonable Effectiveness of Easy Training Data for Hard Tasks.

Aligning LLMs with Domain Invariant Reward Models The Unreasonable Effectiveness of Easy Training Data for Hard Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.252155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.252155Z digest=sha256:94957d2f3147acffda0634c0bb6db8d6b4328949a456e536a276eb964a903172

Observation 0482d9d0-e854-4409-83b1-de60c72f6608 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.302982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.257783Z digest=sha256:eda2d4be8ee40b45a6270a59f58fc842cd2dc66a0ee0cda4bd810017aae08b35

Observation 1b571a44-facb-4844-b079-229b3f4fc963 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligning LLMs with Domain Invariant Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.262601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.262601Z digest=sha256:4ef1fb594d66917878f1fb0fe1032ec4937cecd84a98860bc7aa1985a55e2d08

Observation 1c2111e6-485b-465c-b7c6-982e37b8167e · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.285434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.267497Z digest=sha256:4a1b1f84708eb5654efc6c018a4532581916838007aedfbdd91ec117aabe80d4

Observation f719bc80-b061-4a16-9be3-dd5ca1ac8290 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.268035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.272533Z digest=sha256:027b0ede4842a94e458a5ad548630041e5d6f5d9d6927ff89b974bdf7a4fee08

Observation 5aa80a67-2fb3-40b9-af20-44ae63ebf13b · outbound

This paper cites UDALM: Unsupervised Domain Adaptation through Language Modeling.

Aligning LLMs with Domain Invariant Reward Models UDALM: Unsupervised Domain Adaptation through Language Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.277915Z digest=sha256:7eda1860ff6cf5e639688903dccc0ed5cd4e9c34ef362d2e589307688821361f

Observation f121deb6-27f1-4e45-a3cb-39abe0b41336 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Aligning LLMs with Domain Invariant Reward Models Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.283131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.283131Z digest=sha256:01625776eee47723bd43e1326b90b4c39b3294fb89f5e09a854bac5801441d75

Observation 399c1371-c5ea-41d9-900a-eb1a7ec50e34 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.247219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.288497Z digest=sha256:9b3ac23054a160cc7f04339a9534b1401c81690d4e2b0da614ff73d580079905

Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.293458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.293458Z digest=sha256:c76b60fcff9c3bf8ae82f2f706fba585ec285c3810755474a2b90325b1bf0928

Observation f67971a3-3e53-47bb-bb5f-339ef22d38be · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Aligning LLMs with Domain Invariant Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.299196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.299196Z digest=sha256:1f4e7df219bda411d03cd0f526bd27ef3f251ee340e4e926631b45d7937e5d1f

Observation 2855cfac-44af-460d-9951-9d95e4373c8c · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.304282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.304282Z digest=sha256:ebe38d94e749a08f5a855c08e2d4d29672d70bcc77515b64d2b265e2ea937b55

Observation ad49babe-b993-4957-9fee-352854386105 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.228994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.309736Z digest=sha256:ba58ce6402a0cab62bb587f9b3822f37e374e91f0886eccfe89a9275d3e044d3

Observation adb34e46-df8e-40f3-9fda-79761d91fc62 · outbound

This paper cites Preference Tuning For Toxicity Mitigation Generalizes Across Languages.

Aligning LLMs with Domain Invariant Reward Models Preference Tuning For Toxicity Mitigation Generalizes Across Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.315612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.315612Z digest=sha256:94d44e78f57e1d8707a7e1fc39a8a464e1c223ec7dbe2f1e749099824457c044

Observation 73d1d84f-8ece-4bb0-8598-380f9625dd49 · outbound

This paper cites Decoupled Weight Decay Regularization.

Aligning LLMs with Domain Invariant Reward Models Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.322212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.322212Z digest=sha256:29940630f80c792ab04d3b6d9e8e5f57ee5ed792318413bb4611c4b8a26d6f1d

Observation df829618-1316-49a3-ba2b-111a0c5aa761 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Aligning LLMs with Domain Invariant Reward Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.326988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.326988Z digest=sha256:550d4a6b891f8870cb86ba9b56accaaaf5476743919571cfd7774a84df298629

Observation c9293c1f-65b5-46e0-81f3-3e1e4d1441dd · outbound

This paper cites https://chatgpt.com/ Chatgpt.

Aligning LLMs with Domain Invariant Reward Models https://chatgpt.com/ Chatgpt

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.211159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.332430Z digest=sha256:b0aacf28e0c06d2c1e382321fa5ff80fac666a05a9c1073c995ce4761f7ab05d

Observation aad28e62-e728-4eea-bb76-3be7c1245df8 · outbound

This paper cites Training language models to follow instructions with human feedback.

Aligning LLMs with Domain Invariant Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.338619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.338619Z digest=sha256:0ad6cd84f568353b58cc0f50c2304d7f274effa6d254170d6de4e68ea9187ce9

Observation 8eb05ed5-6849-4bc4-ad4b-7687b0f1e684 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.344286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.344286Z digest=sha256:02f3a3fce8b5ec8c5c5104fd26c757ef6ec1326158e5c564578b3a480d13e946

Observation b2c82572-4120-4f94-971a-a25ff1d87f10 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Aligning LLMs with Domain Invariant Reward Models Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.349688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.349688Z digest=sha256:ca4632a2a59c48a6a82a29f5dffb9d8d6d82320d1be9316eef408399aa3d6cf6

Observation c68407d2-bd69-4615-b5b9-025746aa746f · outbound

This paper cites Aligning Language Models with Demonstrated Feedback.

Aligning LLMs with Domain Invariant Reward Models Aligning Language Models with Demonstrated Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.355146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.355146Z digest=sha256:e3b03ec44c589d8a9dc8619d11b7d8e065fc8271905304a259a78862cc164780

Observation bb8d7799-1e80-43f4-8d2a-29b83943b9e6 · outbound

This paper cites Wasserstein Distance Guided Representation Learning for Domain Adaptation.

Aligning LLMs with Domain Invariant Reward Models Wasserstein Distance Guided Representation Learning for Domain Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.361165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.361165Z digest=sha256:b8a0f47d308548510b5daa685406bd69e443c0e25025722b60e8f7d942d4a6c0

Observation 9d92ae9d-cbe4-41d3-96c7-d9cd1c25a657 · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Aligning LLMs with Domain Invariant Reward Models Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.367288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.367288Z digest=sha256:1038b08d6b953f2f8e66b174a4c3acafecf1e86fea8914e3eb9b8754254e13f9

Observation 2de6c780-9bdd-4dee-835c-7e2d8618b558 · outbound

This paper cites Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment.

Aligning LLMs with Domain Invariant Reward Models Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.372847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.372847Z digest=sha256:82ae139ad825b76df5f4e120f0141db031b7880675b95857afbeafc3b0f0367c

Observation ef303f7c-fb2a-4fb1-8fde-c11ed3a7cf1d · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Aligning LLMs with Domain Invariant Reward Models Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.378181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.378181Z digest=sha256:e57134f811fa6290e4c4d0cfd48e3c752abd366b17d5859bd952dfaa6b26d19c

Observation 15eaf064-48e0-4a6e-8d5e-c6c2f796e857 · outbound

This paper cites Chernova, and Dhruv Batra.

Aligning LLMs with Domain Invariant Reward Models Chernova, and Dhruv Batra

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.181683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.383414Z digest=sha256:58a7fa1ae42c7282b674e9263bf851175ef1c6df18f96b79840ebc34610fc159

Observation d9e7e8de-fa6f-4e8d-a26d-4cccde6dc901 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Aligning LLMs with Domain Invariant Reward Models Deep Domain Confusion: Maximizing for Domain Invariance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.388655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.388655Z digest=sha256:8ef20231c5307420687d698c460d9e18f955588bc92ed5b24ac23b695f26744d

Observation 9e561a12-24a5-4b02-b30d-94e811486f54 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.393710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.393710Z digest=sha256:a06dac61cd3610a0e1066aa2669a59dcf8ff8db9007c9c6d303b32c3b2fe3a7d

Observation 6adc45c0-0464-4e51-a507-fb1f30304374 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.161418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.398638Z digest=sha256:50e2f0b5c8ddb448e0f86770b951e0b750d8becf2d28afd3a41d200647eeb79c

Observation a05d9238-5b9f-42d9-b0e6-ff3ebf2d2d76 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.141030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.404542Z digest=sha256:f2a4eb29af0d4f830f8ee8e7b43284467a892d93fe01d953d0f3d5d9fe3319f8

Observation 1340b498-31a0-45e7-b283-1ec86018401b · outbound

This paper cites Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment.

Aligning LLMs with Domain Invariant Reward Models Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.593278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.409531Z digest=sha256:8f20b97432af5c15540457cca8ece74bb9a0ab0d99eac94c642affe02c1b7e40

Observation 97b509e2-18b3-48b6-9e74-8dfa8fc741ee · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

Aligning LLMs with Domain Invariant Reward Models CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.414979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.414979Z digest=sha256:3dbe96b0a71c79d355424324796ba515cee4ed2cd85d451d01ee494f58bf1e60

Observation 99232992-40d2-465d-94e1-27471d260e93 · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Aligning LLMs with Domain Invariant Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.420151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.420151Z digest=sha256:cb684e7d478b987e63b0c2fa90e83868325e3bb64bb6b009a7c28bcdc3d4d5c9

Observation e724335a-a9ae-42b5-9421-b8240cf682a8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Aligning LLMs with Domain Invariant Reward Models Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.425370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.425370Z digest=sha256:6e97413aaeb547717776d3e4d26357087f345b8539b07fd74a17959b2b42ea3e

Observation 0496f569-b2bf-4586-9b6e-075691c6365c · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.431407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.431407Z digest=sha256:77e33d9524242ff4680abc830fdd24a2fef9cffd031682c3961412b7f8fcaf69

Observation 4a406a80-2238-4f13-bd0a-8ba6f59e3246 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.123721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.436817Z digest=sha256:31f98050148ac6ab6528fba6575c194f069c04b8f67eb4a8c03fe723e9cd97fe

Observation f39061c4-0a81-46a2-bb04-17b0c58d1b40 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Aligning LLMs with Domain Invariant Reward Models Fine-Tuning Language Models from Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.441631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.441631Z digest=sha256:a7ac54cc32f972a221d3f84083c4b9d6afbf40b6a53b036f1ef5f36a0fcba6c1

Observation b527aeef-d4b6-463d-9df6-3949ecd8a411 · outbound

This paper cites online" 'onlinestring :=.

Aligning LLMs with Domain Invariant Reward Models online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.446752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.446752Z digest=sha256:103d7e9f8ad80c741234b3892d64a7d5c251deedeec60e8e896b85db6193557e

Observation 1adb606a-3623-4c4a-a5c1-dfb6f422aee9 · outbound

This paper cites write newline.

Aligning LLMs with Domain Invariant Reward Models write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.452310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.452310Z digest=sha256:575983c93f5cf5b2815d77b7c4558d8d50488a0ff4531a358fe73451fbde9c27

Pith citing papers

No inbound Pith citation observations are available.