Pith. sign in

Paper Citation Record · LEDGER

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2509.03672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03672 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:13.607657Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:49:14.243931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:07:30.215801Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe8cde69-a40f-4ab0-9ee2-9c163d20786b · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences A General Language Assistant as a Laboratory for Alignment

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.716053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.716053Z digest=sha256:5048da6d96d62379a42be05b0b5a2b7cf23005b7e262bd849fa2c89fbf1a8397

Observation de8ce66d-7c79-458c-a04b-ec831e7f9bdb · outbound

This paper cites A Sharp Fannes-type Inequality for the von Neumann Entropy.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences A Sharp Fannes-type Inequality for the von Neumann Entropy

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:54:14.125679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:10.769694Z digest=sha256:dad2108b31ec277f1a0a4987e8b2d2f9129dd446ae7acbd24681f424ce908349

Observation 98653cfa-3c35-42ff-94b9-eda507bd7ace · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.832231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.832231Z digest=sha256:032c5d7772c51c50a8888aa3f64cf29b68fb060d1d1b20e07bf56642a815d093

Observation e687d5f9-1a48-4b56-acfc-308db9130456 · outbound

This paper cites Rank analysis of incomplete block designs: I.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.891723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.891723Z digest=sha256:ee205c38552a5cebe40274c5b5c8b520e0ef313613f99ef12e4cf9fead931481

Observation 97bb6238-1d2c-4042-b3e8-2ea3e4dddebe · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.974544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.974544Z digest=sha256:e5e7441dfcc3400f57f68b7473fcd860280cb7553069039f8db2001503c8ac9b

Observation d5241516-3b74-46d1-9351-296611749a36 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.055317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.055317Z digest=sha256:043d061d41135765f8e40fffc0847089c708c20286e05dceb473200c0a8070a7

Observation 1fb46bd4-3130-4a4c-8d0c-9357b37ad186 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.144244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.144244Z digest=sha256:bfbcf7a43bf4891fe8920d9c58fa2c1e70fd20b4a5655f343b624f3a3094d59e

Observation f4181020-b7af-400b-acc7-19cf002d6b28 · outbound

This paper cites PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.230380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.230380Z digest=sha256:03d138f6e3e25d934295ca965ed62b1538f78afd8ce7a5a7a3ee8a4d151854a6

Observation 60305593-17c6-4416-83df-aa8d0c54f42d · outbound

This paper cites The alignment problem: How can machines learn human values? Atlantic Books, 2021.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences The alignment problem: How can machines learn human values? Atlantic Books, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.339079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.305666Z digest=sha256:1318950ef741c9079b46f6d369efb21cab6cd7ed03710873a87c63f1cb5a986f

Observation 936afcaa-c91e-4cd2-89f3-fb19b297db57 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.358214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.358214Z digest=sha256:e6df98f6d3c4992b12a575070a448a084c4f2e1e6215c1310ca452e4ea05046b

Observation 41632b7a-2b4f-4813-bf37-6d0ee4a79b09 · outbound

This paper cites Active preference optimization for sample efficient rlhf.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Active preference optimization for sample efficient rlhf

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.424298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.424298Z digest=sha256:dddd4d7a84a228c9d43fc1d6e350b90f0d29f887442f73377a83821798ab4348

Observation 88ca93d6-092d-48ca-af65-522c391164de · outbound

This paper cites When personalization meets reality: A multi-faceted analysis of personalized preference learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences When personalization meets reality: A multi-faceted analysis of personalized preference learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.466795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.466795Z digest=sha256:012c5a640a785332c125e928620e83c2172b023a5566ad2427d45662460367a3

Observation 1accefff-7cf8-4e75-a659-1d2262aba6fc · outbound

This paper cites Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.555391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.555391Z digest=sha256:7a82c42bc6466407733873623e41ec38770f9dc97c12e9ce395d082e2c1c3002

Observation 678c3fde-147b-4386-a1d5-01868278dfe5 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.639883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.639883Z digest=sha256:ee546e81cc56df80d70670215e4d1f7cb4f1c8d6d82c047f3e4f1e9cbe494e47

Observation 4b01892f-6150-4dd1-aecd-dc5de55d79c3 · outbound

This paper cites Provably feedback-efficient reinforcement learning via active reward learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Provably feedback-efficient reinforcement learning via active reward learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.314729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.724453Z digest=sha256:45da1c64d7a524ed7fb05a429d5f19dfa39e88272e57a868c29d566826c7b58e

Observation 9e840360-9f7b-424c-835f-e8e8a47a2a0b · outbound

This paper cites Bandit algorithms.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Bandit algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.778287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.778287Z digest=sha256:60a05b36960da8f6b7ca1c59c1f67f02b1e39fc124a48338ce9e294e0c1e6edd

Observation 68bd6346-be17-4508-aee6-fc743a492faa · outbound

This paper cites Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.835948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.835948Z digest=sha256:7b0fbf0cc10adb773a70ea6280e58431eb7f9372c1c1677918726cdf70482647

Observation ddb95464-f354-4e29-985d-ce7d3544a8a2 · outbound

This paper cites Learning word vectors for sentiment analysis.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Learning word vectors for sentiment analysis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.288750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.925172Z digest=sha256:fae2f679ce572497224ac87e9b9af8aeb7ecb8715d80f664d216adbf23f1a62a

Observation 6de299e3-5109-4788-9121-e654899f0bed · outbound

This paper cites Training language models to follow instructions with human feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training language models to follow instructions with human feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.989754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.989754Z digest=sha256:a9233cccbef70b3dc994c683d79a27efa644dca988682c10bb868a1634cfb82c

Observation 20c57c66-534a-42b9-a12b-fb52ed52ceae · outbound

This paper cites Iterative reasoning preference optimization.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Iterative reasoning preference optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.079907Z digest=sha256:e046bec1c8306d32bd40ffddb94544083a43315395abbbe4f8a8f04ac2453a1b

Observation 170584b6-cd83-4933-b97a-d3ddcbacea7a · outbound

This paper cites Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.137540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.137540Z digest=sha256:94f73018dd1ee9a8d58af93fdb993742784a4aae1df896bfa0e3e8ea6472135c

Observation 0d3e55d7-8115-4545-83da-7e1eaffe14c7 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Direct preference optimization: Your language model is secretly a reward model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.237589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.237589Z digest=sha256:9cdc90dde561e93c5c06873f77f692f033fd54d57cf98462eba94900fb51c438

Observation ca89941e-85a1-443f-afb1-22990b46f596 · outbound

This paper cites Group robust preference optimization in reward-free rlhf.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Group robust preference optimization in reward-free rlhf

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.235770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.320047Z digest=sha256:6d05a6156579014dd065921a23c3d033d6346f8c76bce1ea3a09f1b7a2b79dd5

Observation 3b9ce4ec-c2af-4127-ae71-2169cf9c2400 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Dueling rl: Reinforcement learning with trajectory preferences

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.219234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.412731Z digest=sha256:755466796852dad1a27b6187aae10ea662730606f432f079885c0c9e5eef3be9

Observation ef35e711-f3b8-44ef-951b-25f36f72e3b4 · outbound

This paper cites Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.204135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.495223Z digest=sha256:0d4f111f92a22763a450593f5c5bbb76b20c0d649a805fde4324121556842394

Observation fa21e778-d9e4-4fdd-9cc0-0aafdad75656 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.534185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.534185Z digest=sha256:a34c5e9a0b9f1074703bb86c00914b8e34f6a42e481562ac5d3360b536b95f02

Observation 452cded4-fff2-4aa3-98aa-e5c967bb6dc6 · outbound

This paper cites Collective Choice and Social Welfare: An Expanded Edition.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Collective Choice and Social Welfare: An Expanded Edition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.626549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.626549Z digest=sha256:e05f8c7336d8c954b721323a2717fc53d87984366f4feb268df52fdf9b4d558d

Observation ab95073d-b7e5-4df8-95de-abd0a5a34da5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.794995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.794995Z digest=sha256:cd5d6cee2c3ddc6a473ae67cb959b75b25fecd697cf0032456830ffe246e1a35

Observation 7d07935c-e1dc-40c4-8720-48d0455f98e3 · outbound

This paper cites Robust multi-objective controlled decoding of large language models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Robust multi-objective controlled decoding of large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.873574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.873574Z digest=sha256:084bcb2f60ee22ee34ad36d61607f748e79b98266c9909880703c14a10a62b40

Observation 1e85c555-f779-4e4c-bcd1-16de7045dfa7 · outbound

This paper cites Learning to summarize from human feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Learning to summarize from human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.976902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.976902Z digest=sha256:ceb4b57cb3e12b89ad3d991ff22241ef9b17b17b9f2b4ca796577f9f52f5b8a9

Observation 57da1d23-91bb-4077-aafa-9c65c0328894 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Aligning Large Language Models with Human: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.130765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.130765Z digest=sha256:3c619dcb8ffefdf5a2bd2154da6696174ab05af0d397aa0bed2387bd3b2c22b2

Observation 456bd312-3c97-43e2-bcc5-7cff9de0513a · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.249178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.249178Z digest=sha256:fd3cd0d09200709814b095868019958c46fedd1b8ce077faf8835e8a911ccdaa

Observation 50b4bcdc-9e2e-4605-a69f-c186ffb3f573 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.188631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.373392Z digest=sha256:0a3c86532f5d54e7a635c7d3f896c051193a5eaefd8da60bb208f33a99798aa1

Observation 96b151ef-cf89-42e9-823b-18e4492e0c01 · outbound

This paper cites Impact of representation learning in linear bandits.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Impact of representation learning in linear bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.173019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.481709Z digest=sha256:b9feba90cd36a2324ef7c60eb2e216f975a4900ec54dc26753be0f5c205e794c

Observation 243ea628-e798-41c4-acb3-3fda90368e0b · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.592468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.592468Z digest=sha256:c3dc21ae13511c87d19a196c88d97f23d313fb3cf5052b95c0ff3f4ab4777573

Observation 77f38854-3e60-4eee-9a0c-ac68f7e0abe5 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.158270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.607657Z digest=sha256:38f10e4c48f4f18e2f0e9608f48056878363183aeb587fdd9e5f697e404b2ff1

Pith citing papers

Observation d01e6178-4e4e-4035-9ad6-0c8a4e3310de · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.217449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:f9f4f0f267d78dd4f70b49b720b6c07bf97e5f9be00e7d1c1e7589d0645f5a7e