Pith. sign in

Paper Citation Record · LEDGER

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback

As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.00816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00816 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:19:14.894582Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b7ac46a-b531-47ff-b5fc-8ca7ce0599d3 · outbound

This paper cites GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.537312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:10.824627Z digest=sha256:7b6a98611d853e2af30fef4ff8cc165af40341cd665f693e88220114637a85c9

Observation 1b20a2cc-3392-49c1-8446-6fcc17730ddd · outbound

This paper cites Proceedings of the 30th ACM SIGKDD conference on Knowledge Discovery and Data Mining , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 30th ACM SIGKDD conference on Knowledge Discovery and Data Mining , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.287995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:10.924810Z digest=sha256:ce941df4fa799f63d848b911f8f51594d03e2b7934761d723fe8e0f6f14fbf15

Observation 8456ce5d-4208-4bc5-b7a5-eff25ce1bb68 · outbound

This paper cites World Wide Web , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback World Wide Web , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.145515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.058092Z digest=sha256:6363e059c672d1edb34d98eb374f1a1f037bebb2f2e5b73a5744b75bdd073e5c

Observation b70799e8-04ac-4b39-8e3b-fc3eda2a73d4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.200723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.200723Z digest=sha256:7bb5c4e882de1eb6f9e73df5a67d7816b6a59affaea2cfdb12f5b0a6472f18e5

Observation 5cfc5d06-c236-4aaf-a350-5b0554f65af9 · outbound

This paper cites International Conference on Machine Learning (ICML) , year=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback International Conference on Machine Learning (ICML) , year=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.018679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.304724Z digest=sha256:a6099c1b351497d9b2d98392e88436d7fd8659d99e33d44f0ba3b882cd12cf75

Observation 084da148-7a8b-4d78-96e7-1832ff6a7c20 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages =.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 41st International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.832458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.426050Z digest=sha256:dd3f7c710bc360e7366bd888de26adfcb0d60ebb0096f0dbd3ba0b7c68664938

Observation b04b891e-61c5-4812-9284-43979c9be78a · outbound

This paper cites 2018 IEEE international conference on data mining (ICDM) , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 2018 IEEE international conference on data mining (ICDM) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.586804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.586804Z digest=sha256:44cad639ca207b3005d0d0f9e08d863ebabe637a4e29bdeb2e10d6c5434f9e2c

Observation 9a9774e2-2e22-4421-862f-ae355f691df6 · outbound

This paper cites OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.737576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.737576Z digest=sha256:95822981313dc989008ad9f1609015959d2f1c81580cab5ead7aa293ed005641

Observation 5ec8ec22-014b-4036-8fd6-911d5be069c8 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.844487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.844487Z digest=sha256:446e25d366f62dc44e4c2051aa02d1758650ef5a9055991d620339a62019cf58

Observation b0ca95dc-e35c-4856-98e3-c60acac5c8bf · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.956367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.956367Z digest=sha256:5399cce3043260ac6d5ab72de11bc3b4225e6a63ef2a0c10295464848456b72f

Observation b20ad193-e0e7-4071-a486-71c6e07e7445 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.070497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.070497Z digest=sha256:ca6068d6d13dc4e5d9b9ecc53fc7710146fab07ca2cbfde3cbe36a91a3c42c6c

Observation 528aa86a-2d76-45d3-bf5a-4dad349ed9c8 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.211566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.211566Z digest=sha256:4f42e3e7c66f48dc08baf018cbcea8cd56d9de87a9c3e8db5d8e0fc99a6bfc29

Observation 9e6dcee5-181b-44d7-88a8-7d6421722f5a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.292009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.292009Z digest=sha256:1bccd108ad13d440e76c2118ea2c69a87a29d54b3699881b8bbda82573117478

Observation 64151d39-3180-47d6-930c-a763b193f0c0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.295874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.295874Z digest=sha256:b1a22c91d68c9973691e8ed371e62b047e7c2a45ea6157e77b790a438efb066a

Observation ba916abc-18dc-4c4f-b922-4a3fdd784c66 · outbound

This paper cites Local Policy Improvement for Recommender Systems.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Local Policy Improvement for Recommender Systems

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.338670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.368388Z digest=sha256:763412271bb149ae9d8ff2718317d5af5e73faad3d4d3f76f913c68167e3136d

Observation 5f2c86f5-9b0a-4b96-b9af-bfaf21d7ebbd · outbound

This paper cites Proceedings of the 10th ACM conference on recommender systems , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 10th ACM conference on recommender systems , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.466083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.466083Z digest=sha256:630c51926de48f4a5fd1acc89bb53c52cd1611d372130e1058e46df4f3a6dac2

Observation dc58d33e-6c97-4e4b-a486-51e9b9888aab · outbound

This paper cites ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.168696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.591937Z digest=sha256:969179a14fc5f578f5ac1d6d1c1b62886874642317e64760219aca923d21556a

Observation 17e6eb35-c945-4ad1-89d1-579f4e12468c · outbound

This paper cites The Journal of Machine Learning Research , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Journal of Machine Learning Research , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.683494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.679862Z digest=sha256:e8ffbaa8132001f2a1e6957757daa8efab0176c6aade47bfba7f61051e4f5993

Observation 6991502d-713a-42ba-b217-4212fcdcc6ed · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Doubly Robust Policy Evaluation and Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.758286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.758286Z digest=sha256:547c0382ace00d1975936128a630bf7327967b84bba83546238339f8341cc6d7

Observation 4f387d92-364e-41f3-a1e8-1470086c1582 · outbound

This paper cites 1998 , publisher=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 1998 , publisher=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.848680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.848680Z digest=sha256:30890877508bc05a61bb4c5091491ab629e5f0fa9e2bb655d43b2bcf50c21617

Observation 198f9939-aeb8-4fb9-b68a-a5813b2e7b4b · outbound

This paper cites Deep Reinforcement Learning in Large Discrete Action Spaces.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Deep Reinforcement Learning in Large Discrete Action Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.929381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.929381Z digest=sha256:eeb7b02181795c0771387ce18b4fe913133cddac5b6618cf3f5913fb20efe4e6

Observation 44e3588a-c9d6-446d-8267-5959ef8959d5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.460164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.010008Z digest=sha256:ff1198853c536bab42c333b4fed7c0ac22912de608df06aef657c73e69dee6e2

Observation 44ed47ad-1fc1-4067-9902-6b0d347dee88 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.088764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.088764Z digest=sha256:425437349020f50c66028701c0fa973d6b34bd518037af9089227a42f3122f28

Observation 85636644-4893-436f-8d69-8016335f7797 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.185095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.185095Z digest=sha256:43e641c0a6a3c309fd29126e5d4e5c342d93b4953b37f59ebe6429bcf52144ae

Observation 616e8515-de3d-4d1c-b101-39b9d14b4924 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.307778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.307778Z digest=sha256:d34f9921297b51f8fd393122cda20cf906f43a5acc248fd93b5027b200aac21e

Observation 4d122eb5-4151-4b83-9c26-315927ce4600 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback International Conference on Artificial Intelligence and Statistics , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.406677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.406677Z digest=sha256:2d17f7605c04c2798ffee9be181b259d8a1d85852e8acddc344513a9b3e27a19

Observation 74d43e1a-d897-41df-8e2b-2c61583264dd · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.504276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.504276Z digest=sha256:2fdfdbabfd8677c7ca6bf04df200fc56a910b5575b12ad574cce2a7bfe0cdc69

Observation f4729e51-fd38-4421-92ca-9cad72d06ebc · outbound

This paper cites Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.584966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.584966Z digest=sha256:88bb48021607a441ed3711f66df5c61079991086ad5f209f3fc47c205d167eec

Observation fdba5dbc-4326-47d6-a6b0-5bebe7ac1b64 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.674050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.674050Z digest=sha256:3a412c401906baa1705f547222b8dc0b3369e232c4b9fae2c45925ee51fe6049

Observation 95bd82d2-ea9b-42cc-9b35-27aa58f7ffe5 · outbound

This paper cites Acm transactions on interactive intelligent systems (tiis) , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Acm transactions on interactive intelligent systems (tiis) , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.787238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.787238Z digest=sha256:2f3656b7af5ea5115aad4f687ce1a1881eefad17a98c4605a93d0d6f4c2d7200

Observation c8fe50c8-b796-4a01-98bf-6f0a983112bb · outbound

This paper cites Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.257418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.891636Z digest=sha256:6fe2bfec0cc864f972df920af7aedd132a4092256ff37c2c0194d4c4d7185bf7

Observation 71be91dd-66ea-499d-9c11-fd1093792839 · outbound

This paper cites Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining , pages=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.114661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.988717Z digest=sha256:e28c8652eab1233f450e6f5a51066da7db1b7fdca34f34551380d74d59ad715a

Observation 10a7a1bd-3bbb-48ac-a4c4-c44a9b1bad0b · outbound

This paper cites international conference on machine learning , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback international conference on machine learning , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.035733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.035733Z digest=sha256:e99aa8714d7f328f5aeebfa1127ac5fe247b7b53adb24094b913e68dc13721ea

Observation f922640e-028d-42a1-b6ec-6e3cb44bfbd1 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Offline Reinforcement Learning with Implicit Q-Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.216356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.216356Z digest=sha256:3218095e052c023103f998650941101bb8c686b199f582d084147b4ccd885909

Observation 608e2981-d26b-4b75-9a27-e1d918845277 · outbound

This paper cites 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:15.884769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:14.300165Z digest=sha256:91bb008ec193d27013748dd3cfa6fd5a5a12b0556ae65acb89884c252a11d0a7

Observation 9b7f5ebf-58c7-4f58-8ece-ecd70ef2026e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.370525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.370525Z digest=sha256:b8c83021fa738a295f3eedea94c6b8e5ed9dad344117e2bf889090edc7683f8f

Observation 8282fac1-639a-4950-b16e-894fea6c321c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:15.703841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T00:19:14.429470Z digest=sha256:68b29cb7e6d92973955da30256c89ca30143246dc37113493ddb350a7db22d0e

Observation 82d42bf6-2cca-4297-ae8f-a32229b733be · outbound

This paper cites The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.540226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.540226Z digest=sha256:1b85a3dafb79c7bd677a773cd649c9e5714514d7ebb21aacdc06d4f2e3536800

Observation 686cfe5c-6d1e-450a-bbea-0f47829103fe · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.655825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.655825Z digest=sha256:bd6b4a93148f8f10b2438a44960a11707a027e58ce7b4233739b8126ba965d27

Observation 038b2e80-d69c-482d-9f6b-104954b98d58 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.736605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.736605Z digest=sha256:06cc9d55b66b4524e81aa3db365b7d63ac8d2e8768bf9765ba336a794c2e21de

Observation 9e2803f2-28f1-4b10-8431-2b1b00e6c358 · outbound

This paper cites Machine learning , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Machine learning , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.802419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.802419Z digest=sha256:f871bfdf116ca9260b070335fe8cb769ec52e9c4d2245ddc22d00650632d0ab9

Observation 55dc147b-2d7b-4e08-a9d3-90bcf86d6fe8 · outbound

This paper cites Concrete Problems in AI Safety.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Concrete Problems in AI Safety

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.894582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.894582Z digest=sha256:d8dc99123298770528152a72c86eaa4c5d8896c0357ce9ccb2fe9b0772b8776a

Pith citing papers

No inbound Pith citation observations are available.