Pith. sign in

Paper Citation Record · LEDGER

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.00816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00816 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:19:14.894582Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b7ac46a-b531-47ff-b5fc-8ca7ce0599d3 · outbound

This paper cites GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.537312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:10.824627Z digest=sha256:f9aa640b3a1cd1a78fd6a14663301140387f42dc521c6526a78ffc081b019f50

Observation 1b20a2cc-3392-49c1-8446-6fcc17730ddd · outbound

This paper cites Proceedings of the 30th ACM SIGKDD conference on Knowledge Discovery and Data Mining , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 30th ACM SIGKDD conference on Knowledge Discovery and Data Mining , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.287995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:10.924810Z digest=sha256:17baf69a875c3d2116a9cb1f0c1cf5ef4a4085c8cfd9dcbca2c59e09b193c2be

Observation 8456ce5d-4208-4bc5-b7a5-eff25ce1bb68 · outbound

This paper cites World Wide Web , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback World Wide Web , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.145515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.058092Z digest=sha256:6af186cf60a99aaecdf5b2a18ec023fbc1f2feb77f0d6c3fef168ea533b7fa8b

Observation b70799e8-04ac-4b39-8e3b-fc3eda2a73d4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.200723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.200723Z digest=sha256:7bb5c4e882de1eb6f9e73df5a67d7816b6a59affaea2cfdb12f5b0a6472f18e5

Observation 5cfc5d06-c236-4aaf-a350-5b0554f65af9 · outbound

This paper cites International Conference on Machine Learning (ICML) , year=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback International Conference on Machine Learning (ICML) , year=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:17.018679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.304724Z digest=sha256:b1f3aec2d02b1413c03d9a25fa213cd49f0d0745890be6e06c01d6693375d47b

Observation 084da148-7a8b-4d78-96e7-1832ff6a7c20 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages =.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 41st International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.832458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:11.426050Z digest=sha256:39253e26502f6b510774e1f4352e7ae9979242ed3df361e41310c60a3fdc79a9

Observation b04b891e-61c5-4812-9284-43979c9be78a · outbound

This paper cites 2018 IEEE international conference on data mining (ICDM) , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 2018 IEEE international conference on data mining (ICDM) , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.586804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.586804Z digest=sha256:44cad639ca207b3005d0d0f9e08d863ebabe637a4e29bdeb2e10d6c5434f9e2c

Observation 9a9774e2-2e22-4421-862f-ae355f691df6 · outbound

This paper cites OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.737576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.737576Z digest=sha256:95822981313dc989008ad9f1609015959d2f1c81580cab5ead7aa293ed005641

Observation 5ec8ec22-014b-4036-8fd6-911d5be069c8 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.844487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.844487Z digest=sha256:446e25d366f62dc44e4c2051aa02d1758650ef5a9055991d620339a62019cf58

Observation b0ca95dc-e35c-4856-98e3-c60acac5c8bf · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:11.956367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:11.956367Z digest=sha256:5399cce3043260ac6d5ab72de11bc3b4225e6a63ef2a0c10295464848456b72f

Observation b20ad193-e0e7-4071-a486-71c6e07e7445 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.070497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.070497Z digest=sha256:ca6068d6d13dc4e5d9b9ecc53fc7710146fab07ca2cbfde3cbe36a91a3c42c6c

Observation 528aa86a-2d76-45d3-bf5a-4dad349ed9c8 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.211566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.211566Z digest=sha256:4f42e3e7c66f48dc08baf018cbcea8cd56d9de87a9c3e8db5d8e0fc99a6bfc29

Observation 9e6dcee5-181b-44d7-88a8-7d6421722f5a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.292009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.292009Z digest=sha256:1bccd108ad13d440e76c2118ea2c69a87a29d54b3699881b8bbda82573117478

Observation 64151d39-3180-47d6-930c-a763b193f0c0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.295874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.295874Z digest=sha256:b1a22c91d68c9973691e8ed371e62b047e7c2a45ea6157e77b790a438efb066a

Observation ba916abc-18dc-4c4f-b922-4a3fdd784c66 · outbound

This paper cites Local Policy Improvement for Recommender Systems.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Local Policy Improvement for Recommender Systems

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.338670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.368388Z digest=sha256:d302bc344be63abe5a1547b60247a69d36363198f5a9517a663b4de9d46bec07

Observation 5f2c86f5-9b0a-4b96-b9af-bfaf21d7ebbd · outbound

This paper cites Proceedings of the 10th ACM conference on recommender systems , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 10th ACM conference on recommender systems , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.466083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.466083Z digest=sha256:630c51926de48f4a5fd1acc89bb53c52cd1611d372130e1058e46df4f3a6dac2

Observation dc58d33e-6c97-4e4b-a486-51e9b9888aab · outbound

This paper cites ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:19:15.168696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.591937Z digest=sha256:2426a2ce9a0a24536452ff9ddb1980f671813420cb001b9fef241f3b32ec44dd

Observation 17e6eb35-c945-4ad1-89d1-579f4e12468c · outbound

This paper cites The Journal of Machine Learning Research , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Journal of Machine Learning Research , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.683494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:12.679862Z digest=sha256:375bc89de64e2eae3ef27dbec4e65f05d9549bc17b3497865b9c647196690185

Observation 6991502d-713a-42ba-b217-4212fcdcc6ed · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Doubly Robust Policy Evaluation and Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.758286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.758286Z digest=sha256:547c0382ace00d1975936128a630bf7327967b84bba83546238339f8341cc6d7

Observation 4f387d92-364e-41f3-a1e8-1470086c1582 · outbound

This paper cites 1998 , publisher=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 1998 , publisher=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.848680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.848680Z digest=sha256:30890877508bc05a61bb4c5091491ab629e5f0fa9e2bb655d43b2bcf50c21617

Observation 198f9939-aeb8-4fb9-b68a-a5813b2e7b4b · outbound

This paper cites Deep Reinforcement Learning in Large Discrete Action Spaces.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Deep Reinforcement Learning in Large Discrete Action Spaces

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:12.929381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:12.929381Z digest=sha256:eeb7b02181795c0771387ce18b4fe913133cddac5b6618cf3f5913fb20efe4e6

Observation 44e3588a-c9d6-446d-8267-5959ef8959d5 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.460164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.010008Z digest=sha256:61dd956026f4b2c522743647252dfc2a4ad67415926ea9d7c6fc2d4cd88080f6

Observation 44ed47ad-1fc1-4067-9902-6b0d347dee88 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.088764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.088764Z digest=sha256:425437349020f50c66028701c0fa973d6b34bd518037af9089227a42f3122f28

Observation 85636644-4893-436f-8d69-8016335f7797 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.185095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.185095Z digest=sha256:43e641c0a6a3c309fd29126e5d4e5c342d93b4953b37f59ebe6429bcf52144ae

Observation 616e8515-de3d-4d1c-b101-39b9d14b4924 · outbound

This paper cites Advances in neural information processing systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in neural information processing systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.307778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.307778Z digest=sha256:d34f9921297b51f8fd393122cda20cf906f43a5acc248fd93b5027b200aac21e

Observation 4d122eb5-4151-4b83-9c26-315927ce4600 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback International Conference on Artificial Intelligence and Statistics , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.406677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.406677Z digest=sha256:2d17f7605c04c2798ffee9be181b259d8a1d85852e8acddc344513a9b3e27a19

Observation 74d43e1a-d897-41df-8e2b-2c61583264dd · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.504276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.504276Z digest=sha256:2fdfdbabfd8677c7ca6bf04df200fc56a910b5575b12ad574cce2a7bfe0cdc69

Observation f4729e51-fd38-4421-92ca-9cad72d06ebc · outbound

This paper cites Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.584966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.584966Z digest=sha256:88bb48021607a441ed3711f66df5c61079991086ad5f209f3fc47c205d167eec

Observation fdba5dbc-4326-47d6-a6b0-5bebe7ac1b64 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.674050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.674050Z digest=sha256:3a412c401906baa1705f547222b8dc0b3369e232c4b9fae2c45925ee51fe6049

Observation 95bd82d2-ea9b-42cc-9b35-27aa58f7ffe5 · outbound

This paper cites Acm transactions on interactive intelligent systems (tiis) , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Acm transactions on interactive intelligent systems (tiis) , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:13.787238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:13.787238Z digest=sha256:2f3656b7af5ea5115aad4f687ce1a1881eefad17a98c4605a93d0d6f4c2d7200

Observation c8fe50c8-b796-4a01-98bf-6f0a983112bb · outbound

This paper cites Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.257418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.891636Z digest=sha256:639e461eda6c9777831c3d3ee61c628bb563e9fb9b307d6c7d5fe3fbe81248f3

Observation 71be91dd-66ea-499d-9c11-fd1093792839 · outbound

This paper cites Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining , pages=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:16.114661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:13.988717Z digest=sha256:d891baa5b55925fa9fa72c04f6eb5e4997b9a75db526ae0f3cec1bfea16bb13f

Observation 10a7a1bd-3bbb-48ac-a4c4-c44a9b1bad0b · outbound

This paper cites international conference on machine learning , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback international conference on machine learning , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.035733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.035733Z digest=sha256:e99aa8714d7f328f5aeebfa1127ac5fe247b7b53adb24094b913e68dc13721ea

Observation f922640e-028d-42a1-b6ec-6e3cb44bfbd1 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Offline Reinforcement Learning with Implicit Q-Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.216356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.216356Z digest=sha256:3218095e052c023103f998650941101bb8c686b199f582d084147b4ccd885909

Observation 608e2981-d26b-4b75-9a27-e1d918845277 · outbound

This paper cites 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning , pages=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:15.884769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:14.300165Z digest=sha256:5359efb3fdd27168e7e2779ecc5d7fbdff793b0956067379f82a9e567b60e33c

Observation 9b7f5ebf-58c7-4f58-8ece-ecd70ef2026e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.370525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.370525Z digest=sha256:b8c83021fa738a295f3eedea94c6b8e5ed9dad344117e2bf889090edc7683f8f

Observation 8282fac1-639a-4950-b16e-894fea6c321c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:19:15.703841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T00:19:14.429470Z digest=sha256:f072a4c05b5a199da7c151b28915b681924afe18e09fee9526d9ca86b14b6028

Observation 82d42bf6-2cca-4297-ae8f-a32229b733be · outbound

This paper cites The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.540226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.540226Z digest=sha256:1b85a3dafb79c7bd677a773cd649c9e5714514d7ebb21aacdc06d4f2e3536800

Observation 686cfe5c-6d1e-450a-bbea-0f47829103fe · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.655825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.655825Z digest=sha256:bd6b4a93148f8f10b2438a44960a11707a027e58ce7b4233739b8126ba965d27

Observation 038b2e80-d69c-482d-9f6b-104954b98d58 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.736605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.736605Z digest=sha256:06cc9d55b66b4524e81aa3db365b7d63ac8d2e8768bf9765ba336a794c2e21de

Observation 9e2803f2-28f1-4b10-8431-2b1b00e6c358 · outbound

This paper cites Machine learning , volume=.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Machine learning , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.802419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.802419Z digest=sha256:f871bfdf116ca9260b070335fe8cb769ec52e9c4d2245ddc22d00650632d0ab9

Observation 55dc147b-2d7b-4e08-a9d3-90bcf86d6fe8 · outbound

This paper cites Concrete Problems in AI Safety.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Concrete Problems in AI Safety

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.894582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.894582Z digest=sha256:d8dc99123298770528152a72c86eaa4c5d8896c0357ce9ccb2fe9b0772b8776a

Pith citing papers

No inbound Pith citation observations are available.