Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from User Feedback

As of 11 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2505.14946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14946 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:31:19.984724Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:00:17.717041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T06:41:10.730142Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2ab7095-50fb-416c-87bc-9a347c2cacdb · outbound

This paper cites write newline.

Reinforcement Learning from User Feedback write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:18.136449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:18.136449Z digest=sha256:8a5dae8b786343dc8673c0d832565c361a23777b9b06b390e99a2d043062b159

Observation 8f263c6e-0cde-4bb6-b96b-90849af888b9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning from User Feedback Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:18.262530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:18.262530Z digest=sha256:591be1ea191db1f069924302d0b10b5ff9fd4b88b55364a935a65596f9d366e0

Observation 523153a0-7109-4a28-bf9b-bc5488b3d8cd · outbound

This paper cites Deep neural networks for youtube recommendations.

Reinforcement Learning from User Feedback Deep neural networks for youtube recommendations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:31:20.868378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:31:18.404379Z digest=sha256:168cfbfb698825cbd3d49b5a214350b0d69e2ce080d21d4ac6aebc5ef9187494

Observation 7052395b-27cb-4d97-a0d3-ede20d28d8d6 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reinforcement Learning from User Feedback Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:18.545522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:18.545522Z digest=sha256:f46671f1942448f117def50ce18b21acab4eb6229c0c99d4dbb0aa175c9aa65d

Observation 5bf7fe79-51d5-464a-9e17-d20e2c49382d · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning from User Feedback The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:18.652719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:18.652719Z digest=sha256:d473ee1c166e1db18e86e4027f4741f9ec5f1fe480100ac0c6c43ca9d46c5578

Observation 12575196-fd6d-46f4-a34b-dc1f0bcf13dd · outbound

This paper cites Monolith: Real Time Recommendation System With Collisionless Embedding Table.

Reinforcement Learning from User Feedback Monolith: Real Time Recommendation System With Collisionless Embedding Table

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:18.796406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:18.796406Z digest=sha256:7105aa45978e461e057b5b54010c3a29e484e014a1881c81d44e9836db0a23c2

Observation afc134b4-159d-400c-b6bf-a9a9957d135b · outbound

This paper cites Deep learning recommendation model for personalization and recommendation systems.

Reinforcement Learning from User Feedback Deep learning recommendation model for personalization and recommendation systems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:31:20.630150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:31:18.982092Z digest=sha256:16e6446491f4842eaf3a4b8ec6b0d1b859bcc2db5655fc66d0e8caf7afd5b7ed

Observation e0e223e9-45df-4c5b-87da-8dccd58dd865 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Learning from User Feedback GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.088890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.088890Z digest=sha256:6558ceebaf9eb9ab1e674105558b6fbe5115a01cf3e7a9ee099618419f1530db

Observation 187829dc-d5e9-4be6-8728-8a5decf356e1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning from User Feedback Training language models to follow instructions with human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.258698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.258698Z digest=sha256:ac54a77231e188e0682e956149a3738b505fc485acf193853a3ac16db499e38b

Observation 14511249-4d0a-4732-a8e2-58380b960387 · outbound

This paper cites Juicer: A benchmark for open domain dialogue evaluation with diverse negative responses.

Reinforcement Learning from User Feedback Juicer: A benchmark for open domain dialogue evaluation with diverse negative responses

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:31:20.378457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T15:31:19.387326Z digest=sha256:7671a2471e3148c99b28df09f8ebc6c5c96678be4162a9017e4f30f1b8d34aab

Observation 832ffd17-9217-4ad5-9f8d-1669bfd06395 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Learning from User Feedback Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.528623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.528623Z digest=sha256:bcfda0b1d66b95532b020ed4bbc618eb7324411dc7fd690b3458d127507388b5

Observation fd73abf3-90d8-4f4c-a164-675c57203be7 · outbound

This paper cites Improving Open Language Models by Learning from Organic Interactions.

Reinforcement Learning from User Feedback Improving Open Language Models by Learning from Organic Interactions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.651643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.651643Z digest=sha256:9c463f21fb7d48d8660972a8612a3935174f352bc880f593468da97efee79b65

Observation 75b57ebb-2c3e-4acd-9fe1-0523b5f07d79 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

Reinforcement Learning from User Feedback The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.744551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.744551Z digest=sha256:2c9b8dc0874a66b206a586fb8a5b167f0aaca9fce838eb1a67eb8d663a04d465

Observation 6e8cb7dd-c969-43e7-b52d-91bbe4171d6c · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Reinforcement Learning from User Feedback Secrets of RLHF in Large Language Models Part I: PPO

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.864968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.864968Z digest=sha256:18ebf5bfeccf40e4ed1b15203149ce833a989718d46882393e79f61e22b9e72a

Observation ba52866a-1355-43be-aa47-161221142ae4 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reinforcement Learning from User Feedback Fine-Tuning Language Models from Human Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:19.984724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:19.984724Z digest=sha256:c332f218200cfa6efda16c94ce424895800871596eed4962441a37b117e601ef

Pith citing papers

Observation d87a5206-3a8d-43b2-883a-5aa84860bd56 · inbound

Improve Large Language Model Systems with User Logs cites this paper.

Improve Large Language Model Systems with User Logs Reinforcement Learning from User Feedback

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:34:12.887504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T14:34:01.332088Z digest=sha256:7f6a3a49cb8aba6113b9965f463035d921e50014d4aa165aa00eba24f7c5ea55

Observation cc288a59-2e67-4c42-8f07-1a56c70ae9e6 · inbound

Improve Large Language Model Systems with User Logs cites this paper.

Improve Large Language Model Systems with User Logs Reinforcement Learning from User Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:00:17.717041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:00:17.717041Z digest=sha256:528e71019bce52a79277de18df3f92cbc796cd1ad60b89cf058388d6f4459488

Observation 53b82c74-fda3-4e75-b667-04a7ee5080d5 · inbound

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target cites this paper.

Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target Reinforcement Learning from User Feedback

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:18:21.443053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:13:49.666793Z digest=sha256:258833355ec570f7e31895c5051483aea98fcccbdbf2c29635c0c3be261b88ea

Observation be025eea-7c21-41ee-bd3e-2b19cfe6c51b · inbound

Echo: Learning from Experience Data via User-Driven Refinement cites this paper.

Echo: Learning from Experience Data via User-Driven Refinement Reinforcement Learning from User Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.732700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:37:34.840129Z digest=sha256:a51e2635cc8482d95fea321a0b4d387d1a0592cfbb15d0a207cec9264839dfd1