Pith. sign in

Paper Citation Record · LEDGER

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2602.10623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10623 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:06:25.403043Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T14:00:28.579853Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45c0ea56-6d84-4151-aecb-afbb4342e5f5 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:24.938679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.938679Z digest=sha256:c234bd517ed43f542722bb113a08721ee8afeca85f3787d2e401bc3dad2042a5

Observation 8cfc5805-e433-4df9-8115-680fb2a9cae3 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:24.979514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.979514Z digest=sha256:26aae05779f664438aca0f0facc3cabcb556749a8fe996855dc3845a4ec2f5db

Observation 678e5c54-e360-4cfa-aafe-ef1fc4ac4688 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.010890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.010890Z digest=sha256:8d5a5b633f4d993fc196d72e061c6c30c0e22d5c2aac033d25fcf9d105cdb26e

Observation 48de6a4e-50b7-4075-b9db-d0fc5195d55e · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.050054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.050054Z digest=sha256:7bedf88d6babd84f6df8323a00af3e7865f49e2f2d42bc53716ae64820cf3058

Observation e58d9cf4-c98f-48da-88ea-d744c0f00c25 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.071900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.071900Z digest=sha256:358ff52c02eaba7e9b12d7efcb63025e7a836e037b9355693a83328463e15597

Observation a25d2624-1d4a-410a-a32d-ae2982f7173b · outbound

This paper cites 3https://github.com/vllm-project/vllm.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling 3https://github.com/vllm-project/vllm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.123714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.123714Z digest=sha256:9c5a91dc6fbd572aac92ec0f6441d198b529be109ae38de50debbc1f138825d7

Observation 95fd40e9-4e6c-4d89-9b91-9b5dec5f36b8 · outbound

This paper cites good” or “bad.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling good” or “bad

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.180869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.180869Z digest=sha256:ff02ebb63c9a2a363a1db0feff7f35ae540aa53fefe4bb7ea48c1684dc3de194

Observation bb7f70ff-e993-482c-96a7-15a81b7062cc · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.240241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.240241Z digest=sha256:96006ca20bc0d4902f21c9755e8ff089798a2d08fb3c4d9ebf6adbf98b406a72

Observation 1298910b-0d39-499c-a0a1-fbd7433687e1 · outbound

This paper cites Justify your judgement using the high-activation examples above.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Justify your judgement using the high-activation examples above

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.268589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.268589Z digest=sha256:5852090c0cbaccd6159e5717938cd335ff688637038109dda992140463f4cd99

Observation 1fc126a7-d615-46b2-a29a-51bd8291bc25 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.322569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.322569Z digest=sha256:c538e40b4bb47edf514913600b67f337014f6b71e2cb31c1ce00b8c0a34dea2a

Observation 33dabb3f-934b-44cc-b3c8-1e27f6bb26c7 · outbound

This paper cites FactorName.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling FactorName

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:06:25.349105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.349105Z digest=sha256:8e757ebe070c36c41363b35f4c4243e892f72b2cb629e63b2e8b731e936bc7cf

Observation 22b17670-b1d9-45f2-a47f-f8bb9827c25c · outbound

This paper cites Sorry, but I can’t assist with that.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Sorry, but I can’t assist with that

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:25.403043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:25.403043Z digest=sha256:c938bf586e17439f0355fd6a4ac7b7bc585ab464ece1f5ea62b5eb59b1812c67

Observation aa981351-cdd1-4906-8b72-001ca6833084 · outbound

This paper cites Spurious Feature Diversification Improves Out-of-distribution Generalization.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling Spurious Feature Diversification Improves Out-of-distribution Generalization

Reference 281

Resolution
unresolved
no resolver link, observed 2026-08-03T01:06:24.853310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.853310Z digest=sha256:d674d082081a9991734261ade5678c30d9bad29a1da8bed11b92f3eb343f786a

Observation 4182d5ab-7695-4523-805c-a384433e185d · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling A Long Way to Go: Investigating Length Correlations in RLHF

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:06:24.896678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:06:24.896678Z digest=sha256:f0e3a13f539a3de2dcbb20b692b7607570542312208f0466b47e421638793d6a

Pith citing papers

Observation 733f837e-17c2-4e01-bfe8-852524d4de8f · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:38.582535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:cf37becd03f636fa9b862bfd170474ffa45644f9402516e1d81f4fe838f06f95

Observation ebaac4fd-fbd1-4d1a-b479-e99b2ee84248 · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:38.582535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:8ca5fdc0949620bd06d02c1a4dc97e1d869b8f122e3d16ae06fded1d3123a471

Observation 81dcee44-4941-42a0-b217-466f5710df25 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:746004f40e941daee52838e06a7cec5dda45c679b7c5f8e9eaad4911d35a9f32