Pith. sign in

Paper Citation Record · LEDGER

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.10216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10216 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:55:53.556785Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5786330-897c-49a8-9ffd-bc257326029b · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.683896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:4b1570f76f18a0f284c3b7ec1f25f7fb4cb764ef05c0d265333d2e19c0a5cfda

Observation 447d25aa-7459-4144-8226-1b23bbab127d · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.512693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:b610c2e2db6e2d59076eaf106186b3aa04fe597e64e8733d595e6dcf1fbbfdc1

Observation c2be624d-9675-487d-92de-65d134ad6429 · inbound

ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning cites this paper.

ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:55:53.556785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:55:53.556785Z digest=sha256:b3a788f29ce6a16a6f8a8af5a4f374a2840a577285b4097cc8488d1cbce6c1dc

Observation 17375a3c-39b5-4c5f-93b5-1113f5dbca4a · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.851632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.851632Z digest=sha256:7fb5816058e44007acc03c209830ae4e21717abdb582daf1789cf526ce664695

Observation b248e687-6a40-4976-a20b-c4b0a2463933 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.269350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.269350Z digest=sha256:9ba0e885cee8071966865614d9d97b7223bb30bfca077b63fa3e00b2d8fe010e

Observation ca0fdb99-e30d-4830-9981-4e1aa7a27de4 · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.397981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.397981Z digest=sha256:eda9c940e266f911dfb1a1662e628d310da7a6af8007198ff2cece9830944d5f

Observation 73a94515-8b5d-404a-a12a-94383c9bb9c5 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.795751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:b0fea2b8b4fad49518306423b14478e64b002c9b87a5e466472dd0221b440154

Observation a77023ae-d679-40fd-88d7-c47287cfdbbd · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.948767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:26ee4e952afd9816d21e84b27a93eb8381a562a77e5324872bb48a9152ae9729

Observation 49586af3-8237-4e3c-805f-e53f310e390b · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.598816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:8f5c60184c0a3940ee0eee64d0d6139ee0e2c8e3ea65732918d7db48adde8524