Pith. sign in

Paper Citation Record · LEDGER

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.10216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10216 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:32:47.969450Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5786330-897c-49a8-9ffd-bc257326029b · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.683896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:d6bce537bfd1f9e1feaca9a8ced9ca9fb7ab91d75da9cea034ca0c081ac8c157

Observation 447d25aa-7459-4144-8226-1b23bbab127d · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.512693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:f7d293b5dc13fbb3ad88e8990a3c8ea43446c164eebee311f5e9755aca55de7f

Observation 092df6bd-e88b-4aae-885b-4a2b3ff2ef0a · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.969450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.969450Z digest=sha256:df717d914ffbd9c8729ce4f400ce5402293018d178010fb4e4951e9eec1bd4f2

Observation c2be624d-9675-487d-92de-65d134ad6429 · inbound

ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning cites this paper.

ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:55:53.556785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:55:53.556785Z digest=sha256:27365be4a87cdfeb93a793707908270a8049e72b4f980b72bb5f42096d3b378f

Observation 17375a3c-39b5-4c5f-93b5-1113f5dbca4a · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.851632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.851632Z digest=sha256:7fb5816058e44007acc03c209830ae4e21717abdb582daf1789cf526ce664695

Observation b248e687-6a40-4976-a20b-c4b0a2463933 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.269350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.269350Z digest=sha256:9ba0e885cee8071966865614d9d97b7223bb30bfca077b63fa3e00b2d8fe010e

Observation ca0fdb99-e30d-4830-9981-4e1aa7a27de4 · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.397981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.397981Z digest=sha256:eda9c940e266f911dfb1a1662e628d310da7a6af8007198ff2cece9830944d5f

Observation 73a94515-8b5d-404a-a12a-94383c9bb9c5 · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.795751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:e2583ccb8eb2c0d72814199a3a99c473328961e2410088a4efa2acd10cc48b6f

Observation a77023ae-d679-40fd-88d7-c47287cfdbbd · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.948767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:1c72bdfe9f371ee65c88f99c7643b7b246927ce069bb5f0743461ea843eda8ad

Observation 49586af3-8237-4e3c-805f-e53f310e390b · inbound

Addressing Over-Refusal in LLMs with Competing Rewards cites this paper.

Addressing Over-Refusal in LLMs with Competing Rewards Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.598816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:59:12.695984Z digest=sha256:d6d41ca5288b75738b3a212db138c889138f8c2bb94c356af56645a4262fe26c