Pith. sign in

Paper Citation Record · LEDGER

HelpSteer2-Preference: Complementing Ratings with Preferences

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2410.01257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01257 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:15:46.125494Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.542430Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f408b60-18da-45ec-990d-743993458d17 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.666129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:01911a5fc9ddda887b966d89a73976192aa5f70428f73bb6236af15cdc717fed

Observation 899c1cb0-337c-4fad-abbf-406152101b87 · inbound

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning cites this paper.

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T14:15:46.125494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:15:46.125494Z digest=sha256:d0596325fce50683a4f978f670ecc0bf59ed96956aeb360a20dfaa45c8c38a6d

Observation af5f5635-b5f7-40f8-9a80-c7f2e860a756 · inbound

Expect the Unexpected: FailSafe Long Context QA for Finance cites this paper.

Expect the Unexpected: FailSafe Long Context QA for Finance HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:54:43.248027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:54:43.248027Z digest=sha256:7c83a700083154e104626886b5cf1bbe58fadb3f822e44ff56b6092bb6cd0b54

Observation 42666668-3a53-487f-97a0-02efe5af086b · inbound

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning cites this paper.

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:39.220697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:39.220697Z digest=sha256:4aeb3eece6106ac036c1c114fe9198ae51c8357b1235903f8a0646af40bdbb67

Observation 16781b5f-b038-4ecf-9483-0c3d591e4a18 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.659478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:47beb1d1d64f985a5e42baa3d4f2f87deb5fc56b756670d7530a36eb73c6738f

Observation 38ed5144-7e39-40df-b130-03924ff07c98 · inbound

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding cites this paper.

Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:59.874401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:59.874401Z digest=sha256:44cce8e3ad0b93774329f42ab739e98ff8620484a57f4f4edd7006e6ea22b04a

Observation a39d22f1-4ff5-4198-b742-58d1984e87cb · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.189073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.189073Z digest=sha256:f906d6b76fd710bb250d17795af2f2b1f1cace44f28dc367048471ab0af882eb

Observation 755f426c-4980-4b25-9170-c6be89aa72b4 · inbound

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems cites this paper.

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:22:15.929822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:20:12.827871Z digest=sha256:898cf07ec82ae4fe1220cf7b87a33780e6ea5be8b64ec4ad1551cb8f123f2bd1

Observation e9bba9b6-ce48-4fcb-83d2-d29a073e870d · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:19.287789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:19.287789Z digest=sha256:c8a451b66de3b77ad4a2ae8ad69f053bd715b67bbf722b63ece61d79ef1873f4

Observation 0941c192-6f1f-491f-82f9-a7d3820bddc5 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.787636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.787636Z digest=sha256:cb921d83a45538f853de24e187ed6b2f48657769c4d35812bcdc689940b40292

Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.619519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.619519Z digest=sha256:5e10c764b9e8aa470d850c03bd11cdfc82ec26ee99e36b505b7eaf8839246fae

Observation e57d647d-c45f-4c8d-8daf-a9bf780f2275 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.482448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.482448Z digest=sha256:b72b754e9d7d287f3828bb7991ba95511c0814681a68c76961501ba01e7e47ce

Observation 2a379593-88ca-4d08-bb53-4ebb4da394de · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.388592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.388592Z digest=sha256:5be136426b1cbe87f6806b181be6b0fff3626d0e009a20c61256e532c1c1b6a1

Observation b4039719-d289-4f68-abe4-e1b8e027db86 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.416543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.416543Z digest=sha256:b9c03ffd4ef6a656814cf642811e437958771c040c0e3c57742ac789b4441a2e

Observation 49e3ca38-f659-4cf0-8018-470ee97cc75a · inbound

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents cites this paper.

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:30.732674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:57:20.186356Z digest=sha256:799d8e38062d638fc60f749173b3bb4b204c8995cb58caffd3805c5a0690bd76

Observation 0d6e39b8-e452-49dc-b58d-4e4be105394e · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.543769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:02392e1809bde500c686b2939c7bacfeb62e332199c92a06a8157d99314166e4

Observation ce7f2b4a-8299-4273-a58c-2adf792cb670 · inbound

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis cites this paper.

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T10:30:46.777170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:30:46.777170Z digest=sha256:3488cae87b9298dba388e21f7910f976cd5b0881e8f35b67d77478003d8d5614