pith. machine review for the scientific record. sign in

Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems , pages=

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

stat.ML 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

Variance-aware Reward Modeling with Anchor Guidance

stat.ML · 2026-05-12 · unverdicted · novelty 7.0

Anchor-guided variance-aware reward modeling uses two response-level anchors to resolve non-identifiability in Gaussian models of pluralistic preferences, yielding provable identification, a joint training objective, and improved RLHF performance.

citing papers explorer

Showing 1 of 1 citing paper.

  • Variance-aware Reward Modeling with Anchor Guidance stat.ML · 2026-05-12 · unverdicted · none · ref 43

    Anchor-guided variance-aware reward modeling uses two response-level anchors to resolve non-identifiability in Gaussian models of pluralistic preferences, yielding provable identification, a joint training objective, and improved RLHF performance.