Specifically, we generate 16 re- sponses (4 for models with 32k context) per ques- tion using a temperature of 0.6 and a top-p value of 0.95

· 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning

cs.CL · 2025-12-17 · unverdicted · novelty 7.0

SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.

citing papers explorer

Showing 1 of 1 citing paper.

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning cs.CL · 2025-12-17 · unverdicted · none · ref 4
SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.

Specifically, we generate 16 re- sponses (4 for models with 32k context) per ques- tion using a temperature of 0.6 and a top-p value of 0.95

fields

years

verdicts

representative citing papers

citing papers explorer