SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.
• When the speed is s + 2 km/h, the walking time is 9 s+2 × 60 − t = 144 − t minutes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
SCOPE uses step-wise confidence and dynamic subgroups to create finer pseudo-labels in test-time RL, delivering 13.1% relative gains on AIME 2025 over majority-voting baselines.