REVIEW 2 cited by
Dyve: Thinking Fast and Slow for Dynamic Process Verification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present Dyve, a dynamic process verifier that enhances reasoning error detection in large language models by integrating fast and slow thinking, inspired by Kahneman's Systems Theory. Dyve adaptively applies immediate token-level confirmation System 1 for straightforward steps and comprehensive analysis System 2 for complex ones. Leveraging a novel step-wise consensus-filtered process supervision technique, combining Monte Carlo estimation with LLM based evaluation, Dyve curates high-quality supervision signals from noisy data. Experimental results on ProcessBench and the MATH dataset confirm that Dyve significantly outperforms existing process-based verifiers and boosts performance in Best-of-N settings.
Forward citations
Cited by 2 Pith papers
-
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
FlexiVe, a GRPO-trained generative verifier with fast and slow modes, plus an early-detection pipeline, improves AIME math accuracy while reducing tokens versus self-consistency.
-
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
A two-model collaborative reasoning framework (Long⊗Short) scores reasoning thoughts by rollout accuracy, trains one model for important thoughts and one for the rest, and reports over 80% token savings with small acc...
Discussion (0). Continue with ORCID to comment.