Test-time scaling helps only certain medical AI models and only on hard questions, with parallel sampling best for short-reasoning models and sequential revision best for deep-reasoning models.
The Llama 3 herd of models,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
Test-time scaling helps only certain medical AI models and only on hard questions, with parallel sampling best for short-reasoning models and sequential revision best for deep-reasoning models.