Structuring a reasoning model's chain of thought into multi-turn answer steps and optimizing with reinforcement learning reduces token usage and latency by up to roughly 70 percent with only small accuracy losses.
Training large language model to reason in a continuous latent space, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition
Structuring a reasoning model's chain of thought into multi-turn answer steps and optimizing with reinforcement learning reduces token usage and latency by up to roughly 70 percent with only small accuracy losses.