REVIEW 2 cited by
SEM: Reinforcement Learning for Search-Efficient Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in Large Language Models(LLMs) have demonstrated their capabilities not only in reasoning but also in invoking external tools, particularly search engines. However, teaching models to discern when to invoke search and when to rely on their internal knowledge remains a significant challenge. Existing reinforcement learning approaches often lead to redundant search behaviors, resulting in inefficiencies and over-cost. In this paper, we propose SEM, a novel post-training reinforcement learning framework that explicitly trains LLMs to optimize search usage. By constructing a balanced dataset combining MuSiQue and MMLU, we create scenarios where the model must learn to distinguish between questions it can answer directly and those requiring external retrieval. We design a structured reasoning template and employ Group Relative Policy Optimization(GRPO) to post-train the model's search behaviors. Our reward function encourages accurate answering without unnecessary search while promoting effective retrieval when needed. Experimental results demonstrate that our method significantly reduces redundant search operations while maintaining or improving answer accuracy across multiple challenging benchmarks. This framework advances the model's reasoning efficiency and extends its capability to judiciously leverage external knowledge.
Forward citations
Cited by 2 Pith papers
-
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
A DAPO-trained 4B model modestly improves selective evidence adoption on SelectBench-v2, but the gains are not statistically robust and prompt-injection resistance does not improve.
-
Agent Safety Alignment via Reinforcement Learning
RL-based safety alignment with an execute-refuse-verify policy improves reported threat resistance for tool-using agents, but utility preservation is not consistently demonstrated.
Discussion (0). Continue with ORCID to comment.