REVIEW 4 cited by
WildIFEval: Instruction Following in the Wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we introduce WildIFEval - a large-scale dataset of 7K real user instructions with diverse, multi-constraint conditions. Unlike prior datasets, our collection spans a broad lexical and topical spectrum of constraints, extracted from natural user instructions. We categorize these constraints into eight high-level classes to capture their distribution and dynamics in real-world scenarios. Leveraging WildIFEval, we conduct extensive experiments to benchmark the instruction-following capabilities of leading LLMs. WildIFEval clearly differentiates between small and large models, and demonstrates that all models have a large room for improvement on such tasks. We analyze the effects of the number and type of constraints on performance, revealing interesting patterns of model constraint-following behavior. We release our dataset to promote further research on instruction-following under complex, realistic conditions.
Forward citations
Cited by 4 Pith papers
-
Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code
A large dataset of 57.5K GitHub prompts is annotated with a new ontology, enabling quantitative analysis of how transactional prompts are used in real code.
-
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
Agentic coding prompts grow because the reasoning behind old instructions decays, and comments that preserve that reasoning halt the growth and recover instruction-following.
-
UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following
UNSPECIFIC generates constraints common to two similar articles, hardens only too-easy constraints, and scores satisfaction after summarization, yielding a harder and more natural instruction-following benchmark.
-
Revisiting the Reliability of Language Models in Instruction-Following
LLMs exhibit up to 61.8% performance drops on nuanced rephrasings of instruction-following tasks, revealing insufficient nuance-oriented reliability across 46 tested models.
Discussion (0). Continue with ORCID to comment.