Pith. sign in

REVIEW 2 cited by

EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08893 v2 pith:753FT7B5 submitted 2025-03-11 cs.CL cs.LG

classification cs.CLcs.LG
keywords evaltreeweaknesscapabilityprofilinglanguagemodelweaknessescollection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An ideal model evaluation should achieve two goals: identifying where the model fails and providing actionable improvement guidance. Toward these goals for language model (LM) evaluations, we formulate the problem of generating a weakness profile, a set of weaknesses expressed in natural language, given an LM's performance on every individual instance in a benchmark. We introduce a suite of quantitative assessments to compare different weakness profiling methods. We also introduce a weakness profiling method EvalTree. EvalTree constructs a capability tree where each node represents a capability described in natural language and is linked to a subset of benchmark instances that specifically evaluate this capability; it then extracts nodes where the LM performs poorly to generate a weakness profile. On the MATH and WildChat benchmarks, we show that EvalTree outperforms baseline weakness profiling methods by identifying weaknesses more precisely and comprehensively. Weakness profiling further enables weakness-guided data collection, and training data collection guided by EvalTree-identified weaknesses improves LM performance more than other data collection strategies. We also show how EvalTree exposes flaws in Chatbot Arena's human-voter-based evaluation practice. To facilitate future work, we provide an interface that allows practitioners to interactively explore the capability trees built by EvalTree.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

    cs.CL 2025-07 conditional novelty 6.0 of 10

    CLEAR converts per-instance LLM judge critiques into system-level error issues with prevalence counts and an interactive dashboard for exploration.

  2. Data Swarms: Optimizable Generation of Synthetic Evaluation Data

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Data Swarms uses particle swarm optimization over data-generator LLM weights to produce synthetic evaluation data that scores higher on five quantitative evaluation objectives than eight baselines.

Pith tools