Pith. sign in

REVIEW 6 cited by

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06378 v2 pith:73XVJFID submitted 2025-03-09 cs.AI cs.CLcs.CY

classification cs.AIcs.CLcs.CY
keywords powerscalestasksbenchmarksevaluationexplanatorygeneralpredictive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it has offered limited explanatory and predictive power for general-purpose AI systems, given the low transferability across diverse tasks. In this paper, we introduce general scales for AI evaluation that can explain what common AI benchmarks really measure, extract ability profiles of AI systems, and predict their performance for new task instances, in- and out-of-distribution. Our fully-automated methodology builds on 18 newly-crafted rubrics that place instance demands on general scales that do not saturate. Illustrated for 15 large language models and 63 tasks, high explanatory power is unleashed from inspecting the demand and ability profiles, bringing insights on the sensitivity and specificity exhibited by different benchmarks, and how knowledge, metacognition and reasoning are affected by model size, chain-of-thought and distillation. Surprisingly, high predictive power at the instance level becomes possible using these demand levels, providing superior estimates over black-box baseline predictors based on embeddings or finetuning, especially in out-of-distribution settings (new tasks and new benchmarks). The scales, rubrics, battery, techniques and results presented here represent a major step for AI evaluation, underpinning the reliable deployment of AI in the years ahead. (Collaborative platform: https://kinds-of-intelligence-cfi.github.io/ADELE.)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A training-free meta-reasoning loop that tracks remaining cognitive demand and steers an LLM's next action improves average accuracy by about 9 percent over chain-of-thought across three models and six benchmarks.

  2. Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Stochastic sampling within one LLM yields no detectable cross-question error structure (≤1 MP eigenvalue) while a 24-model ensemble yields four, so per-question uncertainty via self-consistency is real but structurall...

  3. 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new spatial reasoning benchmark shows current multimodal models lag humans badly and lack the item-level predictability humans show.

  4. Rethinking the Illusion of Thinking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Reasoning models' Towers of Hanoi failures persist under stepwise prompting, while River Crossing failures mostly vanish when tests are restricted to solvable configurations.

  5. Contextual Memory Intelligence -- A Foundational Paradigm for Human-AI Collaboration and Reflective Generative AI Systems

    cs.AI 2025-05 conditional novelty 4.0 of 10

    Contextual Memory Intelligence reframes memory as dynamic infrastructure and proposes the Insight Layer to preserve decision rationale, detect semantic drift, and support human-in-the-loop reflection.

  6. Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

    cs.CL 2025-07 conditional novelty 3.0 of 10

    An experience report arguing that reliable AI evaluation needs structured volunteer cohorts, statistical error bars, and shared infrastructure rather than just software engineering.

Pith tools