Pith. sign in

REVIEW 7 cited by

HCAST: Human-Calibrated Autonomy Software Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.17354 v1 pith:ZMIFYRSO submitted 2025-03-21 cs.AI

classification cs.AI
keywords taskstakehourshumansagentshcastsoftwaretime
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world effects we care about. We present HCAST (Human-Calibrated Autonomy Software Tasks), a benchmark of 189 machine learning engineering, cybersecurity, software engineering, and general reasoning tasks. We collect 563 human baselines (totaling over 1500 hours) from people skilled in these domains, working under identical conditions as AI agents, which lets us estimate that HCAST tasks take humans between one minute and 8+ hours. Measuring the time tasks take for humans provides an intuitive metric for evaluating AI capabilities, helping answer the question "can an agent be trusted to complete a task that would take a human X hours?" We evaluate the success rates of AI agents built on frontier foundation models, and we find that current agents succeed 70-80% of the time on tasks that take humans less than one hour, and less than 20% of the time on tasks that take humans more than 4 hours.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

    cs.CL 2026-07 conditional novelty 7.5 of 10

    Across ~38,000 hours on 134 ultra-long real-world tasks, aggregate agent performance follows a log-sigmoid of interaction time, and measured learning speed doubles about every three months.

  2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    cs.AI 2025-07 conditional novelty 7.0 of 10

    In a randomized trial of 246 real open-source tasks, experienced developers took 19% longer when AI tools were allowed, despite forecasting 24% faster completion.

  3. Predicting Task Difficulty Without Rollouts

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.

  4. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A 46-task terminal benchmark with subtask-level dense rewards shows frontier agents rarely finish long workflows, with the best model at 28.3% pass@1 (R≥0.95).

  5. InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

    cs.AI 2026-05 conditional novelty 6.0 of 10

    In a new benchmark, AI agents underperform a matched-budget hyperparameter search at LLM server optimization because they converge early on one framework and test very few configurations.

  6. BRIDGE: Predicting Human Task Completion Time From Model Performance

    cs.AI 2026-02 conditional novelty 6.0 of 10

    BRIDGE shows that item-response-theory difficulty estimated from model performance tracks log human completion time, enabling human time prediction and a ~6-month doubling forecast for frontier task horizons.

  7. TextAtari: 100K Frames Game Playing with Language Agents

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.

Pith tools