REVIEW 7 cited by
HCAST: Human-Calibrated Autonomy Software Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world effects we care about. We present HCAST (Human-Calibrated Autonomy Software Tasks), a benchmark of 189 machine learning engineering, cybersecurity, software engineering, and general reasoning tasks. We collect 563 human baselines (totaling over 1500 hours) from people skilled in these domains, working under identical conditions as AI agents, which lets us estimate that HCAST tasks take humans between one minute and 8+ hours. Measuring the time tasks take for humans provides an intuitive metric for evaluating AI capabilities, helping answer the question "can an agent be trusted to complete a task that would take a human X hours?" We evaluate the success rates of AI agents built on frontier foundation models, and we find that current agents succeed 70-80% of the time on tasks that take humans less than one hour, and less than 20% of the time on tasks that take humans more than 4 hours.
Forward citations
Cited by 7 Pith papers
-
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Across ~38,000 hours on 134 ultra-long real-world tasks, aggregate agent performance follows a log-sigmoid of interaction time, and measured learning speed doubles about every three months.
-
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
In a randomized trial of 246 real open-source tasks, experienced developers took 19% longer when AI tools were allowed, despite forecasting 24% faster completion.
-
Predicting Task Difficulty Without Rollouts
Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.
-
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
A 46-task terminal benchmark with subtask-level dense rewards shows frontier agents rarely finish long workflows, with the best model at 28.3% pass@1 (R≥0.95).
-
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
In a new benchmark, AI agents underperform a matched-budget hyperparameter search at LLM server optimization because they converge early on one framework and test very few configurations.
-
BRIDGE: Predicting Human Task Completion Time From Model Performance
BRIDGE shows that item-response-theory difficulty estimated from model performance tracks log human completion time, enabling human time prediction and a ~6-month doubling forecast for frontier task horizons.
-
TextAtari: 100K Frames Game Playing with Language Agents
TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.
Discussion (0). Continue with ORCID to comment.