Pith. sign in

REVIEW 6 cited by

Large Language Models Can Self-Improve At Web Agent Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20309 v2 pith:WOXYTTAQ submitted 2024-05-30 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords modelsagentagentsbenchmarkdatalanguagellmsnavigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31\% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Skill Self-Play uses an evolving skill library to guide LLM self-training, improving tool-call and reasoning accuracy beyond unguided self-play on five model backbones.

  2. Morae: Proactively Pausing UI Agents for User Choices

    cs.HC 2025-08 conditional novelty 6.0 of 10

    Morae, a UI agent that proactively pauses at ambiguous decision points, helps blind and low-vision users complete more tasks and express preferences better than fully autonomous agents.

  3. Build the web for agents, not agents for the web

    cs.LG 2025-06 conditional novelty 6.0 of 10

    The paper proposes a paradigm shift: design a standardized Agentic Web Interface for AI agents, rather than adapting agents to human-facing websites.

  4. InSTA: Towards Internet-Scale Training For Agents

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Automated LLM task generation, agent execution, and judge filtering at 150k-site scale lets a 1.7B model match much larger web agents.

  5. Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models

    cs.CL 2025-12 conditional novelty 4.0 of 10

    Across five reasoning benchmarks, sparse MoE models dominate the accuracy-vs-FLOPs Pareto frontier, inference-compute gains saturate, and wrong answers systematically consume more compute than correct ones.

  6. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools