Pith. sign in

REVIEW 6 cited by

LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.11942 v3 pith:L7Y5VORU submitted 2025-05-17 cs.AI

classification cs.AI
keywords agentslifelonglearninglifelongagentbenchenvironmentsknowledgeoperatingability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Lifelong learning is essential for intelligent agents operating in dynamic environments. Current large language model (LLM)-based agents, however, remain stateless and unable to accumulate or transfer knowledge over time. Existing benchmarks treat agents as static systems and fail to evaluate lifelong learning capabilities. We present LifelongAgentBench, the first unified benchmark designed to systematically assess the lifelong learning ability of LLM agents. It provides skill-grounded, interdependent tasks across three interactive environments, Database, Operating System, and Knowledge Graph, with automatic label verification, reproducibility, and modular extensibility. Extensive experiments reveal that conventional experience replay has limited effectiveness for LLM agents due to irrelevant information and context length constraints. We further introduce a group self-consistency mechanism that significantly improves lifelong learning performance. We hope LifelongAgentBench will advance the development of adaptive, memory-capable LLM agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

    cs.CL 2026-07 conditional novelty 7.5 of 10

    Across ~38,000 hours on 134 ultra-long real-world tasks, aggregate agent performance follows a log-sigmoid of interaction time, and measured learning speed doubles about every three months.

  2. FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A longitudinal benchmark of 120 real-case financial tasks shows four agent scaffolds beat their state-reset controls by 9.33 to 19.37 score points when allowed to retain experience.

  3. Progressive Multimodal Alignment for Continual Instruction Tuning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Progressive Multimodal Alignment expands projector experts only when multimodal distribution shifts are detected, reducing projector-level forgetting and boosting MCIT baselines with sub-linear growth.

  4. SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single RL policy alternating task solving and skill-document curation, with decoupled cross-task credit, improves Pass@1 and cross-task test-time scaling on ALFWorld, WebShop, and ScienceWorld.

  5. Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

    cs.AI 2026-07 conditional novelty 5.0 of 10

    SkillBoost reduces skill overfitting in self-evolving LLM agents by combining failure-localized editing, multi-candidate generation, and an anti-regression acceptance gate.

  6. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

Pith tools