Pith. sign in

REVIEW 8 cited by

AI Agents: Evolution, Architecture, and Real-World Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.12687 v1 pith:L5FINDVB submitted 2025-03-16 cs.AI

classification cs.AI
keywords applicationsagentagentsarchitectureevaluationevolutionreal-worldsystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper examines the evolution, architecture, and practical applications of AI agents from their early, rule-based incarnations to modern sophisticated systems that integrate large language models with dedicated modules for perception, planning, and tool use. Emphasizing both theoretical foundations and real-world deployments, the paper reviews key agent paradigms, discusses limitations of current evaluation benchmarks, and proposes a holistic evaluation framework that balances task effectiveness, efficiency, robustness, and safety. Applications across enterprise, personal assistance, and specialized domains are analyzed, with insights into future research directions for more resilient and adaptive AI agent systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

    cs.AI 2026-03 conditional novelty 7.0 of 10

    SciVisAgentBench provides 108 expert-crafted tasks and a mixed LLM-plus-deterministic evaluation pipeline for benchmarking AI agents that perform scientific visualization workflows.

  2. PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

    cs.CL 2026-07 conditional novelty 6.0 of 10

    PARALLEL learns per-sample LoRA update intensity (skip/light/strong) with a budget-constrained REINFORCE controller, retaining 94–99% of full-adaptation performance at 30% update mass.

  3. Agentic Services Computing

    cs.SE 2025-09 conditional novelty 5.0 of 10

    A position and survey paper that defines Agentic Services Computing, a lifecycle-based framework for engineering LLM agents as governed, first-class services.

  4. Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A survey that groups graph-empowered AI agent research into planning, execution, memory, and multi-agent coordination, plus agents-for-graphs and applications.

  5. How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests

    cs.SE 2026-07 conditional novelty 4.0 of 10

    AI coding agents mostly handle routine, well-scoped development tasks; their pull requests are merged at similar rates and have comparable or lower bug-proneness than human-written pull requests across repository lifecycles.

  6. Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions

    cs.NI 2025-08 conditional novelty 4.0 of 10

    A survey that organizes agentic AI for 6G edge networks into four pillars, compactness, efficiency, knowledge and reasoning, and migration, and illustrates them with prior case studies.

  7. BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web

    cs.MA 2025-08 unverdicted novelty 4.0 of 10

    BetaWeb promises a blockchain-enabled trustworthy agentic web, but the submitted manuscript body is a different mining-robot paper, leaving the proposal without supporting evidence.

  8. Towards Log Analysis with AI Agents: Cowrie Case Study

    cs.CR 2025-08 conditional novelty 2.0 of 10

    A rule-based pipeline processes 313,412 Cowrie honeypot log events into 26,368 labeled attacker sessions and reports that most are automated low-skill probes.

Pith tools