Pith. sign in

REVIEW 3 cited by

KwaiAgents: Generalized Information-seeking Agent System with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04889 v3 pith:HFYWX3BD submitted 2023-12-08 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords agentllmssystemgeneralizedkwaiagentsthemcapabilitieseven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driven by curiosity, humans have continually sought to explore and understand the world around them, leading to the invention of various tools to satiate this inquisitiveness. Despite not having the capacity to process and memorize vast amounts of information in their brains, humans excel in critical thinking, planning, reflection, and harnessing available tools to interact with and interpret the world, enabling them to find answers efficiently. The recent advancements in large language models (LLMs) suggest that machines might also possess the aforementioned human-like capabilities, allowing them to exhibit powerful abilities even with a constrained parameter count. In this paper, we introduce KwaiAgents, a generalized information-seeking agent system based on LLMs. Within KwaiAgents, we propose an agent system that employs LLMs as its cognitive core, which is capable of understanding a user's query, behavior guidelines, and referencing external documents. The agent can also update and retrieve information from its internal memory, plan and execute actions using a time-aware search-browse toolkit, and ultimately provide a comprehensive response. We further investigate the system's performance when powered by LLMs less advanced than GPT-4, and introduce the Meta-Agent Tuning (MAT) framework, designed to ensure even an open-sourced 7B or 13B model performs well among many agent systems. We exploit both benchmark and human evaluations to systematically validate these capabilities. Extensive experiments show the superiority of our agent system compared to other autonomous agents and highlight the enhanced generalized agent-abilities of our fine-tuned LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    InfoDeepSeek is a 245-question benchmark that measures how well AI agents seek information on the live web, with new metrics for answer accuracy, evidence quality, and compactness.

  2. Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.

  3. Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A survey of long video generation that groups methods into frame-by-frame, planned-segment, and all-at-once approaches, but is weakened by inconsistent data tables.

Pith tools