REVIEW 3 cited by
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.
Forward citations
Cited by 3 Pith papers
-
RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems
RecUserSim combines profile, memory, action, and refinement modules in an LLM agent to generate realistic, diverse user utterances and multi-dimensional ratings for evaluating conversational recommender systems.
-
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
A learnable memory cycle with adaptive retrieval, merging, and storage, trained online, improves LLM agent accuracy on HotpotQA and MemDaily for most backbones.
-
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
MemBench introduces a multi-scenario, multi-level memory benchmark for LLM agents, evaluating factual and reflective memory across participation and observation settings.
Discussion (0). Sign in to comment.