Pith. sign in

REVIEW 3 cited by

MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20163 v1 pith:I2FLQKZB submitted 2024-09-30 cs.AI cs.CL

classification cs.AIcs.CL
keywords memsimbayesiandatasetllm-basedmemorymessagespersonaluser
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset. To benefit the research community, we have released our project at https://github.com/nuster1128/MemSim.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems

    cs.HC 2025-06 conditional novelty 6.0 of 10

    RecUserSim combines profile, memory, action, and refinement modules in an LLM agent to generate realistic, diverse user utterances and multi-dimensional ratings for evaluating conversational recommender systems.

  2. Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A learnable memory cycle with adaptive retrieval, merging, and storage, trained online, improves LLM agent accuracy on HotpotQA and MemDaily for most backbones.

  3. MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MemBench introduces a multi-scenario, multi-level memory benchmark for LLM agents, evaluating factual and reflective memory across participation and observation settings.

Pith tools