REVIEW 7 cited by
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Long-term memory plays a critical role in personal interaction, considering long-term memory can better leverage world knowledge, historical information, and preferences in dialogues. Our research introduces PerLTQA, an innovative QA dataset that combines semantic and episodic memories, including world knowledge, profiles, social relationships, events, and dialogues. This dataset is collected to investigate the use of personalized memories, focusing on social interactions and events in the QA task. PerLTQA features two types of memory and a comprehensive benchmark of 8,593 questions for 30 characters, facilitating the exploration and application of personalized memories in Large Language Models (LLMs). Based on PerLTQA, we propose a novel framework for memory integration and generation, consisting of three main components: Memory Classification, Memory Retrieval, and Memory Synthesis. We evaluate this framework using five LLMs and three retrievers. Experimental results demonstrate that BERT-based classification models significantly outperform LLMs such as ChatGLM3 and ChatGPT in the memory classification task. Furthermore, our study highlights the importance of effective memory integration in the QA task.
Forward citations
Cited by 7 Pith papers
-
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Setoka evaluates memory-augmented agents on four levels of user understanding—semantic memory, episodic memory, behavior patterns, and personality traits—over synthesized heterogeneous user data, and finds performance...
-
LMEB: Long-horizon Memory Embedding Benchmark
LMEB is a new benchmark that evaluates embedding models on long-horizon memory retrieval and shows this skill is largely orthogonal to traditional passage-retrieval performance.
-
MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
MemSifter trains a 4B proxy with an outcome-driven, rank-sensitive RL reward to sift LLM memory, and on eight benchmarks it matches or beats embedding, graph, and long-context baselines.
-
MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
MemoCue uses a 5W scenario classifier and Monte Carlo Tree Search to generate cue-rich questions that help LLM agents guide humans through memory recall.
-
How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
A new open-domain dialogue dataset shows that users' own judgments of stylistic similarity correlate with their preference (Spearman r=0.67-0.75), but third-party stylistic similarity judgments do not, indicating a ga...
-
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
MemBench introduces a multi-scenario, multi-level memory benchmark for LLM agents, evaluating factual and reflective memory across participation and observation settings.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
Discussion (0). Continue with ORCID to comment.