Pith. sign in

REVIEW 7 cited by

RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03610 v1 pith:FTEPFJWV submitted 2024-02-06 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords agentsmultimodalplanningapplicationscomplexcurrentdecision-makingexperiences
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Owing to recent advancements, Large Language Models (LLMs) can now be deployed as agents for increasingly complex decision-making applications in areas including robotics, gaming, and API integration. However, reflecting past experiences in current decision-making processes, an innate human behavior, continues to pose significant challenges. Addressing this, we propose Retrieval-Augmented Planning (RAP) framework, designed to dynamically leverage past experiences corresponding to the current situation and context, thereby enhancing agents' planning capabilities. RAP distinguishes itself by being versatile: it excels in both text-only and multimodal environments, making it suitable for a wide range of tasks. Empirical evaluations demonstrate RAP's effectiveness, where it achieves SOTA performance in textual scenarios and notably enhances multimodal LLM agents' performance for embodied tasks. These results highlight RAP's potential in advancing the functionality and applicability of LLM agents in complex, real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Conflicting memory enters early and propagates with weak recovery, producing similar compliance rates across models but larger absolute damage for stronger agents.

  2. Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses

    cs.CR 2026-07 conditional novelty 7.0 of 10

    FARMA forges and self-amplifies an agent's reasoning history with evasive language to induce unsafe skips; SENTINEL's Reasoning Guard reduces ASR to 0% across tested agents and models.

  3. When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents

    cs.CR 2026-07 conditional novelty 6.5 of 10

    GhostWriter poisons tool-using personal agents' long-term memory via untrusted emails/calendar invites (~98% injection, ~60% activation); AM-Sentry policies and retrieval screens sharply reduce success while preservin...

  4. MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    MAFIA poisons RAG agent memory through probing and compact factual cloaks, reaching up to 90.7% attack success while evading low-FPR input audits.

  5. Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Modeling agent trajectories as action-centric probabilistic graphs lets a GNN warn LLM agents of likely step-level errors before execution, improving pass ratio ~14.7% across four benchmarks.

  6. Object-Centric Environment Modeling for Agentic Tasks

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Object-Centric Environment Modeling (OCM) builds an online executable object-and-procedure code model that improves average rank and cuts invalid actions on ScienceWorld, ALFWorld, and PlanCraft.

  7. Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents

    cs.PL 2026-06 unverdicted novelty 6.0 of 10

    FCGraft synthesizes code policies for embodied agents by grafting KV caches from a library of validated functions, claiming 18.31% higher success rate and 2.3x faster synthesis than prompt-level caching.

Pith tools