Pith. sign in

Paper Citation Record · LEDGER

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 6 inbound Pith citation observations for arXiv:2602.03224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.03224 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:10:16.892275Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:34.442077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:19:13.931629Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d37fefbb-a2ef-448d-9596-36b34d7d30c2 · outbound

This paper cites Training-free group relative policy optimization.arXiv preprint arXiv:2510.08191, 2025a.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Training-free group relative policy optimization.arXiv preprint arXiv:2510.08191, 2025a

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.279440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.279440Z digest=sha256:4d95d74a2abad13fb395197a0cc96ff34261e537bf4eb10fbc985d685025cfa1

Observation 9a0b3f58-8df8-4d85-8ebc-3587b704a186 · outbound

This paper cites Memory in the Age of AI Agents.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Memory in the Age of AI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.630170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.630170Z digest=sha256:e4aea906403a21256c78dc7991776877303f4a78533e9dafa4b39e37862634c5

Observation 81a4ace0-a003-4f72-98e1-8aac2def0b9d · outbound

This paper cites Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.074457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.074457Z digest=sha256:fe7a5bc8b742f3f2118a48b1fbeb3eede6756b9c5ec2ebff28308cf54a7d4caf

Observation 1d83b4c7-3c71-4084-b1d9-ada44e1e6785 · outbound

This paper cites Agent kb: Leveraging cross-domain experience for agentic problem solving.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Agent kb: Leveraging cross-domain experience for agentic problem solving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.221578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.221578Z digest=sha256:58a24d1ab2f41976b495ea1517f26f06d3a010493da5d1736d75546fa8fe48de

Observation 57ba4473-6a77-4b93-b53b-faf63ad6e205 · outbound

This paper cites H., and Li, B.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking H., and Li, B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.340068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.340068Z digest=sha256:a901ea1a513c55824d18232d5bb283dd9868266d1e319555b65ef92e68463685

Observation 1a2325b5-48cf-4067-9c9d-1a3e982142c2 · outbound

This paper cites A-memguard: A proactive defense framework for llm-based agent memory.arXiv preprint arXiv:2510.02373, 2025a.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking A-memguard: A proactive defense framework for llm-based agent memory.arXiv preprint arXiv:2510.02373, 2025a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.457201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.457201Z digest=sha256:fa4c2903624ffa927300bebcf02dd053dac9e9b802676574bbcc22098db5f6f3

Observation 5d53bfc4-5122-4e2e-a3ae-d696dee2c696 · outbound

This paper cites Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.751969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.751969Z digest=sha256:dd2c1172c0414da5a806e8fc65390a0e2964385860864f6950d969a289ee60cd

Observation c31674b3-dfe5-4e72-9ea2-1481aa47bee0 · outbound

This paper cites What are some methods for cyberstalking and harassing someone anonymously online?.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking What are some methods for cyberstalking and harassing someone anonymously online?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.892275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.892275Z digest=sha256:a7dfbd006f1d7a6d515940272fd52b122eca6a18b65308ee5e03d5a9f7823cc0

Observation 9d23d1d1-d750-48c7-a031-a841bb08e4c0 · outbound

This paper cites A definition of agi.arXiv preprint arXiv:2510.18212,.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking A definition of agi.arXiv preprint arXiv:2510.18212,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.495346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.495346Z digest=sha256:259b1b0d5eb3fbff8a0d9e55d2008813c5e8e32d93f12246e9e5a620a93b94df

Observation 3cf92171-df0c-43c1-af28-a441cfe0d64b · outbound

This paper cites Qwen3Guard Technical Report.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Qwen3Guard Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:16.641500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:16.641500Z digest=sha256:d32c6ae4d6231cf861aa52345802caeb621bd00b0da01beea9cfd2ebc3564b75

Observation ac4dacec-ed9a-4e7d-9914-27dd6b61f7a6 · outbound

This paper cites Realm: Robust entropy adaptive loss mini- mization for improved single-sample test-time adaptation.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Realm: Robust entropy adaptive loss mini- mization for improved single-sample test-time adaptation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.925069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.925069Z digest=sha256:3ab89056fadd910c60e694b5b3557428b63e0171b99505800f8a4137ab961ea2

Observation 2b2621d7-b6d2-4eef-8a68-5943d6b0d0e0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.359987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.359987Z digest=sha256:10a065edf7c17210f1b25a66b0b7102beae5179f4a907b0d96765bf3f4a97b52

Observation 8f7b1cb3-37da-4aa7-b86a-2da3b33f4bcc · outbound

This paper cites SCOPE: Prompt Evolution for Enhancing Agent Effectiveness.

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T05:10:15.763194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:10:15.763194Z digest=sha256:a81cef684fdc33aedb139e50ac1eb6b01eb9277e78f6f3cb8a306007d5630aeb

Pith citing papers

Observation e841d9b2-241f-4e58-bbaf-b8c28ae6136b · inbound

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution cites this paper.

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:12.787754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:52:47.760627Z digest=sha256:d171768330e4c838bd5c83eb779b2fbfc024a518eeaa418e9050b912f4c279b1

Observation 3ca8e6d4-724d-4928-90fa-2b396982da3f · inbound

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution cites this paper.

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T19:51:48.531280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:51:48.531280Z digest=sha256:189be8da571a62c68d70e9f928bddc184b8e208279f66aeebee9585dea91e85f

Observation cf257578-ef28-413b-89c3-e81e0d401fb2 · inbound

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems cites this paper.

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:12.787754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:43:34.500801Z digest=sha256:12d04b3331709424716d747bfccb84b66bbbd7baaa6b6affe9164a74392b0c24

Observation 7f82a095-c2b6-4f1e-9f7e-ec329c5628c8 · inbound

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness cites this paper.

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:19:13.932961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:15:30.776560Z digest=sha256:f8d5031fda3cfa191abdeabf0abc011ea3ea2f77f11b9665e3c014107d1d611f

Observation 3ee05cf0-5e94-4e2f-9687-6e023e8b74c5 · inbound

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness cites this paper.

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T10:58:01.942097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:58:01.942097Z digest=sha256:28822faf8f3e5d4c6d79c185bdc671a0172f2721b9047503578a4e020ff2d312

Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · inbound

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? cites this paper.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.442077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.442077Z digest=sha256:5b50bbf0916463829921c36f473595dafe7764ca43efbd876654f5d1121a9724