Pith. sign in

Paper Citation Record · LEDGER

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

As of 7 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 0 inbound Pith citation observations for arXiv:2608.00155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00155 v1

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b611d2-1277-424a-96f0-bd6abc117c0c · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.678535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.678535Z digest=sha256:43a9801b081465f44a206209a5861f70fd3b4a5cf98f778811a30636ab8d3d22

Observation 15c3b5f3-0cc4-407d-97dd-53e6f149a015 · outbound

This paper cites Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.740362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.740362Z digest=sha256:9177455233ccb2c80d140ecc944a276b18111e57458ba607b9ab8f99164eea3a

Observation 774e17c1-81ba-4c06-8184-c3942a218a0c · outbound

This paper cites A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.844983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.844983Z digest=sha256:b0071fd3d8bfb69bd1ad48af020e576d0a2842697a2c9d916762a876eae050db

Observation f6d0b54b-6b3b-4420-ba41-a3bb96d97d7c · outbound

This paper cites Agentic context engineering: Evolving contexts for self-improving language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic context engineering: Evolving contexts for self-improving language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.945741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.945741Z digest=sha256:7980633f9f36569650ce040c89fe71bf6cf551c84443875cca4a5035b2c20b92

Observation 4b30595c-b4bf-4321-82f1-c1ac004ba677 · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.047941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.047941Z digest=sha256:2cc5a8d751bcc5122dcace8a6327f4774069b98af117bff4452f90a2769b4862

Observation 4d1cbd62-4880-457e-9d8e-76e4c51ce49d · outbound

This paper cites A-mem: Agentic memory for llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A-mem: Agentic memory for llm agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.168633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.168633Z digest=sha256:c00aa2fc6d93a2a5907f83b33f73f4a8494fe2b8cab196f10adcee4907a785c9

Observation 5b4a68ac-830d-4563-8af2-bc2c42882100 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.238737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.238737Z digest=sha256:652f8315b0bbbf45ee924db89abef03c167459c2b89aac150c15fad1d2cf7ada

Observation 9997acb5-415b-4035-aced-f5f12b482392 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.339997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.339997Z digest=sha256:410297f3fddc5bc9f8496c79e6f84bcdc8ad47172ef4c754bcaef6dd846f58cd

Observation 2ae5e5f3-dae5-48a1-9f88-e580831d52be · outbound

This paper cites Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.448963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.448963Z digest=sha256:3c9b835431f8853c2e5adf8c634359dae160cdd09fe149c9e5b561fba6e58c96

Observation 669b2ad1-81a2-4075-bd74-1379c5d0b903 · outbound

This paper cites Memento-skills: Let agents design agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Memento-skills: Let agents design agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.553680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.553680Z digest=sha256:928f85f74cf003a9d790ed46dc6ba8a70f8aef0f5c70fed9b8d228e2c3a854ed

Observation aff5ba51-e69a-4fe0-958c-0ca5c67e030e · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.613003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.613003Z digest=sha256:b3ae0e8cff8649e334d642268623157844d0840817347643522bba30883a781c

Observation 8650ca55-dd31-4e2b-b6c2-2e8233c3b8e2 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.709243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.709243Z digest=sha256:bc0c50ee27b3447bf31c444fdb1426b9985301c2522fd51174eecd85841c03de

Observation 0313bafd-50e0-4c88-95b5-24fa3f54c936 · outbound

This paper cites Gonzalez.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gonzalez

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.826496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.826496Z digest=sha256:e088a25c6b8cabf5989fc3773d18faf1f9522144f025698fc51389448a7fe3b7

Observation 192ac56d-3eb2-4e4d-966f-b596268400d9 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In Proc.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Swe-bench: Can language models resolve real-world github issues? In Proc

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.918060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.918060Z digest=sha256:b48e32e2a1e1e4007457e62351afc8eed2e16b7a27a7a4fff5f1fab018a1fd00

Observation c22fc5a9-a231-4699-b27a-21e74d18ae93 · outbound

This paper cites Humanity's Last Exam.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Humanity's Last Exam

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.996592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.996592Z digest=sha256:6c478ed6a766c5dcc90032b104658054a398291a8d26fc8db48f788e66dc1105

Observation 45ed4434-2e42-4aac-a302-79bb5630afb5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.087301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.087301Z digest=sha256:bcb34c6f848a960ece6356a0d1862cac65261d16bdf2ddd40ef6b9669ce0597e

Observation 0ab7ed43-9bc6-413b-8b5e-082f746f58f4 · outbound

This paper cites Stream- bench: Towards benchmarking continuous improvement of language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Stream- bench: Towards benchmarking continuous improvement of language agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.156197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.156197Z digest=sha256:6f547a751ee65726dca3f47f768365daf117d1342838934fb316a138217c3570

Observation 19de0ee3-c2cd-4dac-84f9-1710b859a36c · outbound

This paper cites Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.201860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.201860Z digest=sha256:db9ba26a54e0daf36384ae9ca24d43b4872dd9b7b6936fc13096031bce1de033

Observation b8cc89d3-74d9-495a-a8d0-64d85f8c11ec · outbound

This paper cites OpenAI GPT-5 System Card.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenAI GPT-5 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.267699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.267699Z digest=sha256:cf5a62fd6182661a869142784e755979cef413d16ea151f3911faed426e46227

Observation 5ba34156-2f11-4e2f-be16-4be50f196bea · outbound

This paper cites Gemini 3.1 Pro model card, February 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gemini 3.1 Pro model card, February 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.359163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.359163Z digest=sha256:c303b3fdb862d8c315fc6152ca3bdf271c02dfa1914daa867fafc8a5a82f61b8

Observation b9ba58dd-a90c-4480-9b54-3e6fb70502a1 · outbound

This paper cites Introducing Claude Opus 4.7, April 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Introducing Claude Opus 4.7, April 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.441241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.441241Z digest=sha256:8512efade09d0290d11475920a3d25becf31d33fd8dbae7ef43c0aaa66f177c3

Observation 1e4e000b-6c7e-4e9c-b814-b4951cdeda99 · outbound

This paper cites Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.517774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.517774Z digest=sha256:00875ccf8174be5d262a129956cdbbe9e866c908bfde6542bf51b98a2658289a

Observation a9c9edd8-577d-471a-81e4-899bf8f7293c · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training with self-supervision for generalization under distribution shifts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.557621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.557621Z digest=sha256:36be1c009b0c9eb87753ce7d335f00d53d5e1a6685f12da7ef4a518b0e235d4b

Observation 1ef7ec98-9596-452d-b766-b5a37395abac · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.616665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.616665Z digest=sha256:efb88664b1cf172a2bd8c460c64bfdb9fed4cd8eb57e88bbba861a04be03c284

Observation 4ec2d7e4-606f-45df-bcbf-08412cdd5a1a · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.706158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.706158Z digest=sha256:c5571ad3203c3fede11ef7af419b500aa0495b0d9a955d36f3f931fdecbb172d

Observation 7128a647-c5e2-4fc7-bee4-44afb379adaa · outbound

This paper cites Gradient episodic memory for continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gradient episodic memory for continual learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.785239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.785239Z digest=sha256:f285f079bc3a54ae71d722d8d779d450b225b2e62398762146162448753c9335

Observation 13b29ec8-639f-4f57-9e71-ec2d6c5b41ed · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Scaling llm test-time compute optimally can be more effective than scaling model parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.834609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.834609Z digest=sha256:0075620a5829b1c886e81a495e09136205d0ee030fd2ef9773eb29c4da8044e8

Observation ea24b3aa-c55c-429d-8f36-4f6fd65c80e7 · outbound

This paper cites Test-time training on nearest neighbors for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training on nearest neighbors for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.907311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.907311Z digest=sha256:0413ea494b01bf5d2c57e0954b2aa850d11d6fabd3f51bf3ed38c6b2efa61d7c

Observation 69b76d30-5c03-4755-a817-87bedd5539bf · outbound

This paper cites Efficiently learning at test-time: Active fine-tuning of llms.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Efficiently learning at test-time: Active fine-tuning of llms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.947220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.947220Z digest=sha256:176cd3223fc1df170f1196f3456d478181919ed173db4579a51a920e1e6d9063

Observation 8353d4e8-7813-4986-b32e-8b3fd29c1545 · outbound

This paper cites The surprising effectiveness of test-time training for few-shot learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The surprising effectiveness of test-time training for few-shot learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.990733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.990733Z digest=sha256:6b88528063040ab6585fe6df9fe2767aad06e368da56fa8c7c395f0f824f089b

Observation 1aaae17b-f559-45b8-b4aa-042734c11429 · outbound

This paper cites In-place test-time training.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? In-place test-time training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.052856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.052856Z digest=sha256:3035292dac90501ba91c6fdd1b1ef17034c4a19d3740f631e902938fdd800dcd

Observation 78bb23fd-0340-46e7-a979-5c9cb6c60f57 · outbound

This paper cites Test-time adaptation for llm agents via environment interaction.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time adaptation for llm agents via environment interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.147161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.147161Z digest=sha256:5f80b2b57d66aa7a6c11136badd43cd87ee824850f547412712ef2bae8922dbc

Observation 8c6d34b6-56e8-42fd-be9c-40b104d6fb5b · outbound

This paper cites Test-time learning for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time learning for large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.266388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.266388Z digest=sha256:3307b6f3931154394dded5fe0e42136ec85ed4e3100e2a80d84fc3a23236746e

Observation 31c98bfe-fc30-4da1-b482-26138dfd8ce1 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttrl: Test-time reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.347224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.347224Z digest=sha256:7a7b3521d3b7d53a636d18912e1abd3481128e9d6f5173fc3e2fd1e989fb6af2

Observation 560c579e-c5e5-4904-893b-432431fa2cb1 · outbound

This paper cites Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.407176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.407176Z digest=sha256:50ca727ff348fb184ece44992b48608dbbce693a65f3770b567ef8b421714fcb

Observation 729ca982-0b5d-48a6-bc45-3aa1aa448f4f · outbound

This paper cites Learning to discover at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to discover at test time

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.486960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.486960Z digest=sha256:ec3f1240350a8a38597eac532b79a44cd15e7e49cedbf4e4c59d67b40d069fe9

Observation 0dbd5315-bd55-49ce-a095-a6213b44dc44 · outbound

This paper cites Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.584703Z digest=sha256:2a466e860d443ad23291af9f8d9adfb9ced3dfe73e3e854c1ce5e92c4caa3524

Observation 8a992a21-8df4-4314-91b8-8d6506facb9c · outbound

This paper cites What if consensus lies? selective-complementary reinforcement learning at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? What if consensus lies? selective-complementary reinforcement learning at test time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.672842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.672842Z digest=sha256:6a33d8424bdc255934b9a57b8bb6553618fd433fbff1f41e14bd8026cf005d32

Observation 0cb42fec-7ea7-47d8-add2-4a5c603aa6c1 · outbound

This paper cites Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.783687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.783687Z digest=sha256:b04d6aba550085eef91f89c5cc9a848d481b78a715e8559c8895336476df5f49

Observation d76d5b23-7867-48c6-ba76-c7b3adbf5fd3 · outbound

This paper cites Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.899505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.899505Z digest=sha256:46f3062324c4881f6e7fb7294d6f447f62ad639dbabd87c54190cb4500a485e2

Observation b01a108e-71ff-4206-a9bc-190f1a4e571e · outbound

This paper cites Test-Time Learning with an Evolving Library.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-Time Learning with an Evolving Library

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.009288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.009288Z digest=sha256:a98ae5807e1d6d01d6f6f3687a5384c172523426497c82e5095da9346ce980ab

Observation eb9d85b7-402a-49fc-aec2-7efb0b250981 · outbound

This paper cites Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.114246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.114246Z digest=sha256:ab1d50cc0d1aef269210b3164e3fa869753fe06e0e639016f5725c5c970bc090

Observation 4b2262f2-6162-4fac-abde-a3da7f3f3208 · outbound

This paper cites Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.226463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.226463Z digest=sha256:fd7af04ddbef6efb0e80e6420b91e75578b5313dad90dde42d5338c3e5af4724

Observation 26fb1c9a-50dd-4827-84ac-1aa5afaa48dd · outbound

This paper cites Agentic plan caching: Test-time memory for fast and cost-efficient llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic plan caching: Test-time memory for fast and cost-efficient llm agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.337320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.337320Z digest=sha256:c16d544a23562e8261450f679ca8928ddbed1a989bbd81468b9a44ddad08605f

Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · outbound

This paper cites TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.442077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.442077Z digest=sha256:c55bd4eb8676018e2001702fd6bb7741fcb2fccc83d15d2b09b06719d6524e10

Observation 18107843-204b-4eb5-b983-e6c2028238e2 · outbound

This paper cites Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.540281Z digest=sha256:65c9877bd95e687c9d91f66f64fdf7f0895cb402150396c54093c26e86186ac7

Observation b09b02a8-d8b8-453a-b705-88ffabf8a7c3 · outbound

This paper cites Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.652568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.652568Z digest=sha256:525473273e03f2ff1e1780c72a6d35a43e93439eb2c80fb38d0cbc895d20c8ec

Observation 9f5ea4cc-4e98-4228-9cbd-eb3796899c71 · outbound

This paper cites Panini: Continual learning in token space via structured memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Panini: Continual learning in token space via structured memory

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.761418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.761418Z digest=sha256:0946ac28cbcfe1d6666c3ea1cdfcea8f2c2745904ed1e7b9a34724c54a1cf749

Observation 5a6a2359-872e-42b0-bcc8-368eb0a6533b · outbound

This paper cites Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.866944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.866944Z digest=sha256:8e0c8f953fb6a1a110eb8825166d2759b518b94a37d1e36138e320ba621be1ce

Observation df12cecf-038a-4dc1-8a2b-c5be36a6d665 · outbound

This paper cites Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.939316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.939316Z digest=sha256:50dd82d3177aa722568e83a51c964620f33e48fcd914335e6fa6f74268d6771f

Observation fec8e04f-7212-49fe-a4ed-8b154e68aa71 · outbound

This paper cites Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.017911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.017911Z digest=sha256:bec0e71d13f3dc6126b296bf0d6ce5ae9ded810fe083ca9daffd004bb6dc8388

Observation fb473b65-f10e-480b-bb9b-04a615958777 · outbound

This paper cites Xskill: Continual learning from experience and skills in multimodal agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xskill: Continual learning from experience and skills in multimodal agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.087726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.087726Z digest=sha256:9187d3a0d16a4dea32eb68c7909298e24269fa23ff7de9cb3aeac241e0ce7cd9

Observation da43f976-c6ce-463c-9799-361ce994f6aa · outbound

This paper cites Online Experiential Learning for Language Models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Online Experiential Learning for Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.159727Z digest=sha256:2d6669e96fdb9ecc8da27fa4875df762b9a53f8da9ecca9f0d0a4ff11a3b6b53

Observation ddc78d07-c649-491c-9b06-7196fe1fb0e7 · outbound

This paper cites Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.227678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.227678Z digest=sha256:fce9a9e1772588464e403bdcbbffeb853dd9ab00f45c53816fadd6e3a080a556

Observation d42845b4-b90e-46aa-b1b5-2dbf86837931 · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.336093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.336093Z digest=sha256:6856d0d2ae61dc43503dabaf95a0313abfd8df5141107b1c05cc5a1ca9da37e3

Observation 227b17f5-290a-48c0-94bd-9da4a8674c98 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.459460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.459460Z digest=sha256:c60fbe21f54257c956f47b40679f0d11ec13a50dff36bb767196b6e3f05c948a

Observation e6271b35-7585-4cff-ac8d-17152f2132e8 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.579166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.579166Z digest=sha256:c0e92534eb77dff24ea416ec1cd3587ad3ee9623418925c703d596db260c6a2b

Observation 1fcce961-c4c5-4a04-8cba-613fc43b0b43 · outbound

This paper cites SkillOS: Learning Skill Curation for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOS: Learning Skill Curation for Self-Evolving Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.690973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.690973Z digest=sha256:85fad6b087f35b8ba504054f0083109d30b5c1b36d587917ecc109b2e3a97a34

Observation 60d83db8-2aa5-4635-b235-d2f7d73e311e · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.780877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.780877Z digest=sha256:a0b5bd8cfec782e6ffdc45606ac7bdc380264579629fba54ec93669033132d68

Observation 77a4cd23-7cf0-42d9-bfbc-c364071b1d2d · outbound

This paper cites OpenSkill: Open-World Self-Evolution for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenSkill: Open-World Self-Evolution for LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.900154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.900154Z digest=sha256:896da4f17387db1946cb4044a6a77fc080d80809657cbac42227105032765a60

Observation 3406787a-4e71-4b33-9ab7-b8a54d9deec6 · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.013114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.013114Z digest=sha256:9072cbc912ff43c653ed2b1d9ae5295e862d716bad5963db12c0667611ba53d3

Observation 9f8206b5-8e87-4d2a-82ea-3561f5a75200 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.133344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.133344Z digest=sha256:acb4dd910d7102853797ca78a242616136bf2ffb22ae1162d3eeb0c070fed8c6

Observation f4deb2ae-ff8d-44e6-a47e-4bf4271f2963 · outbound

This paper cites Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.247731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.247731Z digest=sha256:57576e4755c8d4b18b3489b2b0392dd1b3354b2f8d9231f4098082663ba664c3

Observation fd65faba-7aaf-4848-ae01-57b277ee676c · outbound

This paper cites Selaur: Self evolving llm agent via uncertainty-aware rewards.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Selaur: Self evolving llm agent via uncertainty-aware rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.328391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.328391Z digest=sha256:8a96545f14bb409f29df1a52b7227cf78b381a75910eebd2f5e72c818ed6b1a3

Observation 78298233-2985-415e-a37a-9f4e5254b1ca · outbound

This paper cites Self-Improving Language Models with Bidirectional Evolutionary Search.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-Improving Language Models with Bidirectional Evolutionary Search

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.464200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.464200Z digest=sha256:bb0f4d9549b43688cea4516d84e7ecea2e35c968e12b69f01d0f52a333c33990

Observation 42b75e4e-19e0-460c-86e8-ec6acc6fd764 · outbound

This paper cites Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.602393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.602393Z digest=sha256:494413cd1b4b067e91584071b82df2d59b05ecef9a49d6aef348d64fbbab061b

Observation c2e9eb84-99d0-47dc-beda-b6654d619d6a · outbound

This paper cites Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.737392Z digest=sha256:8056bdd241c4fb7ad23852176d9cf3f3388383a9b56467a461e341c0df855036

Observation 8d4c3a23-e484-4c81-82a6-2206d90e5e79 · outbound

This paper cites Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.901016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.901016Z digest=sha256:5374dd4b2bc525cb3203f0dec9131f3e3c6dbe14b5382f709f1803d23da2c90a

Observation 81c2eb88-de5c-4ea8-b3e8-4ef04ac48b49 · outbound

This paper cites Self-evolving multi-agent collaboration networks for software development.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-evolving multi-agent collaboration networks for software development

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.006298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.006298Z digest=sha256:e4b30797ade473e995b86db36e355e98534779739000808a76856cc37fd8d275

Observation fb119d41-1db2-491c-b863-2c7cd2db0aad · outbound

This paper cites SEW: Self-Evolving Agentic Workflows for Automated Code Generation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.085301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.085301Z digest=sha256:21c684f9fbd87e750dbc21deb3df18ef04bc7ae2284720841a974f22297e44d3

Observation a438e6fe-bc99-4020-aadc-cd4a1d5f5693 · outbound

This paper cites Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.179134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.179134Z digest=sha256:ae95a8f1bb3d7aa13a903849ee6eb1cc873b434ba5d4094c9460b557845ef20f

Observation 3ecd6a5f-8ebc-4011-8b8c-1248324a388d · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.248349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.248349Z digest=sha256:19f419ab2c4f7dfc08692f637d819692070e41740904eb90b965eb1fb3991d72

Observation 6a19d7cd-4a82-425a-8a93-a521749a1200 · outbound

This paper cites Evotest: Evolutionary test-time learning for self-improving agentic systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotest: Evolutionary test-time learning for self-improving agentic systems

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.318781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.318781Z digest=sha256:798195c820e31dddbf88280bf077596f7985512cbefae1cd99bf1687f8ab55ba

Observation 3a3b0b20-44fa-43b8-a733-cfb5ed07a0be · outbound

This paper cites Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.425048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.425048Z digest=sha256:f442eabf5bbc0732293afa40d9116b531b6c4e6b47e5bf33cb9632f87593b90c

Observation b76f30c6-aa60-4f0a-8914-2712ff2475be · outbound

This paper cites Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.498006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.498006Z digest=sha256:c0d08bbe26fb315094ec8bd5d1dc999f5dbdae5676c73ea34c92ae3a17ee2a80

Observation 73fce81a-b1bf-4152-9555-ee4a82465c9c · outbound

This paper cites Your agent may misevolve: Emergent risks in self-evolving llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Your agent may misevolve: Emergent risks in self-evolving llm agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.591245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.591245Z digest=sha256:472d9f9a3e1eebabb2740fa51b58d5bd63eabf176005f27bf77040149bbe7644

Observation d6b8e116-6b94-4a8e-8c9e-eafca8f43605 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.698493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.698493Z digest=sha256:60e14a8789d76e169fe1241633e5267abb5cc8964778ec9ca9c45b64c9ae7fb3

Observation 2935ea0f-764a-4a2a-8341-5d42cd473aad · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.811240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.811240Z digest=sha256:8ba35d03904da4def3a89c9d893e588faf6c83d7e24ef72c1316d098a2874141

Observation 188f85ae-edff-4a3f-aa09-91fc9df16cc3 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.916387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.916387Z digest=sha256:3788b374db1933f9c7570e26d5a1f534d371d27d5d2d26197cef26743ae000d2

Observation bf78fd0c-0c32-48c7-86ad-fc0033c497ef · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.996496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.996496Z digest=sha256:deb1370a524f442d9530c9280a5f211cb2f5de87706c24bded6ad29af29b75ea

Observation 592bab4f-14f5-43de-8234-c160ebd7d284 · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.075464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.075464Z digest=sha256:e6a3900ebe7d674d73e6bc26e7f2e029d16beef8fdb5c43a9b806598b3fbc7d6

Observation f05cad0a-6ce5-48e5-82e9-5d21a17a8a02 · outbound

This paper cites The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.158997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.158997Z digest=sha256:9ca29268e024f50075bf27b810d9eeaab42af8cf73228247e01f95f920885ddc

Observation 02c2e747-2463-4d6e-b47b-e9c0e5f07032 · outbound

This paper cites Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.214251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.214251Z digest=sha256:d6d14fb8b175805a51cd76afe976ac03b4283bbf9510b2c9d75fcd1f14b610bc

Observation 4cc7197c-1adf-4548-8a30-f846e95381f2 · outbound

This paper cites Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.287770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.287770Z digest=sha256:cca1740ed9f50bdde58eb35e2f238c045b9cce3bb4b5b8d0c43bcec7d049ed09

Observation 81771bf5-edef-4fed-94ad-808de9210a98 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.372036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.372036Z digest=sha256:30cdabce908cbd59208336f30920429e6c01a262519b462b649716c0afb1c670

Observation b6e25965-c2c8-4b22-b36c-349ab6a533c4 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.455621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.455621Z digest=sha256:46445770db94af529796921ece7324264fffac9f9455862af61de50cebbc8684

Observation b47a9d7b-74e2-4fe3-8fbd-963edfbf6074 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gaia: a benchmark for general ai assistants

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.536812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.536812Z digest=sha256:f41a4bca67dc9dcb931a6b73f91600eaced21df29fd5e51d2ed92c985f3cfa53

Observation b512849c-9191-4ddc-a38b-e8badd0f1281 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.625544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.625544Z digest=sha256:5f67f84760300a7d7ebff7a26d1ba1d0ac48785f68e7195f1d49e03f5194ae1b

Observation 6f0dc2de-6651-456f-a1aa-1c902b2c9a9f · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.705418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.705418Z digest=sha256:b00fdcf10ecb445d6b2fcf6df9588a1e3ad4af2be2f9c984c46ba7579efa3a9d

Observation 9f015896-0a35-463c-a30a-23394db146bc · outbound

This paper cites Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.779498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.779498Z digest=sha256:b9261c1f622286b57741f23c3529b4290069080ed9ffc0000dcb4f080a528de2

Observation ba0c0772-e8a4-4962-8322-b8fd75fcc637 · outbound

This paper cites Agent workflow memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent workflow memory

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.857487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.857487Z digest=sha256:cae341da28e28a12663e0de902507b59d3f709bea2b6cee3caecaca81a52bf54

Observation 62aa1b8c-60ad-45aa-9b9b-2d5cc21f14e3 · outbound

This paper cites none identified.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? none identified

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-04T01:12:38.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.939920Z digest=sha256:c7f23a2dbebf22f8638e57e09ca8de3cb5fb8d5c7852978f0162876627998e0e

Observation 3377049d-f4bc-4c76-8fb7-557fa03b9c9e · outbound

This paper cites reasoning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? reasoning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.018297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.018297Z digest=sha256:8c34adac91e48398243c9683d59c081c9155e50d880d4cb12a4abbf8f045232b

Observation e2a5d1a1-0e80-419f-9eef-e66ae93408dc · outbound

This paper cites Order from most to least important.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Order from most to least important

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.092541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.092541Z digest=sha256:27520c6ce58d18d6b00048beeeb18b6694be21982d5ee29fb2c93f0ab324a751

Observation f10004a5-26a9-4d39-95d1-277af5e60038 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.153063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.153063Z digest=sha256:90e9fd0f2f85b9815b3730a879391728ecfd660b815fdeb0cea1f97e768087dd

Observation 90dc9e6e-7d6c-4c87-8624-ff0085958f7c · outbound

This paper cites Status:.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Status:

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.230523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.230523Z digest=sha256:9849e65904b338ff7045e5fa3355591fc1ea4c0623bc1559c0556f60826034fc

Observation 53aba8b1-c42c-45e0-9c21-a9a2afb989b4 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.310212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.310212Z digest=sha256:9ecb8ea1d54f0f2e208e1b6f6ace718f417ce49cd694bd0e559aeac16bd414ba

Observation 09cd7c8b-48af-4bba-ade3-477b10cbe395 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.419533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.419533Z digest=sha256:3807726f198efb3234fc32157946ba92ee9062a257f2ca78e22b96efc80a0e71

Observation 00a09f38-aa8c-413d-a4b4-396c781e0e26 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.476352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.476352Z digest=sha256:89b43dfdfffc33bdd525c8ce29bb378d9c1a3af910b685f4e47562a0e8f262af

Observation 6984840c-21ed-47f3-a1db-0713b084276d · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.559055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.559055Z digest=sha256:306dc5adc1ec04e7d963a599805a4854d980146f5603aa9d9316730ae2c848b4

Pith citing papers

No inbound Pith citation observations are available.