Pith. sign in

Paper Citation Record · LEDGER

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

As of 10 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 0 inbound Pith citation observations for arXiv:2608.00155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00155 v1

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:12:39.559055Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b611d2-1277-424a-96f0-bd6abc117c0c · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.678535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.678535Z digest=sha256:84f4b813edb572641fb4e5c2c4cc7436acbf386afef6771c9496536be78728dc

Observation 15c3b5f3-0cc4-407d-97dd-53e6f149a015 · outbound

This paper cites Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Position: Agentic evolution is the path to evolving llms.arXiv preprint arXiv:2602.00359, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.740362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.740362Z digest=sha256:f8e2067650680d32bf104be9a3b764487b8062e01f011cb8fde43d6e389d9169

Observation 774e17c1-81ba-4c06-8184-c3942a218a0c · outbound

This paper cites A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.844983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.844983Z digest=sha256:22b46ee7fe17643dccc84430a13ed76b0b2f3f4ee2c552f28f3786852575c1ff

Observation f6d0b54b-6b3b-4420-ba41-a3bb96d97d7c · outbound

This paper cites Agentic context engineering: Evolving contexts for self-improving language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic context engineering: Evolving contexts for self-improving language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:30.945741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:30.945741Z digest=sha256:3aa2f65bd596aa34cd82f86a886cdb0ed62b9720bbfa14ad42060710cbdcc75b

Observation 4b30595c-b4bf-4321-82f1-c1ac004ba677 · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.047941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.047941Z digest=sha256:489017180138e8e9ccb232c2963a9833c7c2aa4c384a135b386c87f0460a0577

Observation 4d1cbd62-4880-457e-9d8e-76e4c51ce49d · outbound

This paper cites A-mem: Agentic memory for llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? A-mem: Agentic memory for llm agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.168633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.168633Z digest=sha256:6073d22bc8c7dedde578e97d2450f87ae4bdf410cc2aa950363255d3eb4df195

Observation 5b4a68ac-830d-4563-8af2-bc2c42882100 · outbound

This paper cites MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.238737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.238737Z digest=sha256:289c5e6d6d45cb8e315dfb94cc51ed0df046d07d417b9069b638583a7865c8f6

Observation 9997acb5-415b-4035-aced-f5f12b482392 · outbound

This paper cites Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Autoskill: Experience-driven lifelong learning via skill self-evolution.arXiv preprint arXiv:2603.01145, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.339997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.339997Z digest=sha256:06e1e260fabb38821904d1a779c8d64ab2997802a4ae86628f563ed79cc9dd4a

Observation 2ae5e5f3-dae5-48a1-9f88-e580831d52be · outbound

This paper cites Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.448963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.448963Z digest=sha256:b00b4e2b51a0e128c739f3d9da4ee8a0711f991b043b1489d868bf5221ecf6cf

Observation 669b2ad1-81a2-4075-bd74-1379c5d0b903 · outbound

This paper cites Memento-skills: Let agents design agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Memento-skills: Let agents design agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.553680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.553680Z digest=sha256:8161437c55bf941b0edf92c25c79eae2865d58aeddbb1c25cdf4950787b3e526

Observation aff5ba51-e69a-4fe0-958c-0ca5c67e030e · outbound

This paper cites Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.613003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.613003Z digest=sha256:012f714bf983c5872cab6498b1c11a1f87e76fe3dc47102c7b29feba0f13275d

Observation 8650ca55-dd31-4e2b-b6c2-2e8233c3b8e2 · outbound

This paper cites Appworld: A controllable world of apps and people for benchmarking interactive coding agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Appworld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.709243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.709243Z digest=sha256:380276e5e37876c4a9fba77ea854f383d8de25ffc1f0bd1e55c5e5fd141d8a1d

Observation 0313bafd-50e0-4c88-95b5-24fa3f54c936 · outbound

This paper cites Gonzalez.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gonzalez

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.826496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.826496Z digest=sha256:b9ac1efa5f9daf3eec9a334dd31f810ad11e366f93a650afc7378ecdbf51d10b

Observation 192ac56d-3eb2-4e4d-966f-b596268400d9 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In Proc.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Swe-bench: Can language models resolve real-world github issues? In Proc

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.918060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.918060Z digest=sha256:0a60b9fc701904117bb7d9306c28c3f6e681727588c6ab0f6e8136acdbe45914

Observation c22fc5a9-a231-4699-b27a-21e74d18ae93 · outbound

This paper cites Humanity's Last Exam.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Humanity's Last Exam

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:31.996592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:31.996592Z digest=sha256:c082b3d6d462b5e8f1f853e3ca8ac35dad8818ca1abb281605fa3948e4856453

Observation 45ed4434-2e42-4aac-a302-79bb5630afb5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.087301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.087301Z digest=sha256:a0161d9fd11cfe4e95c255dcf13672316a270c00ee732ab088cfff0bfeb933a5

Observation 0ab7ed43-9bc6-413b-8b5e-082f746f58f4 · outbound

This paper cites Stream- bench: Towards benchmarking continuous improvement of language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Stream- bench: Towards benchmarking continuous improvement of language agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.156197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.156197Z digest=sha256:8848bbeb0142be401bdf015d1a60bb67b073067e99c9198b441f5f4ec2dc1ac1

Observation 19de0ee3-c2cd-4dac-84f9-1710b859a36c · outbound

This paper cites Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.201860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.201860Z digest=sha256:42b52d1d18a1203cc55d971b0c28d82fa756e43600a28bdb0a554f78556f37d1

Observation b8cc89d3-74d9-495a-a8d0-64d85f8c11ec · outbound

This paper cites OpenAI GPT-5 System Card.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenAI GPT-5 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.267699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.267699Z digest=sha256:f6e05ee26ecbdd4d73cdbff90e4a75dddf530605cd76d2ced30bd5fcd9e882fc

Observation 5ba34156-2f11-4e2f-be16-4be50f196bea · outbound

This paper cites Gemini 3.1 Pro model card, February 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gemini 3.1 Pro model card, February 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.359163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.359163Z digest=sha256:f684ae23aef1ceb29a1ee9db134b591fb441b9d003be141ccd404ac758c1b388

Observation b9ba58dd-a90c-4480-9b54-3e6fb70502a1 · outbound

This paper cites Introducing Claude Opus 4.7, April 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Introducing Claude Opus 4.7, April 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.441241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.441241Z digest=sha256:732ed95966bd379ec16e3dbf26aa93ec0f9f9d1a658df457975fa52615191152

Observation 1e4e000b-6c7e-4e9c-b814-b4951cdeda99 · outbound

This paper cites Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Browsecomp-plus: A more fair and transparent evaluation benchmark of deep-research agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.517774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.517774Z digest=sha256:79b20dd6cf22d69d60158e3fa1416bdbaaede1ba4d0248c20c1b1381f4352549

Observation a9c9edd8-577d-471a-81e4-899bf8f7293c · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training with self-supervision for generalization under distribution shifts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.557621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.557621Z digest=sha256:38620bb301db21ec9ff14cdb5ff57a9a90862618c96ba9a2e4e3bb714f208cff

Observation 1ef7ec98-9596-452d-b766-b5a37395abac · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.616665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.616665Z digest=sha256:823ec3c17b1c52a01ce09cd95f6f70e7167fb7dcbb311d3b03357e91f0a21abf

Observation 4ec2d7e4-606f-45df-bcbf-08412cdd5a1a · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13):3521–3526, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.706158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.706158Z digest=sha256:eb3d97b17a73eb6ab50e9b368e640a971472c556668b5041cf419886153e54a6

Observation 7128a647-c5e2-4fc7-bee4-44afb379adaa · outbound

This paper cites Gradient episodic memory for continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gradient episodic memory for continual learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.785239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.785239Z digest=sha256:a316bc71e67c8cc8678e8acad2977741b130b9cc7cde43a76fb2cf5d61369cc6

Observation 13b29ec8-639f-4f57-9e71-ec2d6c5b41ed · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Scaling llm test-time compute optimally can be more effective than scaling model parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.834609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.834609Z digest=sha256:fa3f593e41157f391fcf9fdceffbf46d964acd3e11e43304d048c9a76fccefee

Observation ea24b3aa-c55c-429d-8f36-4f6fd65c80e7 · outbound

This paper cites Test-time training on nearest neighbors for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time training on nearest neighbors for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.907311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.907311Z digest=sha256:ea851c0a790317eed2c870e7e37174a97b0adf1dad03930c3fbb671108f433f8

Observation 69b76d30-5c03-4755-a817-87bedd5539bf · outbound

This paper cites Efficiently learning at test-time: Active fine-tuning of llms.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Efficiently learning at test-time: Active fine-tuning of llms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.947220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.947220Z digest=sha256:3ddc22034b488a4692d4b2028a88d6e18ee5d976de909f33b128f834d5d37fc9

Observation 8353d4e8-7813-4986-b32e-8b3fd29c1545 · outbound

This paper cites The surprising effectiveness of test-time training for few-shot learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The surprising effectiveness of test-time training for few-shot learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:32.990733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:32.990733Z digest=sha256:c80b8148548cad3c5eeda6af171e3d91d64eff246af329df74754b66ae2377ab

Observation 1aaae17b-f559-45b8-b4aa-042734c11429 · outbound

This paper cites In-place test-time training.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? In-place test-time training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.052856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.052856Z digest=sha256:672eae99d72f595fe23f984a584c0c88216e91e5bbc23a9769050db3730b0b74

Observation 78bb23fd-0340-46e7-a979-5c9cb6c60f57 · outbound

This paper cites Test-time adaptation for llm agents via environment interaction.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time adaptation for llm agents via environment interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.147161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.147161Z digest=sha256:5a2eb212fbe5d738e21fdd6b4094dd96adab82158e563bd81ac7d04bd9d91821

Observation 8c6d34b6-56e8-42fd-be9c-40b104d6fb5b · outbound

This paper cites Test-time learning for large language models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-time learning for large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.266388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.266388Z digest=sha256:b0107d6bc328cafa1e3fb805fa421e44749d13f0e6ee9e6ce2313f1e35a90c21

Observation 31c98bfe-fc30-4da1-b482-26138dfd8ce1 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttrl: Test-time reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.347224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.347224Z digest=sha256:4e5980c2338e382ed194ddcd856dd9f604324b6627fcd90c830a48b32ef7de34

Observation 560c579e-c5e5-4904-893b-432431fa2cb1 · outbound

This paper cites Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learn- ing on the job: Test-time curricula for targeted reinforcement learning.arXiv preprint arXiv:2510.04786, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.407176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.407176Z digest=sha256:0d029c0b503c7cd9fa48d2368cc8a67b3bd9ddc33a24dc019645c4c5f4c0efb6

Observation 729ca982-0b5d-48a6-bc45-3aa1aa448f4f · outbound

This paper cites Learning to discover at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to discover at test time

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.486960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.486960Z digest=sha256:93c42da4161a73c39e92e5589f12185480158a6c3319b6895b75dfabdd9fa20b

Observation 0dbd5315-bd55-49ce-a095-a6213b44dc44 · outbound

This paper cites Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Collaborative multi-agent test-time reinforcement learning for reasoning.arXiv preprint arXiv:2601.09667, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.584703Z digest=sha256:f4f7f4e45a2fd1a6ffc16131a58cfdaf62af14f3167569ff5dd1b2ce1bd193b5

Observation 8a992a21-8df4-4314-91b8-8d6506facb9c · outbound

This paper cites What if consensus lies? selective-complementary reinforcement learning at test time.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? What if consensus lies? selective-complementary reinforcement learning at test time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.672842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.672842Z digest=sha256:19a88746e96738c742e75f8405c016987faf2c0a28fba72aff4c82449bce3b04

Observation 0cb42fec-7ea7-47d8-add2-4a5c603aa6c1 · outbound

This paper cites Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttsr: Test-time self-reflection for continual reasoning improvement.arXiv preprint arXiv:2603.03297, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.783687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.783687Z digest=sha256:1dd0d9bbb029d29e2e949adbcb446574f551820908a170ffda92d1b484c0e19e

Observation d76d5b23-7867-48c6-ba76-c7b3adbf5fd3 · outbound

This paper cites Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Ttcs: Test-time curriculum synthesis for self-evolving.arXiv preprint arXiv:2601.22628, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:33.899505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:33.899505Z digest=sha256:d5a2f9f2fd1adbfbf18c2819294e1f3a5c99bbf30d846f8c22aa30ece967c4b9

Observation b01a108e-71ff-4206-a9bc-190f1a4e571e · outbound

This paper cites Test-Time Learning with an Evolving Library.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Test-Time Learning with an Evolving Library

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.009288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.009288Z digest=sha256:6e069ac6d3aa27fa8e9b86cd94093d5aecee8a41d1dc0da6237924733099fa61

Observation eb9d85b7-402a-49fc-aec2-7efb0b250981 · outbound

This paper cites Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.114246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.114246Z digest=sha256:334c8fcdd0ac12ef03df73bf459021ebec739e04ae61fa7480078f61077d723b

Observation 4b2262f2-6162-4fac-abde-a3da7f3f3208 · outbound

This paper cites Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tarse: Test-time adaptation via retrieval of skills and experience for reasoning agents.arXiv preprint arXiv:2603.01241, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.226463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.226463Z digest=sha256:f98a4fea96d20549469792ac767d1c3c8954875fe3cd54d9705252e6abb5928f

Observation 26fb1c9a-50dd-4827-84ac-1aa5afaa48dd · outbound

This paper cites Agentic plan caching: Test-time memory for fast and cost-efficient llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agentic plan caching: Test-time memory for fast and cost-efficient llm agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.337320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.337320Z digest=sha256:e5bbc5ffe1270d96d79bf8384ade8f0102d25b170e47947de17c7551b9cc02c2

Observation ed2c9bfc-f745-4cf6-b322-a65549161db3 · outbound

This paper cites TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.442077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.442077Z digest=sha256:5b50bbf0916463829921c36f473595dafe7764ca43efbd876654f5d1121a9724

Observation 18107843-204b-4eb5-b983-e6c2028238e2 · outbound

This paper cites Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-improving llm agents at test-time.arXiv preprint arXiv:2510.07841, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.540281Z digest=sha256:9dbe1213ce2dde5d9338d077ddd0e225c7deb2cfeb10d41dc55c35333ac35492

Observation b09b02a8-d8b8-453a-b705-88ffabf8a7c3 · outbound

This paper cites Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Just- in-time reinforcement learning: Continual learning in llm agents without gradient updates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.652568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.652568Z digest=sha256:92a885120fc7b377e1534fe0ba4bc66379bad0a95a2cfdd0e1fbe49d08256b90

Observation 9f5ea4cc-4e98-4228-9cbd-eb3796899c71 · outbound

This paper cites Panini: Continual learning in token space via structured memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Panini: Continual learning in token space via structured memory

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.761418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.761418Z digest=sha256:d245b01bf3b85b802453fc1c0040601905838e1b59f856e16085d2437fa85cfa

Observation 5a6a2359-872e-42b0-bcc8-368eb0a6533b · outbound

This paper cites Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.866944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.866944Z digest=sha256:983914ddd3b3f0f06eb0b7647a0e61d5fa15d1b5b87ccec73822e95054692af7

Observation df12cecf-038a-4dc1-8a2b-c5be36a6d665 · outbound

This paper cites Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mssr: Memory-aware adaptive replay for continual llm fine-tuning.arXiv preprint arXiv:2603.09892, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:34.939316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:34.939316Z digest=sha256:9841a122366631aa710506746f8e56d3f7e123810fce2c81561172a4adb1bfb3

Observation fec8e04f-7212-49fe-a4ed-8b154e68aa71 · outbound

This paper cites Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Learning to continually learn via meta-learning agentic memory designs.arXiv preprint arXiv:2602.07755, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.017911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.017911Z digest=sha256:3de6b14607992f2ca50b94d6df3e0177305342dd4aa94de95f6a75a95af41eb4

Observation fb473b65-f10e-480b-bb9b-04a615958777 · outbound

This paper cites Xskill: Continual learning from experience and skills in multimodal agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xskill: Continual learning from experience and skills in multimodal agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.087726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.087726Z digest=sha256:01d197ae92bf0a56294a4638490cd12f8e7a32149e833eb7eeacef6cfb4cfa07

Observation da43f976-c6ce-463c-9799-361ce994f6aa · outbound

This paper cites Online Experiential Learning for Language Models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Online Experiential Learning for Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.159727Z digest=sha256:4b43388598708a126292e704af853bc49278908c9ab6a66bf657df33ed7cfb91

Observation ddc78d07-c649-491c-9b06-7196fe1fb0e7 · outbound

This paper cites Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Adaptive collaboration with humans: Metacognitive policy optimization for multi-agent llms with continual learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.227678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.227678Z digest=sha256:86c769f26547fda6997a69a06e14e89923b77bb3d5095d0b32eeff2e90793609

Observation d42845b4-b90e-46aa-b1b5-2dbf86837931 · outbound

This paper cites MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.336093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.336093Z digest=sha256:644e7de876d5b42ad0f065c688789cdbb58a03a2db87601c475b25cdf63d7bff

Observation 227b17f5-290a-48c0-94bd-9da4a8674c98 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.459460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.459460Z digest=sha256:5df54bb739e537e5a81024e3fd24135b6233d2aa28b99cdd5e0095f8650e3864

Observation e6271b35-7585-4cff-ac8d-17152f2132e8 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.579166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.579166Z digest=sha256:0cadd347c2004c4430b66a711192e626d454cf1b1b37eac36bac7daf5ad32675

Observation 1fcce961-c4c5-4a04-8cba-613fc43b0b43 · outbound

This paper cites SkillOS: Learning Skill Curation for Self-Evolving Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOS: Learning Skill Curation for Self-Evolving Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.690973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.690973Z digest=sha256:4f7f60a11608bab86d5d01f78cd9848efafd18b325e0214ce3c8ec4c4c957420

Observation 60d83db8-2aa5-4635-b235-d2f7d73e311e · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.780877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.780877Z digest=sha256:990287ecfc8236bca33d50f7ef015f871af76ade3f30afe399f68df1465ee229

Observation 77a4cd23-7cf0-42d9-bfbc-c364071b1d2d · outbound

This paper cites OpenSkill: Open-World Self-Evolution for LLM Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? OpenSkill: Open-World Self-Evolution for LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:35.900154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:35.900154Z digest=sha256:255dca5a0bfb80d4f2960399d0a55baf44e99892cbf8945e5ea3c143070277c8

Observation 3406787a-4e71-4b33-9ab7-b8a54d9deec6 · outbound

This paper cites SkillOpt: Executive Strategy for Self-Evolving Agent Skills.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.013114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.013114Z digest=sha256:acc955256b997f54efdef7fee39fb27a17409a86b3db1ba8aaa83b50b37e627e

Observation 9f8206b5-8e87-4d2a-82ea-3561f5a75200 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.133344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.133344Z digest=sha256:014ec49c6fb0bebc52ab2aff8e69bb4c1deda0e1d961c27b5a253815dacabdb7

Observation f4deb2ae-ff8d-44e6-a47e-4bf4271f2963 · outbound

This paper cites Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evoconfig: Self-evolving multi-agent systems for efficient autonomous environment configuration.arXiv preprint arXiv:2601.16489, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.247731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.247731Z digest=sha256:ddbc163a6d868f19fe1b24155cda1113c107c81c784e5f05ed954e6a038b2b74

Observation fd65faba-7aaf-4848-ae01-57b277ee676c · outbound

This paper cites Selaur: Self evolving llm agent via uncertainty-aware rewards.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Selaur: Self evolving llm agent via uncertainty-aware rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.328391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.328391Z digest=sha256:b5d9a0b0387257c7db666cb84a93340162a9d259ec76da7db5667f4618b09f73

Observation 78298233-2985-415e-a37a-9f4e5254b1ca · outbound

This paper cites Self-Improving Language Models with Bidirectional Evolutionary Search.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-Improving Language Models with Bidirectional Evolutionary Search

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.464200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.464200Z digest=sha256:00d9220647bec3f30545688e1530304f2cd818a41ed73bdb32d4a1ae46945adc

Observation 42b75e4e-19e0-460c-86e8-ec6acc6fd764 · outbound

This paper cites Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.602393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.602393Z digest=sha256:d9a7ef248048a4a849419db6c4eeddbda887e16d87d37e23dd6e39005e2c5e91

Observation c2e9eb84-99d0-47dc-beda-b6654d619d6a · outbound

This paper cites Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Rein- forcing chain-of-thought reasoning with self-evolving rubrics.arXiv preprint arXiv:2602.10885, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.737392Z digest=sha256:f4521579b75f8d95ac936a135b10ea1cecd83b3bd9240100cf2fb6941f27945b

Observation 8d4c3a23-e484-4c81-82a6-2206d90e5e79 · outbound

This paper cites Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Metagen: Self-evolving roles and topologies for multi-agent llm reasoning.arXiv preprint arXiv:2601.19290, 2026

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:36.901016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:36.901016Z digest=sha256:08460f0ca27fcf513b335172bfcf43fe3a0bf566f4772ff8f897a9421127674d

Observation 81c2eb88-de5c-4ea8-b3e8-4ef04ac48b49 · outbound

This paper cites Self-evolving multi-agent collaboration networks for software development.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Self-evolving multi-agent collaboration networks for software development

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.006298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.006298Z digest=sha256:55cee944677071d7514f68c04add6658f5a85c44189a89709f7cdc733ef0f982

Observation fb119d41-1db2-491c-b863-2c7cd2db0aad · outbound

This paper cites SEW: Self-Evolving Agentic Workflows for Automated Code Generation.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SEW: Self-Evolving Agentic Workflows for Automated Code Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.085301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.085301Z digest=sha256:65f5f011a1e95cf4220b98c00b93807b78bb6854cb772d1fdc6d62ac66cbdb2c

Observation a438e6fe-bc99-4020-aadc-cd4a1d5f5693 · outbound

This paper cites Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection.arXiv preprint arXiv:2603.04900, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.179134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.179134Z digest=sha256:32b0b699266db833003fa2406babde4d2ab232342def20301ed1039c51c08f8f

Observation 3ecd6a5f-8ebc-4011-8b8c-1248324a388d · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.248349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.248349Z digest=sha256:8191519f3616a305becff104dd6650ccf81f4c38bbf01b09ecdeb6abbd632c95

Observation 6a19d7cd-4a82-425a-8a93-a521749a1200 · outbound

This paper cites Evotest: Evolutionary test-time learning for self-improving agentic systems.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Evotest: Evolutionary test-time learning for self-improving agentic systems

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.318781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.318781Z digest=sha256:3a322347f0393b86762444fe4f3d8a14404e43aed66fd39df2d1323c5e49bf4c

Observation 3a3b0b20-44fa-43b8-a733-cfb5ed07a0be · outbound

This paper cites Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Building self-evolving agents via experience-driven lifelong learning: A framework and benchmark.arXiv preprint arXiv:2508.19005, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.425048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.425048Z digest=sha256:9c2dae669b9b5119372ae0b31c7c0aebf03d0b19a2302f72818ca01a4a4d7e33

Observation b76f30c6-aa60-4f0a-8914-2712ff2475be · outbound

This paper cites Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Optimizing generative ai by backpropagating language model feedback.Nature, 639:609–616, 2025

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.498006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.498006Z digest=sha256:8cb6c539f252775b429d428a7d2727d6a62e6b5e37a803cfd215904bd326a582

Observation 73fce81a-b1bf-4152-9555-ee4a82465c9c · outbound

This paper cites Your agent may misevolve: Emergent risks in self-evolving llm agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Your agent may misevolve: Emergent risks in self-evolving llm agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.591245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.591245Z digest=sha256:58226d92f5d6cb0a80956d55e329bd381bec6403b38bf4c913fbaf0b93298e42

Observation d6b8e116-6b94-4a8e-8c9e-eafca8f43605 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.698493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.698493Z digest=sha256:054a2b0e3ca740dc20ae139585370b2cf09032ac41fba6834d950db067775b4b

Observation 2935ea0f-764a-4a2a-8341-5d42cd473aad · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.811240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.811240Z digest=sha256:7d3f23c74cb256cfbe6b78c526c7630b33c043ee5ad94a9705e7657133186825

Observation 188f85ae-edff-4a3f-aa09-91fc9df16cc3 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.916387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.916387Z digest=sha256:4aeb4f1ffdb5bd1eb26a5626d5df90c615d1c592f03d4887f6e19849cf0d21d1

Observation bf78fd0c-0c32-48c7-86ad-fc0033c497ef · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:37.996496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:37.996496Z digest=sha256:1282f63a041f77a44220da7c3cf8982ff492dd99312274e479af63e29951fdfe

Observation 592bab4f-14f5-43de-8234-c160ebd7d284 · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.075464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.075464Z digest=sha256:343e424fac0e7758663e1457f2f9bc4fb67fe4e59299448a9b025b8f5e6d592f

Observation f05cad0a-6ce5-48e5-82e9-5d21a17a8a02 · outbound

This paper cites The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.158997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.158997Z digest=sha256:87a4200b2168ce532f49dd3d5e4cfe555b96fa3b829194ff79952211305317f9

Observation 02c2e747-2463-4d6e-b47b-e9c0e5f07032 · outbound

This paper cites Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.214251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.214251Z digest=sha256:a7cd5f2325bf8072f1f16a2a496fa5c4719ae9b184565425b6160f67cee35cbc

Observation 4cc7197c-1adf-4548-8a30-f846e95381f2 · outbound

This paper cites Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Cybergym: Evaluating AI agents’ real-world cybersecurity capabilities at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.287770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.287770Z digest=sha256:6a88bfcfc434b398302d74d7b7b5f6caaab72155cb637578768e266252fe482b

Observation 81771bf5-edef-4fed-94ad-808de9210a98 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.372036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.372036Z digest=sha256:e51e34edf3d1397773dd323b93b3d657205b1714888b523af65458b10a93c43d

Observation b6e25965-c2c8-4b22-b36c-349ab6a533c4 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.455621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.455621Z digest=sha256:746941a071f24176e4ad3868375b47a295e6fe39df8a1da34ac995ef4472c14c

Observation b47a9d7b-74e2-4fe3-8fbd-963edfbf6074 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Gaia: a benchmark for general ai assistants

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.536812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.536812Z digest=sha256:7cff75eecdedc4e3f88bc10379803474cf29211753de7c93b43fe2f71ee3c1d3

Observation b512849c-9191-4ddc-a38b-e8badd0f1281 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.625544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.625544Z digest=sha256:536cf64996f23551adbb2a39709399b96e68611bf3bd1f1db6495be91a304507

Observation 6f0dc2de-6651-456f-a1aa-1c902b2c9a9f · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.705418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.705418Z digest=sha256:a13721b28f2f1bdd330538a396dc07a4904226e463b21c1af547fec625e46c9d

Observation 9f015896-0a35-463c-a30a-23394db146bc · outbound

This paper cites Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.779498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.779498Z digest=sha256:7ff23ca049e643e6a7e73272a64caf210b67a894bf6d5e61be6a2622a9d34a8d

Observation ba0c0772-e8a4-4962-8322-b8fd75fcc637 · outbound

This paper cites Agent workflow memory.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Agent workflow memory

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.857487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.857487Z digest=sha256:f081e044a388383b6fda71dbb4589f01df1ec3afb64500ac584f2876e80d08a6

Observation 62aa1b8c-60ad-45aa-9b9b-2d5cc21f14e3 · outbound

This paper cites none identified.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? none identified

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-04T01:12:38.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.939920Z digest=sha256:8fc1283a3d25332381b8c5a86553899739a1aa2ce35e40aca0e4c853c9700a1a

Observation 3377049d-f4bc-4c76-8fb7-557fa03b9c9e · outbound

This paper cites reasoning.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? reasoning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.018297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.018297Z digest=sha256:d7fda506f50f0c30b6069011a58ecf033042bf04bf8cf74d19291bb69db7ef42

Observation e2a5d1a1-0e80-419f-9eef-e66ae93408dc · outbound

This paper cites Order from most to least important.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Order from most to least important

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.092541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.092541Z digest=sha256:2db56d166e65ba7da9631e3ac387cc8bb7482b135914119d1123a5395ecfe113

Observation f10004a5-26a9-4d39-95d1-277af5e60038 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.153063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.153063Z digest=sha256:0d4b2672f8d1cd0c4bd16e36325b9778bbf9527d514b5f7bcbb6c9e514933069

Observation 90dc9e6e-7d6c-4c87-8624-ff0085958f7c · outbound

This paper cites Status:.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Status:

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.230523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.230523Z digest=sha256:38f49287fb973717305fac9f42a2edbd32653fbad92697a95b12241c415c32c6

Observation 53aba8b1-c42c-45e0-9c21-a9a2afb989b4 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.310212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.310212Z digest=sha256:e6de1a8decefb80212ac8d310bef1e6709dadbd1b5e2d58f1e261dab5c168b9e

Observation 09cd7c8b-48af-4bba-ade3-477b10cbe395 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.419533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.419533Z digest=sha256:ec46fe8c39517987731f01fd1e6ca3d8e0a756f9c3c20ba2f943b5b7896006d1

Observation 00a09f38-aa8c-413d-a4b4-396c781e0e26 · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.476352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.476352Z digest=sha256:8a2546434734ecd31931d9fdfa7dfc9862487cb75470192409a63f315d22a46b

Observation 6984840c-21ed-47f3-a1db-0713b084276d · outbound

This paper cites an unresolved cited work.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:39.559055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:39.559055Z digest=sha256:443a4fbc81f358f46f37e0580b97e7338e5192f472daa62b70888b49b3d17447

Pith citing papers

No inbound Pith citation observations are available.