Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:50.356815Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 100 of 124 outbound references and 8 inbound Pith citation observations for arXiv:2506.07976.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:50.356815Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:11.770737Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T11:56:55.498403Z
100 of 124 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c63f7383-6871-4a1f-9e9f-a6ca34e0960a · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Webvoyager: Building an end-to-end web agent with large multimodal models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ad165f-ce4f-4875-a54d-190c4ef15920 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e934bc-4905-4d1a-bff2-1ce860e60c16 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b960abc-246f-4b64-8ca4-22f1d4193945 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing operator, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb1e39a-110e-4f02-ac44-7521796337d1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Browser use: Enable ai to control your browser, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708cb064-5c0e-47b6-a6ba-913ce2ae3efb · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Cogagent: A visual language model for gui agents, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695a43d4-9e00-495d-b373-8ec44cebb11b · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Your code’s new collaborator, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3845aa36-fb6e-4bd0-93e2-223a6e589eda · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53650ddf-b229-4629-83b1-d4d6810ccd85 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea17f895-9845-4de0-9566-7df6100d870f · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Emergence of Pragmatics from Referential Game between Theory of Mind Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed5a7012-8a04-4d51-a993-5f10dc643ac1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Iterative teacher-aware learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13cb217-ff8e-47c5-a4bf-d67e8501f29f · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16608187-a392-4131-8883-d20194c97ac9 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ff3fcd-35ba-4d17-90e2-c390eab1f002 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1729e8e2-2550-4430-b0d3-2403302b1346 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mind2web: Towards a generalist agent for the web
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba8c202-d633-4d52-b917-2d5fb310d650 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Dery, Corey Staten, Mikhail Khodak, Graham Neubig, and Ameet Talwalkar
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a910c5-f1bb-4445-b180-046e47fffb87 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Tag-llm: Repurposing general-purpose llms for specialized domains, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4c3a1a-8f5e-4dab-8277-e5e7ef31774a · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ce8759-e329-4903-ba95-328e5c1928d3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d4bed2-1749-4bfe-ac01-17c9343ec65b · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70197658-5e2a-4f7a-9cbd-a67f16006c06 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ef3580-6878-45eb-a35f-ac20acc09bae · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e782046-7782-462f-9b1e-06e866525791 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33019fc2-cecc-4746-977a-048adb9447fe · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Claude takes research to new places, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7220bd2f-52ce-4fb8-8113-eee5b8cc2c5c · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing deep research, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b7f33b-8613-4bcc-b7c5-3272ab79e734 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gemini deep research, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c6c100-0c31-4f6e-bfb5-bca2489afee1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction s1: Simple test-time scaling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad2dafa-1b85-43e4-a524-8efb16084ca3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258a1186-09a0-490a-92f7-cce3cec100c2 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference-aware fine- tuning for best-of-n sampling in large language models.ArXiv, abs/2412.15287, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2fabcf-f09e-47ba-911f-3879eb3b7a72 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c762179-83b1-41c5-af07-945c628de556 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Webglm: Towards an efficient web-enhanced question answering system with human preferences, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ef7536-7bf3-44c5-9e0d-e20627c4c691 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Multimodal web navigation with instruction-finetuned foundation models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f84bef39-2991-40df-b917-33be9f082d02 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4693a926-9a90-412a-855c-ac7ae680c583 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Multimodal Web Navigation with Instruction-Finetuned Foundation Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e73f28-b6c8-4382-92c3-ba05488d486c · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction UPS: Efficiently building foundation models for PDE solving via cross-modal adaptation.Transactions on Machine Learning Research, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ec6c1c-3e1d-450e-ace3-c8a49551fcd0 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc420f93-efe5-43e7-ae33-b33d2b5899f3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210c2e1d-ef26-42b9-b227-58920dcdad78 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction CAT: Content-Adaptive Image Tokenization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31994f1e-3fe7-43c4-8214-0e7b2ae186fb · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Specialized Foundation Models Struggle to Beat Supervised Baselines
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36464d1b-0192-4e76-a387-d80269b47128 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Beyond Browsing: API-Based Web Agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1673a028-4cea-4639-9415-9c471ed362ed · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Codepde: An inference framework for llm-driven pde solver generation, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec5d41f-f833-48d6-b881-eaf8ef0aa3f1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mathematicalreconstruction of patient-specific vascular networks based on clinical images and global optimization.IEEE Access, 9:20648–20661, 2021
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6afcc4d2-0016-432e-894f-dc72072953cb · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Autonomous Evaluation and Refinement of Digital Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1c94fc-bec5-4f24-97dc-837b8fc42c2f · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, and Yu Su
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b1add7-ab27-4766-a8d4-f37558244ef0 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction NAS-bench-360: Benchmarking neural architecture search on diverse tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f94752-a38d-4e18-a4c9-34f70b9535a6 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Efficient architecture search for diverse tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d07b871-658c-438a-bac7-013d87e93ae1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction GPT-4 Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53190f75-8c12-4061-819b-83c5ce2b16e4 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a34aa73-9276-4d55-bce1-999ed8db26ca · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction SteP: Stacked LLM Policies for Web Actions
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8804b0b2-2df3-496f-b15f-b57655461877 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing the next generation of claude, 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9b53d9-5f6b-44ea-921d-8b4f8965159e · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction L2g: Repurposing language models for genomics tasks.bioRxiv, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54d8756e-8452-4363-af5e-f212a5e09ba6 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a70ffe-bdf9-42cc-820c-2990f5b3de1d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Training Language Models to Self-Correct via Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aadaef5b-d9b0-4b7e-914c-8d70c7b2e53d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Plan-and-act: Improving planning of agents for long-horizon tasks, 2025
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0e1509-7fa7-46fc-bb9c-d5d47fae60f6 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Synapse: Trajectory-as-exemplar prompting with memory for computer control
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685318d4-30dd-4025-b6f6-9a43ff82a03c · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Tree search for language model agents.arXiv preprint arXiv:2407.01476, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71d6456-f4ba-494c-ad2b-c75f9fd7a9a8 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Infogent: An Agent-Based Framework for Web Information Aggregation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced04181-8e51-4d3f-a178-d46bc6713b93 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent Workflow Memory
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f7f655-bb64-4f40-afcd-152e367c51e6 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d456d3ae-fae2-4423-93f2-60eae900fdc8 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction BAGEL: Bootstrapping Agents by Guiding Exploration with Language
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d55730ce-385f-4afc-a5e0-314eacca5c59 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d96973f-8934-484e-96ed-85db71ef9fe8 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction InSTA: Towards Internet-Scale Training For Agents
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676a71e2-dc54-4ff5-a7d5-6368c6efeed1 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b248dad-5bae-4169-9705-11936e0f99f3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Autowebglm: A large language model-based web navigating agent
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d7fb0f8-0099-41bb-b17b-ec095a7d0434 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbffbb6f-11ba-4bf9-ad32-37096acd9e55 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce53008-52cc-494b-86ab-20ecb88491ac · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c9bf6e-b80f-4b5b-a185-9ab64d74dab3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152980e0-619a-450c-b3ea-1fcf3e11b14d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa309089-4b99-4995-a905-15f4f8aa3190 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Training Verifiers to Solve Math Word Problems
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2ad100-cff4-41f3-bf36-373daa431c22 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20f4c7f-9d7b-48aa-9f8d-bad1bab99af3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling Test-Time Compute Without Verification or RL is Suboptimal
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091291ed-fe60-426f-8e65-4cd840c0a7cb · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851db09b-90e5-403f-b4e4-5bf910310f59 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd30e87b-425d-466d-9efe-fae298b462b5 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Doing: Agents that Reason by Scaling Test-Time Interaction Li Fei-Fei, Lijuan Wang, Yejin Choi, and Manling Li
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4440afe6-8dfc-4d5f-abc1-c2c0225e657a · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ReAct: Synergizing Reasoning and Acting in Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ffb19f0-0140-48e7-8833-c767f0609863 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f797f3bc-1f46-4808-baba-a0518dc4ac7a · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7decc3f5-3045-4f9f-83c9-60a869edd0d9 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86bc6b7-2cb0-4312-a672-33b57d87547b · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Metaxas, and Tong Che
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77754e79-1cbf-4898-8201-c5dc351d6c47 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gemma 3 Technical Report
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e75664d-242b-41ec-8893-3ca685a33ef3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 220a18c6-fc7e-4132-af06-b6ec616c7f4b · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ce756ea-20d6-49bd-aad0-b5ff70f3b8a7 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inducingprogrammaticskills for agentic tasks
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed329a50-5655-44a7-b005-4169833f405b · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b390b4ba-9c37-45ec-80a1-172b04fe4507 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Recursive introspection: Teaching language model agents how to self-improve.Advances in Neural Information Processing Systems, 37:55249–55285, 2024
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0126c2-e4f5-4c92-9d46-635f7f59f1e8 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7697213-ceef-4bd7-8e44-b5797144cee3 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488, 2022
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2e72d6-2e18-4952-8cb9-b2e854f50841 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Curriculum learning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd4f959-4e53-4ac8-83e4-5f27f0f4ae15 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2b60f5-332e-4c81-b660-39eed0222f2d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Analysis and improvement of policy gradient estimation.Neural networks : the official journal of the International Neural Network Society, 26:118–29, 2011
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609c6309-f788-4fb4-a157-f8c10543abb8 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Policy gradients with variance related risk criteria
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f767eee-302f-4806-9f1e-e76b6adfd824 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Abbeel, and Wojciech Zaremba
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172aac69-c4d4-48ae-b132-497d208d9d69 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafc7a9e-4033-4ceb-9c39-b632f5253095 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction A survey on curriculum learning.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 44:4555–4576, 2021
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f9382b-0b5b-4fd6-91eb-98eda6e4019d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Teacher–student curriculum learning.IEEE Transactions on Neural Networks and Learning Systems, 31:3732–3740, 2017
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e863c26d-93fc-486a-968c-e70c501dd59d · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd4ebbe-7c12-42b7-a678-d5fcd0503597 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gonzalez, Hao Zhang, and Ion Stoica
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ee4768-89ff-4c53-a214-2e799dc215f0 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e80e1f36-d90b-427f-af5d-37cd4c7e0990 · outbound
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 254e48c4-098f-4121-97f6-fc85d6ece7e4 · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 78856db5-fe51-411c-96bf-e5a93e222d04 · inbound
SSRL: Self-Search Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf96ff8-46c8-4bb8-baa2-5b44b270689c · inbound
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 946e1a85-769b-4a5d-8d47-a77f080b9904 · inbound
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b221f6b0-b3c2-4192-9046-202a2fbf1df0 · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bae768d1-b8db-4710-b31f-14a59cad3690 · inbound
PRO-CUA: Process-Reward Optimization for Computer Use Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b97dea4-7a53-4863-84ca-2707b6338858 · inbound
AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29cf167a-bdf0-43fa-ae28-3af81c202ecb · inbound
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.