Pith. sign in

Paper Citation Record · LEDGER

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

As of 10 August 2026, this Paper Citation Record lists 100 of 124 outbound references and 8 inbound Pith citation observations for arXiv:2506.07976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07976 v2

Coverage vector

measured 100 of 124 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:50.356815Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:17:11.770737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:56:55.498403Z

Reference resolution

100 of 124 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c63f7383-6871-4a1f-9e9f-a6ca34e0960a · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models,.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Webvoyager: Building an end-to-end web agent with large multimodal models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.868438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.868438Z digest=sha256:e2288df0be8a919c3136314ca049b9e147b2610ebe19314aaa3fc96a02bfaf14

Observation 09ad165f-ce4f-4875-a54d-190c4ef15920 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.878800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.878800Z digest=sha256:09021df9d448c19d093f6501d371005d16b76880bff249403d43574acdcb3c9b

Observation b7e934bc-4905-4d1a-bff2-1ce860e60c16 · outbound

This paper cites Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.883427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.883427Z digest=sha256:1d843191d043e61bb686f2db3a405f678e692371c20566bafd7df78280001293

Observation 2b960abc-246f-4b64-8ca4-22f1d4193945 · outbound

This paper cites Introducing operator, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing operator, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.887573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.887573Z digest=sha256:4f58b05b6323d53e0a5c844aa94f735f362a54c796f2944f4914655c3a19c675

Observation 2cb1e39a-110e-4f02-ac44-7521796337d1 · outbound

This paper cites Browser use: Enable ai to control your browser, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Browser use: Enable ai to control your browser, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.892125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.892125Z digest=sha256:ad199e35e16c95897ea6231969454bfceee472518fd6b5378a8b7df299fe3d68

Observation 708cb064-5c0e-47b6-a6ba-913ce2ae3efb · outbound

This paper cites Cogagent: A visual language model for gui agents, 2023.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Cogagent: A visual language model for gui agents, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.896875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.896875Z digest=sha256:509bee2426da0e21cc809907b2aae7e453496cdbae54b1372b6d9ac1b158938a

Observation 695a43d4-9e00-495d-b373-8ec44cebb11b · outbound

This paper cites Your code’s new collaborator, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Your code’s new collaborator, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.901258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.901258Z digest=sha256:62bfe73c314f1db99fe8d22902da3b819ea579c42db578ba3eb3b787c6e2ffc8

Observation 3845aa36-fb6e-4bd0-93e2-223a6e589eda · outbound

This paper cites an unresolved cited work.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.906388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.906388Z digest=sha256:276bec750079d1b6f0e3281d370880a5925635cb356f4d15b11ec37563b6877e

Observation 53650ddf-b229-4629-83b1-d4d6810ccd85 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.911434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.911434Z digest=sha256:286785a8b8395ba26a912e9edd0fdc1c0d1117e5799535a71f4414a4bac1b57d

Observation ea17f895-9845-4de0-9566-7df6100d870f · outbound

This paper cites Emergence of Pragmatics from Referential Game between Theory of Mind Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Emergence of Pragmatics from Referential Game between Theory of Mind Agents

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:27:51.687826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:49.916142Z digest=sha256:6bf5efecd5206e1085b81d31ed08be62d183f18cc8f191357c4357d377efadb2

Observation ed5a7012-8a04-4d51-a993-5f10dc643ac1 · outbound

This paper cites Iterative teacher-aware learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Iterative teacher-aware learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.920980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.920980Z digest=sha256:be99963d319bad2c37df976397cba68ca1a235c2b9160b69e40a91f90806d0c4

Observation f13cb217-ff8e-47c5-a4bf-d67e8501f29f · outbound

This paper cites Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:27:51.667553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:49.926091Z digest=sha256:2bdb098c551e52d65cc6723da11ede51ea6fc06fafc4f570654a9c0fedd71217

Observation 16608187-a392-4131-8883-d20194c97ac9 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.931593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.931593Z digest=sha256:81a950358bd2a4eb172995b7beedd90fdb76d25a722d7920c7863a59a9ca56f9

Observation 50ff3fcd-35ba-4d17-90e2-c390eab1f002 · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.936632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.936632Z digest=sha256:b28638b094b9f6c57497870fd15c7d4b2429775d7db7a275a2647e9226517121

Observation 1729e8e2-2550-4430-b0d3-2403302b1346 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mind2web: Towards a generalist agent for the web

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.941398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.941398Z digest=sha256:02111736e80579148800a13c6b37403e55022b3a36cfc53e300da8af42974f4d

Observation aba8c202-d633-4d52-b917-2d5fb310d650 · outbound

This paper cites Dery, Corey Staten, Mikhail Khodak, Graham Neubig, and Ameet Talwalkar.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Dery, Corey Staten, Mikhail Khodak, Graham Neubig, and Ameet Talwalkar

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.945786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.945786Z digest=sha256:3e0ca3c20794a2f11443c6fa6ef6dcbf1992bf185cf4f0b119d29a4566c37cfb

Observation 39a910c5-f1bb-4445-b180-046e47fffb87 · outbound

This paper cites Tag-llm: Repurposing general-purpose llms for specialized domains, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Tag-llm: Repurposing general-purpose llms for specialized domains, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.950181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.950181Z digest=sha256:7f0d931bb52b769d5c676c4a491a41e10053ac5630041468db2c86401ad3577c

Observation 4e4c3a1a-8f5e-4dab-8277-e5e7ef31774a · outbound

This paper cites Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Symbiotic Cooperation for Web Agents: Harnessing Complementary Strengths of Large and Small LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.954653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.954653Z digest=sha256:963b2077b136497054b4a030691504fd576d1435961c96e9de153178c19b4040

Observation 70ce8759-e329-4903-ba95-328e5c1928d3 · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.959454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.959454Z digest=sha256:98a2a2398b8928ebd65f01bad15404e18ead874d259dd5552d7500f64e9226b1

Observation c4d4bed2-1749-4bfe-ac01-17c9343ec65b · outbound

This paper cites DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.964683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.964683Z digest=sha256:e4a62cf6e7ad182f4da11a4bb1dfff890c041036ee604d7617e183cb18ea4242

Observation 70197658-5e2a-4f7a-9cbd-a67f16006c06 · outbound

This paper cites Digi-Q: Learning Q-Value Functions for Training Device-Control Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Digi-Q: Learning Q-Value Functions for Training Device-Control Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.969110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.969110Z digest=sha256:cac9be25b86b696f94c042c7b62829f923e7bf2ae2bc4b05f987bc816df05228

Observation a5ef3580-6878-45eb-a35f-ac20acc09bae · outbound

This paper cites an unresolved cited work.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.973768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.973768Z digest=sha256:5aae76aa64dabdbb85935626cfc8fb3780f60dee7251f0d2458b9c4edfdc7ab1

Observation 2e782046-7782-462f-9b1e-06e866525791 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.978039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.978039Z digest=sha256:9639711abe978e10048c4aa332e529fd6e4b274ad2f18a3d0ea57af4b9bc05a2

Observation 33019fc2-cecc-4746-977a-048adb9447fe · outbound

This paper cites Claude takes research to new places, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Claude takes research to new places, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.982960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.982960Z digest=sha256:c46e79c94edef772d73503d39c809b3e19898296b77dde569ca9904f8b7b6883

Observation 7220bd2f-52ce-4fb8-8113-eee5b8cc2c5c · outbound

This paper cites Introducing deep research, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing deep research, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.987539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.987539Z digest=sha256:4adf93ed1fb79f9b65c8faedc6af43c8c257326dcc3e6fde906ee2cc4aa441f3

Observation 29b7f33b-8613-4bcc-b7c5-3272ab79e734 · outbound

This paper cites Gemini deep research, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gemini deep research, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.992631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.992631Z digest=sha256:6f44d2b8ee0f68292c041c05c8f7f369e7208c131fac86156a72cb4aa20808df

Observation d1c6c100-0c31-4f6e-bfb5-bca2489afee1 · outbound

This paper cites s1: Simple test-time scaling.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction s1: Simple test-time scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:49.997734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:49.997734Z digest=sha256:faea1849f088c2bb68a12929812e6072d18819064157bb7ca968438c5d5b002e

Observation 2ad2dafa-1b85-43e4-a524-8efb16084ca3 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.003014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.003014Z digest=sha256:45639150ad33ea82be8cd51f5603b799ccfd2f03163207b999bfc7013e702631

Observation 258a1186-09a0-490a-92f7-cce3cec100c2 · outbound

This paper cites Inference-aware fine- tuning for best-of-n sampling in large language models.ArXiv, abs/2412.15287, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference-aware fine- tuning for best-of-n sampling in large language models.ArXiv, abs/2412.15287, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.008187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.008187Z digest=sha256:688a594addec823fa5fed75d5601e172583f436e3ef66f53715a2b737f6699d9

Observation 7d2fabcf-f09e-47ba-911f-3879eb3b7a72 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.013716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.013716Z digest=sha256:4fb72ae729d483613672339bc850406e9a2241ac778efaf03c7403a6bc9251a6

Observation 5c762179-83b1-41c5-af07-945c628de556 · outbound

This paper cites Webglm: Towards an efficient web-enhanced question answering system with human preferences, 2023.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Webglm: Towards an efficient web-enhanced question answering system with human preferences, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.018456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.018456Z digest=sha256:dd06a9576a8ec9b6cfbc5b7e84d21d7a5cd01a797dd351bc383f92ff302da479

Observation 09ef7536-7bf3-44c5-9e0d-e20627c4c691 · outbound

This paper cites Multimodal web navigation with instruction-finetuned foundation models,.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Multimodal web navigation with instruction-finetuned foundation models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.022921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.022921Z digest=sha256:de5ff59e78aa9dc885b75f28596c099d4492e4d75127ddc68f4a0811f7eb941c

Observation f84bef39-2991-40df-b917-33be9f082d02 · outbound

This paper cites AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.033332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.033332Z digest=sha256:0a88a736b264e21067431ef4113d56254004cb019a9cc157f731124c388dcca8

Observation 4693a926-9a90-412a-855c-ac7ae680c583 · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.027736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.027736Z digest=sha256:35fc1ec15224a97ed5f3188e7e85f71e07731a92c997b7c690d32d6aa7b52fda

Observation d8e73f28-b6c8-4382-92c3-ba05488d486c · outbound

This paper cites UPS: Efficiently building foundation models for PDE solving via cross-modal adaptation.Transactions on Machine Learning Research, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction UPS: Efficiently building foundation models for PDE solving via cross-modal adaptation.Transactions on Machine Learning Research, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.042472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.042472Z digest=sha256:322a2c0d13b07bdf69e5d7fc69afb85d982c993c58845f76ea3415d467f905ed

Observation 44ec6c1c-3e1d-450e-ace3-c8a49551fcd0 · outbound

This paper cites WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.038026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.038026Z digest=sha256:e4dda1e48d0398bab82961ceee1f767c242b80730943aca84731cd97e8e12499

Observation dc420f93-efe5-43e7-ae33-b33d2b5899f3 · outbound

This paper cites AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.052967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.052967Z digest=sha256:3610c046698a4b436d6a3bf7d0b2d293a097cb955098bf64e05b5ce6c496fef9

Observation 210c2e1d-ef26-42b9-b227-58920dcdad78 · outbound

This paper cites CAT: Content-Adaptive Image Tokenization.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction CAT: Content-Adaptive Image Tokenization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.047548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.047548Z digest=sha256:f4d2abe91902218d05df65ee9834b5e22a3ee7befa326f11f9e36fbe5e3ee45e

Observation 31994f1e-3fe7-43c4-8214-0e7b2ae186fb · outbound

This paper cites Specialized Foundation Models Struggle to Beat Supervised Baselines.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Specialized Foundation Models Struggle to Beat Supervised Baselines

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.063572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.063572Z digest=sha256:52c8af23d81162125e8bdbd043f6725832bb7c03d92d917baf32c13815d592d1

Observation 36464d1b-0192-4e76-a387-d80269b47128 · outbound

This paper cites Beyond Browsing: API-Based Web Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Beyond Browsing: API-Based Web Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.058573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.058573Z digest=sha256:4dc6c006b1a4a592ed62c6b44954526b673b029ff83d4e14f8929e82328c1b31

Observation 1673a028-4cea-4639-9415-9c471ed362ed · outbound

This paper cites Codepde: An inference framework for llm-driven pde solver generation, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Codepde: An inference framework for llm-driven pde solver generation, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.072815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.072815Z digest=sha256:5320a27cbab6f0755488f1f8d4a2f0b1cb540fc14532941c8c2808e3cca96b7e

Observation eec5d41f-f833-48d6-b881-eaf8ef0aa3f1 · outbound

This paper cites Mathematicalreconstruction of patient-specific vascular networks based on clinical images and global optimization.IEEE Access, 9:20648–20661, 2021.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mathematicalreconstruction of patient-specific vascular networks based on clinical images and global optimization.IEEE Access, 9:20648–20661, 2021

Reference 42

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:27:51.359828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:50.068419Z digest=sha256:0be45603a8040d727d56b2b0072de7e9570952a9192b8b520e7aebe371d94b9b

Observation 6afcc4d2-0016-432e-894f-dc72072953cb · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Autonomous Evaluation and Refinement of Digital Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.081792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.081792Z digest=sha256:65f7057be3cc00bff29177e02881866918812696b59367f5139182b40dad9ae7

Observation fa1c94fc-bec5-4f24-97dc-837b8fc42c2f · outbound

This paper cites Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, and Yu Su.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, and Yu Su

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.077240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.077240Z digest=sha256:14a9e74fc4828e066554654aa511ff28f7848bfc8ca192e968d6a7e0a05e08ea

Observation 49b1add7-ab27-4766-a8d4-f37558244ef0 · outbound

This paper cites NAS-bench-360: Benchmarking neural architecture search on diverse tasks.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction NAS-bench-360: Benchmarking neural architecture search on diverse tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.091732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.091732Z digest=sha256:b19e9e0463393cd232989213a8a81c8619067e93d09ca43ea847ba6bdffaa282

Observation a3f94752-a38d-4e18-a4c9-34f70b9535a6 · outbound

This paper cites Efficient architecture search for diverse tasks.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Efficient architecture search for diverse tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.086748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.086748Z digest=sha256:f67178e2e989cee03ac2849ae42d96176674e7dfbb2fd464dd85cb2e118eac2a

Observation 2d07b871-658c-438a-bac7-013d87e93ae1 · outbound

This paper cites GPT-4 Technical Report.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction GPT-4 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.101914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.101914Z digest=sha256:4727a2ef11eda6e77978a6aef8bd5b655f5e847870be2c7205274e42c97e9477

Observation 53190f75-8c12-4061-819b-83c5ce2b16e4 · outbound

This paper cites AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.097047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.097047Z digest=sha256:7f6343c1bf698b9bf48668376082e1aec286ce6b6debc7e8a1cbb4f5bd6da345

Observation 1a34aa73-9276-4d55-bce1-999ed8db26ca · outbound

This paper cites SteP: Stacked LLM Policies for Web Actions.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction SteP: Stacked LLM Policies for Web Actions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.112620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.112620Z digest=sha256:ee89289133dcbd0765f8bfb686293c16467c0b7bfc52dd743b4d7d91df4a1408

Observation 8804b0b2-2df3-496f-b15f-b57655461877 · outbound

This paper cites Introducing the next generation of claude, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Introducing the next generation of claude, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.107600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.107600Z digest=sha256:2149fbd99a7dcee94f31576a0bc8292c6738f0d56b26bcdf1f44033d37581b37

Observation 4f9b53d9-5f6b-44ea-921d-8b4f8965159e · outbound

This paper cites L2g: Repurposing language models for genomics tasks.bioRxiv, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction L2g: Repurposing language models for genomics tasks.bioRxiv, 2024

Reference 51

Resolution
verified exact
doi, observed 2026-08-07T05:27:50.519540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:50.122811Z digest=sha256:8bf33160ee162ff1ec8ded4ca30bdc698136b34a074b4a736ffa41da7f60c69d

Observation 54d8756e-8452-4363-af5e-f212a5e09ba6 · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.117392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.117392Z digest=sha256:65d1a9feae75729f9d3d539b20331126bc763c7d0679d34394f8010fdd769c0c

Observation 13a70ffe-bdf9-42cc-820c-2990f5b3de1d · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Training Language Models to Self-Correct via Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.132845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.132845Z digest=sha256:31977ae9a6f6ec353752ea183c6717e2978af33f42d55fd883cbcad206238c1d

Observation aadaef5b-d9b0-4b7e-914c-8d70c7b2e53d · outbound

This paper cites Plan-and-act: Improving planning of agents for long-horizon tasks, 2025.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Plan-and-act: Improving planning of agents for long-horizon tasks, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.127713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.127713Z digest=sha256:995832ff72bfe5f0e419cf4f0626f31753a4e6ac30f6f24095a0a25177627b08

Observation 2b0e1509-7fa7-46fc-bb9c-d5d47fae60f6 · outbound

This paper cites Synapse: Trajectory-as-exemplar prompting with memory for computer control.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Synapse: Trajectory-as-exemplar prompting with memory for computer control

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.141934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.141934Z digest=sha256:2c23bd9e64723a94a0d34c86dfedd87588340c339797e8116c59295423f96687

Observation 685318d4-30dd-4025-b6f6-9a43ff82a03c · outbound

This paper cites Tree search for language model agents.arXiv preprint arXiv:2407.01476, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Tree search for language model agents.arXiv preprint arXiv:2407.01476, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.137448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.137448Z digest=sha256:b97ee95fcc4515ec06a0b8f345fa85e3bb29dd4eed82729c0cd3e1e0cac73fe0

Observation d71d6456-f4ba-494c-ad2b-c75f9fd7a9a8 · outbound

This paper cites Infogent: An Agent-Based Framework for Web Information Aggregation.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Infogent: An Agent-Based Framework for Web Information Aggregation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.151450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.151450Z digest=sha256:9d7970ca54671b460bce7ca2deab3216e7f834b893e22c53cc1b1a3498fd6ed4

Observation ced04181-8e51-4d3f-a178-d46bc6713b93 · outbound

This paper cites Agent Workflow Memory.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent Workflow Memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.146703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.146703Z digest=sha256:326537035ba651d1ca90c3e21d7f53a2f41a2b57681f1c8ef532e9f5f0d6a892

Observation 06f7f655-bb64-4f40-afcd-152e367c51e6 · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.161004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.161004Z digest=sha256:cb247f83a1e6c926d98eaa4587631d3b9af36b0183d57c92d4188e9b6736d188

Observation d456d3ae-fae2-4423-93f2-60eae900fdc8 · outbound

This paper cites BAGEL: Bootstrapping Agents by Guiding Exploration with Language.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction BAGEL: Bootstrapping Agents by Guiding Exploration with Language

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.156194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.156194Z digest=sha256:2364cf27b24a80aeb7e2a776264f38ba2dddb7dc401a0f511e5293a31c9b6bef

Observation d55730ce-385f-4afc-a5e0-314eacca5c59 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.171210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.171210Z digest=sha256:735739518be195770063455f526257728a920684b466b5a49edc1e71eb8e8aca

Observation 1d96973f-8934-484e-96ed-85db71ef9fe8 · outbound

This paper cites InSTA: Towards Internet-Scale Training For Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction InSTA: Towards Internet-Scale Training For Agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.165917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.165917Z digest=sha256:0b4a31416cb38f8aa7353a849d9d55d323572e2b4d9683dc7693da10b367f85c

Observation 676a71e2-dc54-4ff5-a7d5-6368c6efeed1 · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.180946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.180946Z digest=sha256:37ad834c92829e44d89f8343b987a2b171c24f3dc89b0aec04f5bb525e3232f2

Observation 8b248dad-5bae-4169-9705-11936e0f99f3 · outbound

This paper cites Autowebglm: A large language model-based web navigating agent.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Autowebglm: A large language model-based web navigating agent

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.190059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.190059Z digest=sha256:ba2ec62322f222719db4ecade86949d57ee4e49b556f87bad460b4856ccc6ea7

Observation 7d7fb0f8-0099-41bb-b17b-ec095a7d0434 · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.185457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.185457Z digest=sha256:6fd31ac875bf00792586737572cf00288cccecee534e9da7a5041aae88e2d59f

Observation bbffbb6f-11ba-4bf9-ad32-37096acd9e55 · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.199462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.199462Z digest=sha256:23d31dd1aa987f36a8ce16cc9e4af090df714ad3dd1324bc62bcfc4f82d490bc

Observation 1ce53008-52cc-494b-86ab-20ecb88491ac · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.194755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.194755Z digest=sha256:f0d6a5d40c3892e821457af26a0b9096c39f09efccccbae8041ea120333fd106

Observation c0c9bf6e-b80f-4b5b-a185-9ab64d74dab3 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.215014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.215014Z digest=sha256:5c2465af3308f9ef42329e4d5ba656745550b245813ecb207d9dce8db45fd45e

Observation 152980e0-619a-450c-b3ea-1fcf3e11b14d · outbound

This paper cites an unresolved cited work.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.203966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.203966Z digest=sha256:2a8d6fe02cface43ee0001511f68d65772a49e75b1aeae33ff4140b180ba013f

Observation fa309089-4b99-4995-a905-15f4f8aa3190 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Training Verifiers to Solve Math Word Problems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.209389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.209389Z digest=sha256:32cbc0c9f75bf9c967c819cdc11a1634634db8e66d1530de57531421e0d709b2

Observation 8d2ad100-cff4-41f3-bf36-373daa431c22 · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.229084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.229084Z digest=sha256:a9860c91347315d2bf8314e16a2449457ef7f82f7b61138c856c98fe358e3fe4

Observation f20f4c7f-9d7b-48aa-9f8d-bad1bab99af3 · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.219788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.219788Z digest=sha256:0c2481fda915c7e95b2e5226a150801ad671907d6b0989f55606947862324ea3

Observation 091291ed-fe60-426f-8e65-4cd840c0a7cb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.224504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.224504Z digest=sha256:8c9b10cb58ab7ddaec68a6954bbf93df749e6eb37cf5b6e2f15b3bdd638a9ab1

Observation 851db09b-90e5-403f-b4e4-5bf910310f59 · outbound

This paper cites ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.242692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.242692Z digest=sha256:740fb82fcb306ecb6f8ec7e66469baf11c725f8ee0270ddc73dea0c69e01b191

Observation cd30e87b-425d-466d-9efe-fae298b462b5 · outbound

This paper cites Doing: Agents that Reason by Scaling Test-Time Interaction Li Fei-Fei, Lijuan Wang, Yejin Choi, and Manling Li.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Doing: Agents that Reason by Scaling Test-Time Interaction Li Fei-Fei, Lijuan Wang, Yejin Choi, and Manling Li

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.233594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.233594Z digest=sha256:bffd9138339029b9fd57100f683ad6ef424b90602651a993225c029562bd274b

Observation 4440afe6-8dfc-4d5f-abc1-c2c0225e657a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ReAct: Synergizing Reasoning and Acting in Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.238031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.238031Z digest=sha256:def71ad070725f4e007b03ebd9758bc295163617b808012962439dad6a756f91

Observation 3ffb19f0-0140-48e7-8833-c767f0609863 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.261014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.261014Z digest=sha256:bdf500e049d18ce57e125800273849826af8ce4930ab4ede80f5e86c117340a9

Observation f797f3bc-1f46-4808-baba-a0518dc4ac7a · outbound

This paper cites an unresolved cited work.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.248044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.248044Z digest=sha256:2a6d93ec0d51e9435d4d211676ab43f7e9b18d16f11dc1f7a9feed940d8cb67d

Observation 7decc3f5-3045-4f9f-83c9-60a869edd0d9 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.252400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.252400Z digest=sha256:6064e7d43b6272d2a643233362ebb90b0690a6836ddc5cd1856e21f656879af5

Observation f86bc6b7-2cb0-4312-a672-33b57d87547b · outbound

This paper cites Metaxas, and Tong Che.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Metaxas, and Tong Che

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.256725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.256725Z digest=sha256:055a8105f26b3b408ece3d68a273350aa26c458106bf5b84f14e67e766451414

Observation 77754e79-1cbf-4898-8201-c5dc351d6c47 · outbound

This paper cites Gemma 3 Technical Report.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gemma 3 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.279407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.279407Z digest=sha256:e1d26b4c1735ab0e31d40ca748bcb494e025a70ebb8753d2b5513414d54922be

Observation 3e75664d-242b-41ec-8893-3ca685a33ef3 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.265548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.265548Z digest=sha256:011e4fde9ae8870e142dbf96f9ccc50c2c98e2e54703d373cf5b99cdb80a596b

Observation 220a18c6-fc7e-4132-af06-b6ec616c7f4b · outbound

This paper cites Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:27:50.656331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:50.270831Z digest=sha256:38f078ce0de729aa1329c9cc20712993f5c99af128fa58ef337cca2995dde53e

Observation 0ce756ea-20d6-49bd-aad0-b5ff70f3b8a7 · outbound

This paper cites Inducingprogrammaticskills for agentic tasks.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Inducingprogrammaticskills for agentic tasks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.275181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.275181Z digest=sha256:d0696d2b1647e2de5aaf80c3a1e1a30c6e3a47f55b3679015e1bb6a563932c54

Observation ed329a50-5655-44a7-b005-4169833f405b · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.296869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.296869Z digest=sha256:b4897fa2b0634d990f3450b8a59e55c142bec38c5628ce7b0c58406ef1a7407c

Observation b390b4ba-9c37-45ec-80a1-172b04fe4507 · outbound

This paper cites Recursive introspection: Teaching language model agents how to self-improve.Advances in Neural Information Processing Systems, 37:55249–55285, 2024.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Recursive introspection: Teaching language model agents how to self-improve.Advances in Neural Information Processing Systems, 37:55249–55285, 2024

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.284269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.284269Z digest=sha256:81c0a50596817228f811e8326dc229f47db3bc70ea90db32e095fab36b9e2a7f

Observation ce0126c2-e4f5-4c92-9d46-635f7f59f1e8 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.288515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.288515Z digest=sha256:8fbcd9979157c55b9e6690c40215ac6f5a46f93c3d5c5c1eaffbd85b54ad5503

Observation f7697213-ceef-4bd7-8e44-b5797144cee3 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488, 2022.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488, 2022

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.292630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.292630Z digest=sha256:745681c15b63ce4a100d1bcad703fdd3650bd30ba9b233cfbec041458633b9c4

Observation 7d2e72d6-2e18-4952-8cb9-b2e854f50841 · outbound

This paper cites Curriculum learning.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Curriculum learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.315759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.315759Z digest=sha256:6a8977150ad2101e81f37ead48b0d912b12d2d6614935d475f604c66816c3872

Observation cbd4f959-4e53-4ac8-83e4-5f27f0f4ae15 · outbound

This paper cites Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.301153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.301153Z digest=sha256:d006f4c8424e610e1131e316eb1625a1d6fcaef732422e3889d2cb18a490a40c

Observation 6d2b60f5-332e-4c81-b660-39eed0222f2d · outbound

This paper cites Analysis and improvement of policy gradient estimation.Neural networks : the official journal of the International Neural Network Society, 26:118–29, 2011.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Analysis and improvement of policy gradient estimation.Neural networks : the official journal of the International Neural Network Society, 26:118–29, 2011

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.306536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.306536Z digest=sha256:d8ff48a7fcdf14284f2c1f7ab433b6ca011618c45210dc8b0b681af0417b995b

Observation 609c6309-f788-4fb4-a157-f8c10543abb8 · outbound

This paper cites Policy gradients with variance related risk criteria.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Policy gradients with variance related risk criteria

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.311083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.311083Z digest=sha256:f3f54cd0202b6e2438250d883d3455af71919a908aadfe78d2f67f46880c79c9

Observation 9f767eee-302f-4806-9f1e-e76b6adfd824 · outbound

This paper cites Abbeel, and Wojciech Zaremba.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Abbeel, and Wojciech Zaremba

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.337973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.337973Z digest=sha256:7d6a4aa33c3d2be4f19481698e374e1672255aab7c8cf7971eea810185209307

Observation 172aac69-c4d4-48ae-b132-497d208d9d69 · outbound

This paper cites Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.321017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.321017Z digest=sha256:c9404f653b5410912b413cf1b38387b0d98501d1686300a8d936047633640bdd

Observation cafc7a9e-4033-4ceb-9c39-b632f5253095 · outbound

This paper cites A survey on curriculum learning.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 44:4555–4576, 2021.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction A survey on curriculum learning.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 44:4555–4576, 2021

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.325733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.325733Z digest=sha256:6ac6ba0591022566e92fd566a1fcf25cd735973cf369b709f7c7ceeb78f35341

Observation b1f9382b-0b5b-4fd6-91eb-98eda6e4019d · outbound

This paper cites Teacher–student curriculum learning.IEEE Transactions on Neural Networks and Learning Systems, 31:3732–3740, 2017.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Teacher–student curriculum learning.IEEE Transactions on Neural Networks and Learning Systems, 31:3732–3740, 2017

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.333188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.333188Z digest=sha256:59a2394d5929344b0f2c32d9d9ae553aadd155363044f7166b017633534a7273

Observation e863c26d-93fc-486a-968c-e70c501dd59d · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.356815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.356815Z digest=sha256:80415a9320200087e8d53023f9e39b0201494f2c143d23895260c5411f3fb1ea

Observation 2fd4ebbe-7c12-42b7-a678-d5fcd0503597 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Gonzalez, Hao Zhang, and Ion Stoica

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.342750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.342750Z digest=sha256:394104bba0a02c32d4ef69f5a72d61fe0c495be356c926838d41b207afae4fbd

Observation 44ee4768-89ff-4c53-a214-2e799dc215f0 · outbound

This paper cites an unresolved cited work.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:27:52.029289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:50.347566Z digest=sha256:00cc41009c548540ac1cf8180db6ded89c744b73b38df01d711b5386ab81d017

Observation e80e1f36-d90b-427f-af5d-37cd4c7e0990 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.352214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.352214Z digest=sha256:8980fdc6871dde889f0110d500d4508a2a648268f65dbe2492483e6baa761e24

Pith citing papers

Observation 254e48c4-098f-4121-97f6-fc85d6ece7e4 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:10.825837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:e548f5975b4ccef5208dd0f4e844524d93f63b3ed28ac29c1dbbf0508b442055

Observation 78856db5-fe51-411c-96bf-e5a93e222d04 · inbound

SSRL: Self-Search Reinforcement Learning cites this paper.

SSRL: Self-Search Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:11.770737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:17:11.770737Z digest=sha256:241df56f6bef031244b305126ae9bc5d84076c9ccfdb39f9b364c6e033dc79e2

Observation 1cf96ff8-46c8-4bb8-baa2-5b44b270689c · inbound

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents cites this paper.

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:06:42.761774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:06:12.349285Z digest=sha256:f4f303235d23f1b162ed2dfac12fc5bfb830eadb0cc9fa2177e3a919999e2080

Observation 946e1a85-769b-4a5d-8d47-a77f080b9904 · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.987463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:034acba64d5b1529af5819333a11f0af1e815d742a53955d5314aeff67bd9083

Observation b221f6b0-b3c2-4192-9046-202a2fbf1df0 · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.874836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:6c13164e893e694df4186326faaf687e590c4ec944e7e7d1edb0651cfe5a2631

Observation bae768d1-b8db-4710-b31f-14a59cad3690 · inbound

PRO-CUA: Process-Reward Optimization for Computer Use Agents cites this paper.

PRO-CUA: Process-Reward Optimization for Computer Use Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:24.210583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T11:46:38.569528Z digest=sha256:88c04fe539b4ab88c6559fdbb40c1a1f5395664dc46dbeb14ddf5e5b3e0b8bec

Observation 2b97dea4-7a53-4863-84ca-2707b6338858 · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.499801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:16cb542283e47ff753a33ba2e79255cc58a36815ea4695a92f8f9d90416cb8b0

Observation 29cf167a-bdf0-43fa-ae28-3af81c202ecb · inbound

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments cites this paper.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:f443820cccf032db9cf488c82cd5682c9023821fe89a82ee86766add5e6f9c2a